Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260822b-echo-threshold/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260822b-echo-threshold
statusfrozen
created2026-08-22
updated2026-08-22
sensesstyle-correspondence
provisionaltrue
internal-judgment-onlytrue
linkswiki/arms/ARM-echo-threshold.md, wiki/findings/results/RS-20260822-synonym-reach.md, workshop/translations/gulistan/R41-v1/translation.md, workshop/translations/gulistan/R42-v1/translation.md, tools/rhyme_pairs.py, config/models.md, config/budget.md, framework/v0.2/README.md

E-20260822b — how much sound-relation does an English pair need before it is heard as an echo

ARM-echo-threshold step 1, track T3. Frozen before dispatch. Nothing below is revised after the first call; amendments accepted from the pre-run critic are recorded in critic-response.md with the commit that carried them, and the frozen text stands as written.

1. The question, and why it is not a question about our instruments

RS-20260822-synonym-reach §4 printed seven English pairs that a mechanical near-echo rule counts as chimes — spread / tend, creatures / entities, scattered / suspended, put on / set, eat / are bought, does not persist / is inappropriate, delight / comfort — and asserted, in the lead's voice, with nothing behind it:

No reader hears any of them as a chime.

That sentence is load-bearing. framework/v0.2 §7.25 item 3 tells a practitioner, on the strength of it, that the relaxed form of the enumeration instruction points at a property of English rather than at his author. The question here is whether it is true, and the unit of analysis is a reader's report about a passage of English, not a property of tools/rhyme_pairs.py, whose arithmetic is not in doubt and which is used here to define the independent variable, not to measure anything.

What this teaches about translating literature: at what degree of phonetic relation an English pair standing at two phrase-ends is recoverable as an echo by a competent reading instrument that is looking for one — which is what decides whether a translator answering a source rhyme with a slant is carrying his author's music across or decorating.

2. The estimand, stated so it cannot be overclaimed later

Detectability of a planted phrase-end sound relation by three non-Anthropic reading seats whose attention has been directed at sound.

The word reader does not appear in any prediction, criterion or consequence below (critic BLOCKING 1, amendment A1). The seats are models; every figure is a figure about models reading English, and the handbook consequence in §6 carries its warrant on its face.

Three things this is not, said now rather than in the limits:

  1. It is not what an unprompted seat picks up. Stage U measures that rather than assuming it (critic MAJOR 4, amendment A4): the original text of this section claimed the prompted figures were an "upper bound" on ordinary reading, which was an assumption dressed as a measurement. The claim is withdrawn and the quantity is bought.
  2. It is not a human measurement. No human readers exist to buy (NEXT.md, named-not-built) and the three seats are models. Every item and every seat verdict is printed in the result page's appendix so that a human reader can check them, which is the most the project can do here.
  3. It is not a claim about any translation's quality. No seat is asked to judge anything.

3. Materials

138 items over 25 loci, in a 2 × 3 crossover (critic BLOCKING 2, amendment A2). Each locus is one short English passage with two parallel phrase-ends. One end takes the fixed word, the other the filler; the variants of a locus differ in exactly one word from their row- and column-neighbours.

Each locus supplies two fixed words and two or three fillers:

filler X filler Y filler Z
fixed a NEAR — the only slant cell NONE STRICT — the only full cell
fixed b NONE NONE NONE

b is a second ordinary word for the same slot, chosen so that tools/rhyme_pairs.py scores it NONE against every filler at that locus. All levels are assigned by the tool and verified by the verifier before dispatch.

Why the crossover is the design and not a refinement. In a plain three-level comparison the phonetic level travels with the lexical identity of the filler — its meaning, frequency, register, collocation, morphology and syntactic fit — and no naturalness control can separate them. Here filler X stands once in a NEAR cell and once in a NONE cell; filler Y stands in two NONE cells; fixed words a and b each stand in every column. Any main effect of either word cancels in the interaction, and the phonetic relation is present in exactly one of the four core cells.

19 loci carry all six cells; six loci (PN1–PN6) carry four, because STRICT is unreachable there from any ordinary rendering — which is S211's own headline — and their X/Y fillers are drawn whole from the S211 blind-seat pool, including the three pairs RS-20260822 §4 named (scattered / suspended, creatures / entities, spread / tend).

Counts: NEAR 25 cells, STRICT 19, NONE 94. By class: N1 ×5, N2 ×17, N3 ×3.

3.1 Strata

3.2 The items

The full manifest is materials/items.json — every carrier, both fixed words, every filler, every rule verdict, the bearers the rule used, the phrase-end word list, the unplanted-echo screen and the pool provenance flag — and it is printed whole in the result page's appendix so that a human can check any cell. Two loci, given here so the shape is legible:

PR4 (PROSE) — "The one comes {fixed} to artifice; the other stands {filler}." fixed a = near, b = close; fillers X = far (N2), Y = remote, Z = clear (STRICT).

VE10 (VERSE) — "What is the tongue in the mouth, O wise man? / The key to the treasure-door of a man of {fixed}; / and while the door stays shut, who is to know / whether he trades in jewels or in {filler}?" fixed a = skill, b = craft; fillers X = wool (N2), Y = silk, Z = twill (STRICT).

VE2 is the one locus whose filler slot is at the first phrase-end rather than the second; nothing depends on which end varies, and it is named so the asymmetry is on the record.

3.3 The NEAR classes are not balanced, and that is declared here

N1 ×5, N2 ×17, N3 ×3. The imbalance is structural, not a choice: N1 arises almost only between polysyllables, and STRICT between polysyllables is nearly unreachable — S211's headline again. The run measures NEAR as a class, with a per-class breakdown reported and no class carrying a claim on its own.

3.4 Unplanted phrase-end relations, screened mechanically

Critic MAJOR 5, amendment A5. Every passage is segmented at line breaks and at ; : , . ? !, the last word of each segment taken, and every pair of those words scored by the rule.

14 of 138 items, spanning 5 of 25 loci — PR6, PN3, VE1, VE8, VE11 — carry an unplanted phrase-end relation. Every one is NEAR; not one is STRICT.

A registered sensitivity analysis recomputes every criterion with those five loci dropped, and both figures are reported side by side.

4. Procedure

Five stages. All calls temperature 0, one item per call, isolated context, dispatched in one shuffled stream per stage with seed 20260822.

Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, QR qwen/qwen3.7-max. P3 is excluded on the cost finding of S211 (NEXT.md), P4 on note (bps), P5 on note (bne), GL on LONG prompts. There is no fourth reading seat to buy and the result will say so rather than implying a panel.

Stage D — prompted detection (138 items × 3 seats = 414 calls)

The seat sees the passage and nothing else. No mention of Persian, of Sa'di, of translation, of rhyme regimes, or that variants exist.

Read this passage.

<PASSAGE>

Does the passage contain a sound echo -- two words that chime with each other in
sound -- standing at the ends of its phrases or lines?

Answer in exactly this form and add nothing else.
ECHO: yes
PAIR: <word> / <word>
STRENGTH: full
-- or --
ECHO: no

If you hear more than one such pair, give at most two, separated by a semicolon.
STRENGTH is "full" or "half".

max_tokens 300, reasoning cap 100. The yes branch of the template is shown first, which can only raise the yes rate; it is identical for every item, so it moves the floor and the cells together and cancels in the interaction. Failure criterion 4 guards the case where it saturates.

Stage U — unprompted, added on critic MAJOR 4 (57 items × 1 seat = 57 calls)

The aZ, aX and aY cells of the 19 three-level loci, to P1. The word sound does not occur in this prompt, and neither does echo, rhyme or chime.

Here is a short passage.

<PASSAGE>

Is there anything you notice about how it is written? Answer in one sentence,
then stop.

Coded mechanically: does the sentence name both planted words? Secondary coding: does it use any of rhyme / chime / echo / assonance / alliteration / sound? max_tokens 400, reasoning cap 120.

Stage R — naturalness ranking (25 loci × 2 seats = 50 calls)

The four core cells of each locus (aX, aY, bX, bY) shown together, letters assigned by a per-(locus, seat) shuffle from the seed.

Below are 4 versions of the same short passage.

A) ...
B) ...
C) ...
D) ...

Rank them by how natural and ordinary the English reads, best first.

Answer in exactly this form and add nothing else, using each letter once.
RANK: <letter> > <letter> > <letter> > <letter>

Seats P1 and P2. The answer template names no concrete order, and the letter→cell map is shuffled per (locus, seat), so any residual letter preference spreads over cells rather than onto one of them. The Z cells are not ranked; M1 is a gate, not a finding, and that is declared.

Stage S — fluency screen (138 variants + 4 planted breaks = 142 calls, seat P2)

Is the following a fluent, ordinary piece of English that makes sense?

<PASSAGE>

Answer in exactly this form and add nothing else.
FLUENT: yes
-- or --
FLUENT: no
REASON: <at most twelve words>

The 4 planted items replace the filler with a word that cannot stand in the slot — a syntactic break, not a semantic oddity.

5. Scoring, fixed before dispatch

Per (item, seat):

Unparseable or empty → one re-dispatch, then dropped and counted as dead. Dead cells are reported and are not imputed.

6. Predictions and criteria, registered

Write hit(f g) for the mean hit over (locus, seat) cells with fixed word f and filler g. The primary quantities are interactions, so that any main effect of either word cancels (amendment A2):

Δ_N = [hit(aX) − hit(aY)] − [hit(bX) − hit(bY)] Δ_F = [hit(aZ) − hit(aY)] − [hit(bZ) − hit(bY)]

Both are computed on the 20 loci that survive the unplanted-echo screen (amendment A9, critic round 2 MAJOR 3 — the round-1 refusal gave a false reason and the critic caught it). Δ_F uses the 15 of those that carry a Z cell. The 25-locus and 19-locus figures are reported beside them, and if the two disagree about which criterion fires, the 20-locus figure governs and the disagreement is the headline.

Uncertainty (amendment A3, narrowed by A10). The primary statement is a bootstrap over loci — 10,000 resamples with replacement, seed 20260822, percentile interval on Δ. Loci are the resampled unit; seats and cells are not.

What that interval is (A10). The loci are constructed, not sampled from a population of English phrase-ends, so the interval describes instability over these loci and nothing wider. Every conclusion under P1 and F1 is stated about these constructed loci and these pairs; a wider reading is marked in the text as an inference, never as a measurement.

A within-locus label-reshuffle reference distribution (20,000 draws, same seed) is also computed and is reported under that name. It is not a randomisation p-value: the levels are determined by the words, not assigned at random, so no permutation of them can be one. No criterion above depends on it.

7. Controls and their criteria

8. Failure criteria — what would make this run worthless

Written before dispatch, per charter §6:

  1. M1 fails. Three seats cannot recover a full rhyme at a phrase-end above the floor → the task does not measure what it is named for and nothing else is reported as a finding.
  2. C1 fires on the aX cell. The primary is disclosed as possibly confounded.
  3. Dead-cell rate above 10% after the single re-dispatch → reported as a pilot.
  4. The floor is at ceiling. If hit(bX) or hit(aY) exceeds 0.60, the seats are naming the two phrase-end words whatever they are, there is no floor left, and P1 and F1 are both withheld.
  5. A stratum collapses. If C2 drops more than 5 of 25 loci, the run is underpowered and F1 may not fire.
  6. The crossover degenerates. If hit(bX) > hit(aX) and hit(bY) > hit(aY), the fixed word is driving everything and the interaction is not interpretable as a sound effect; reported as such.

9. Budget

Pre-flight built from max_tokens plus the reasoning cap, not from an assumed output length — note (abc), the one estimate this project has overrun failed exactly there.

stage calls cap (out) worst-case
stage D, 138 items × 3 seats 414 400 $0.91
stage P, 12 probe items × 3 seats (amendment A6) 36 400 $0.08
stage U, 57 items × 1 seat (P1) 57 520 $0.19
stage R, 25 loci × 2 seats 50 400 $0.11
stage S, 142 items × 1 seat (P2) 142 400 $0.24
pre-run critic, 2 rounds 2 8500 $0.13
re-dispatch allowance, 15% at the doubled cap ~100 — $0.22
total ~800 $1.88

Declared experiment ceiling $1.90. Runner ceiling $1.55. Stop-loss $1.35 — the runner halts and writes what it has if billed cost passes the stop-loss.

Expected actual, from S211's measured per-call rates (P1 $0.00127, P2 $0.00079): about $0.85. The ceiling is a guard, not a forecast.

Today's UTC ledger stands at $1.067599850 of $5.00 after S211, leaving $3.932400150. The ceiling fits and leaves $2.03.

QR has not been dispatched by this project since selection and its billed rate is unknown. Per note (bqk) the runner prices it from its own first twenty stage-D calls and, if the measured rate would carry the run past the runner ceiling, drops it and completes on P1+P2 — with the arithmetic written down at the moment of the change and nothing already bought re-bought. That decision is made before any stage-U, R or S cell exists, so no outcome can be steered by it.

10. What is deliberately not measured

11. The amendments this design carries

Round 1 of the pre-run critic returned NEEDS-REDESIGN, 5 findings, 2 BLOCKING, before any data call. All five are accepted; critic-response.md carries the reasoning and the one refusal.

# severity amendment
A1 BLOCKING reader struck from every prediction and consequence; the handbook change, if F1 fires, carries its warrant on its face
A2 BLOCKING the design becomes a 2 × 3 crossover and the primary becomes an interaction; 69 items → 138
A3 MAJOR the permutation p-value is demoted to a named reference distribution; bootstrap over loci becomes the uncertainty statement
A4 MAJOR stage U added — the unprompted rate is bought instead of assumed; "upper bound" withdrawn
A5 MAJOR every passage screened for unplanted phrase-end relations

Round 2 also returned NEEDS-REDESIGN, 5 findings, 2 BLOCKING — and buying it was right: it found the confound that would have made the whole run about spelling, and it caught a false statement in one of round 1's refusals.

# severity amendment
A6 BLOCKING stage P, the orthography/phonology probe — 12 items, 36 calls, with a registered veto (P7) over the phonetic reading of every figure in the run
A7 BLOCKING shared final-letter run recorded on all 138 items as a covariate; hit within the NONE cells reported against it
A8 BLOCKING per-locus Δ and a sign test across loci (P8). The critic's remedy — comparable lexical pairings in both NEAR and NONE cells — is refused in writing: no word pair can be simultaneously a rhyme and not a rhyme, so the confound is a property of the question. See §12.
A9 MAJOR the primary moves to the 20 unflagged loci; the round-1 refusal's stated ground was factually wrong and the critic corrected it
A10 MAJOR Δ and its interval declared descriptive of these constructed loci; no class-level threshold claim
A11 MAJOR P3 split by rescoring every named pair; only rule-NONE pairs are called manufactured

12. The one thing no version of this design can fix

The sound relation is present in exactly one cell, and that cell is also the only one containing that particular pair of words. No stimulus can separate them, because a pair cannot be simultaneously a rhyme and not a rhyme. The crossover removes every main effect of either word; what remains is the interaction of the two, and a sound relation is one kind of interaction between two words.

What the design substitutes for a fix is replication across lexis. Twenty loci, twenty unrelated pairs of English words, nothing in common but the planted relation. If Δ is carried by one or two loci it is a fact about those words; if it is broad and consistent in sign, the alternative explanation needs twenty independent coincidences. P8 reports exactly that, and the result page states the limit in these words rather than implying the crossover disposed of it.