Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260823-run-depth/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260823-run-depth
statusfrozen
created2026-08-23
updated2026-08-23
sensesstyle-correspondence
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-run-depth.md, workshop/regimes/R44-run-depth.md, workshop/translations/gulistan-bab1b/passages-frozen.md, workshop/translations/gulistan-bab1b/collation.md, wiki/findings/results/RS-20260822b-echo-threshold.md, wiki/findings/results/RS-20260822c-persian-hands.md, config/models.md, config/budget.md, tools/rhyme_pairs.py

E-20260823-run-depth — the same rhyme at two distances

Frozen before any English rendering exists, before any published hand is opened for this span, and before any API call is dispatched. ARM-run-depth step 1, track T2.

1. The question

Sa'di uses two verse forms inside the same tale. The مثنوی bayt rhymes with itself: in English, with one line to a hemistich, that is a rhyme at adjacent line-ends. The قطعه holds one rhyme across the ends of three or four consecutive bayts and leaves the lines between them unrhymed: in English that is a rhyme at every second line-end. The published English tradition replaces the second with the first — it rhymes each bayt within itself — which supplies more rhyme than the Persian has and puts none of it where the Persian put it.

Is a rhyme held at every second line-end registered by a reader at the rate an adjacent rhyme is? And at the places where Sa'di holds one rhyme three and four bayts deep, does any published English hand hold it at all?

What this is not. It is not a judgment of any rendering: no seat is asked whether anything is good. It is not a claim about human readers; the seats are models and every figure is a figure about models reading. It is not a test of tools/rhyme_pairs.py, which supplies the ground truth by a rule fixed on 2026-08-22 and is not under test here.

2. Why it is worth a session

wiki/goodness-senses.md §style-correspondence is evidenced on device presence (+3.524 for sixteen formal devices carried rather than flattened) and has no evidence about device extent. framework/v0.2 §7.26 measured that a chime is registered at +0.542 at a verse line-end and +0.222 at a prose colon-end — a difference of position. Nothing in the project has measured a difference of distance, and run length is the property R42's own log flagged and R42 rule 4 forbade the hand from attempting.

3. Materials

Source. «گلستان» باب اول حکایات ۶–۱۳, workshop/translations/gulistan-bab1b/source-ganjoor-bab1-h6-13.txt, 59 blocks, 19 prose and 40 bayts. Copy-text and its single-witness limit: workshop/translations/gulistan-bab1b/collation.md.

Structure, frozen from the Persian before any English existed. workshop/translations/gulistan-bab1b/passages-frozen.md — 21 passages, 17 adjacent rhymes and 14 distance links across 9 قطعه runs, two of them four deep. Exclusions E1–E3 are on that page and are not revisited here.

Contamination, declared in advance and high. All four published hands are public domain and certainly in the lead's training data; the Gulistan is among the most translated books in the language. The primary of this study does not depend on the lead's independence: it is a within-lead, within-passage contrast between two renderings the lead makes himself, one under R44 and one under the tradition's rule, and the seats never see a published hand. The published census (§5, P5) is a count of what four printed books do and the lead's column is not pooled with them. Overlap against all four hands is measured with tools/dependence_check.py after the log is frozen and is reported on the translation page whatever it says. The one direction in which recall could manufacture a result is ruled out by the prediction itself: P4 predicts that no published hand carries the distance rhyme, so there is no published deep-run rendering for the lead to recall.

One priming is declared on the passages page: the lead has read حکایت ۱۰'s بنی آدم bayts (V12) in English many times. V12 is a MATHNAWI passage contributing three adjacent rhymes and no distance link — it falls entirely on the control class.

4. Procedure

Stage A — the lead's rendering, free, no API

حکایات ۶–۱۳ whole under R44 (workshop/regimes/R44-run-depth.md): prose as prose, verse as verse, and the English rhyme placed where the Persian rhyme is, at the distance the Persian put it. Every passage graded HELD / SHORT / REFUSED / UNREACHABLE. The log is frozen and committed before stage B exists and before any hand is opened.

Stage B — the two counterfactual arms, free, no API

NEAR-PAIR / FAR-PAIR — the matched minimal pair, and the arm the primary is measured on (amendments A1 and A6). A1's whole-passage permutation was itself confounded — the round-2 critic's BLOCKING 1: putting every rhyming line together turns an alternating pattern into a contiguous block, which changes cluster structure and coherence as well as distance. A6 replaces it with a stimulus in which those cannot vary.

A6's four-line pair was itself confounded — the round-3 critic's BLOCKING 1: in FAR the lines y₁ x₁ are the two hemistichs of one bayt and stay semantically continuous, and in NEAR neither bayt survives. That difference biases towards the registered prediction, so it could not be declared and kept. Amendment A9 removes it by giving the target pair no bayt-mate at all.

Each of the 14 distance links is one item, and each is presented as a six-line stimulus built from the same six lines in both conditions:

condition order presented rhyme partners at
FAR f₁ y₁ f₂ y₂ f₃ f₄ positions 2 and 4 — distance 2
NEAR f₁ y₁ y₂ f₂ f₃ f₄ positions 2 and 3 — distance 1

y₁ and y₂ are the two rhyme-bearing English lines of the link — the lines that render the two Persian bayt-ends the قطعه holds one rhyme across. f₁–f₄ are filler lines drawn from elsewhere in this same span's R44 rendering, so they are the same hand, the same register and the same source, and they belong to neither of the item's two bayts.

The two conditions differ by a single transposition of two adjacent lines — positions 3 and 4 swap. Held identical: the six lines and their words, the line count, y₁'s serial position, the first and last two positions, the rhyme density (one pair), the cluster size (2), and — the round-3 finding — the number of within-bayt adjacencies, which is zero in both, because no filler is a hemistich of either of the item's bayts. Neither target is the final line, so the recency asymmetry A6 carried is gone too.

Filler rule, fixed here and executed by the script, not chosen by hand. The filler pool is every unrhymed first-hemistich line (x) of the span's QITA passages. For item i, fillers are taken from the pool in order starting at offset 4i (mod pool size), skipping any line that (a) belongs to either of the item's own bayts, (b) is graded above NONE against y₁, y₂ or an already-chosen filler by tools/rhyme_pairs.py, or (c) is already used in this item. So each stimulus contains exactly one rhyming pair and the tool says so before dispatch.

What A9 costs, declared. A six-line stimulus of one couplet's worth of sense and four unrelated lines is not verse anybody would print. It is a psychophysical stimulus and the result page calls it one. The literary reading is carried by S2, where the passages are whole and in order.

COUPLET — the ecological arm, secondary. For each of the 9 QITA passages, a second rendering of the same bayts under the tradition's rule: each bayt rhymed within itself, couplets, and no rhyme at any distance-2 position. Same content, same line count, same hand — but different words and twice the rhyme density, which is exactly why it cannot carry the primary (round-1 critic, BLOCKING 1).

Order and what it costs, declared. Stage A is written first and frozen first, so the R44 rendering cannot be anchored on the couplet one. The couplet rendering is therefore a re-rendering of content the hand has already put into English, and may inherit wording — a limit the result page states. It is the conservative direction: shared wording makes the two arms more alike, not less. Every couplet passage is graded mechanically before dispatch and any accidental distance-2 rhyme is re-rendered, because a couplet arm that also carries the run is not the tradition's move.

Stage S — the registration measurement (paid)

Two sub-stages, one instrument, one prompt:

The seat sees numbered English lines and nothing else. No mention of Persian, of Sa'di, of translation, of rhyme placement, of this study, and no indication that any two bodies are related. The instruction is fixed:

Below are numbered lines of English verse. Look at the last word of each line. Give every line a letter. Two lines get the same letter when their last words rhyme — judge by how the words sound, not by how they are spelled. A line whose last word rhymes with no other line in the passage gets a letter of its own. Answer with one line per input line, in the form 3: b, and nothing else.

(Amendment A2: the round-1 critic's MAJOR 2 — the instruction originally read "chime, fully or nearly" while the scoring kept STRICT only. A7, round-2 MAJOR 2: the instruction still does not operationalise STRICT, and cannot without teaching the seat the scoring rule. What answers the objection is that NEAR and FAR contain the identical words, so whatever private threshold a seat uses is the same in both conditions and cancels in the paired difference. It does not cancel in S2, which is one more reason S2 is secondary. The clause judge by how the words sound, not by how they are spelled is added on note (bqy).)

Scoring. For every unordered pair of lines at distance 1 or 2 as presented, the seat links the pair if it gave both lines the same letter. Ground truth for the pair is tools/rhyme_pairs.py on the two line-final expressions, computed before dispatch and not revised after:

class definition
A+ distance 1 as presented, tool says STRICT
D+ distance 2 as presented, tool says STRICT
A− distance 1 as presented, tool says NONE
D− distance 2 as presented, tool says NONE

Pairs the tool grades NEAR-*, IDENTICAL or UNKNOWN are excluded from S2's classes and reported separately; the NEAR-inclusive recomputation is reported as a sensitivity analysis on every figure (amendment A3). In S1 the target pair is the same words in both conditions whatever the tool grades it, so S1's primary keeps every link whose tool grade is STRICT and reports the NEAR and UNKNOWN links separately. Dispatch order is a single shuffle of all 174 calls from seed 20260823.

On transitivity (amendment A3, round-1 critic MAJOR 2, second half). A letter is an equivalence class and a phonetic relation is not, so a NEAR pair inside a passage can pull a STRICT pair's letters around. Between NEAR and FAR the words are identical, so any such artefact is present in both conditions in the same amount and cancels in the primary difference. It does not cancel in S2, which is a further reason S2 is secondary.

The primary statistic, and why it is not a pooled rate (amendment A8, round-2 critic MAJOR 3). Links inside one run are not independent — a seat that gives one letter to a four-member set produces several pairwise hits mechanically — and three model seats are three calls to mutable services, not a sample of readers. So:

Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. P4 is out on note (bps), P5 on note (bne), GL on LONG prompts, and QR is excluded by name on note (bqy) — a seat that reads spelling rather than sound must not sit on a design whose whole question is where a sound falls. P3 is a cost problem and is priced from its own first twelve calls before the rest of the stage is dispatched (note (bqk)); if its measured rate breaks the arithmetic it is dropped and the result reports two seats. There is no fourth reading seat to buy and the result will say so rather than implying a panel.

Stage F — the addition check on the forced renderings (paid)

A four-member rhyme set is where a translator pads. 18 items — the 9 QITA passages in both arms — plus 6 planted items, go to two seats (P1, P2) with the Persian blocks themselves as the standard and the question: does the English state anything the Persian does not? 48 calls.

The six plants are the same passage with one added modifier, intensifier or evaluation of at most three words — the shape an addition takes when a translator buys a rhyme — constructed and frozen before dispatch, never a new event or character. If the seats do not catch at least 4 of the 6 plants pooled, stage F clears nothing, the R44 grades are reported as self-certified, and the result page says so.

What a stage-F clearance means, at its true strength: no addition was detected by two seats reading the Persian, on an instrument shown to catch four of six three-word plants. It is not a certification by an established Persian reader; the project has none, and NEXT.md has carried independent human readers as named, not built for weeks.

Stage P — the published census (free, mechanical, no API)

After stages A and B are frozen, the four hands are extracted for this span exactly as S213 extracted them (materials/hands.md, to be written then). At each of the three deep runs (V07 depth 3, V09 depth 4, V11 depth 4) and, if the extraction is clean, at the six depth-2 runs, for each hand: the English word rendering each Persian rhyme-bearer is identified — bearers, not clause-ends, which is note (bra)'s requirement and the reason S213's S11 did not become a false four-for-four.

Amendment A5 (round-1 critic, MAJOR 4). A source rhyme position is counted only when the English word rendering its bearer stands at an English line-end, and the hand's depth at that run is the largest mutually-STRICT set among those positioned ends. Three codes, not one scale:

Every extracted cell's raw text is preserved in raw/.

5. Predictions, registered

id statement bar
F1 gate — hit rate on the target pair in the NEAR condition, pooled over 14 links × 3 seats (the matched instrument floor) ≥ 0.80
F2 gate — link rate on A− and D− items (non-rhyming pairs), pooled over S1 and S2 ≤ 0.20
P1 primary, matched, identical words — the paired NEAR − FAR target-pair hit rate over the 84 S1 cells, with the registered test an exact one-sided sign test over the 9 runs (A10) and reporting conditional on three retained seats (A11) rate difference ≥ 0.25
P2 secondary, ecological — hit(A+, COUPLET arm) − hit(D+, R44 arm) in S2, same nine passages, different words and twice the rhyme density same sign as P1
P3 secondary, natural — hit(A+, the 17 MATHNAWI adjacent rhymes) − hit(D+, R44) same sign as P1
P4 descriptive — hit(D+) by position in the run (link 1 vs links 2–3 of the deep runs) no bar
P5 the published census — no published hand holds the source's rhyme at English line-ends to the source's depth at any run of depth ≥ 3 0 of 12 cells

P1's direction is registered: the adjacent placement is predicted to be heard more often.

What a miss means, in the critic's own words (amendment A4, round-1 MAJOR 3). With 14 distance links, nine passages and links correlated inside a run, a difference below 0.25 licenses exactly one sentence — "the preregistered ≥ 0.25 contrast was not observed" — and the result page is forbidden the sentence distance is not the obstacle. A real adverse effect smaller than 0.25 is fully compatible with a miss.

6. Failure criteria, registered

  1. F1 below 0.80 → the task does not detect an adjacent rhyme on the same words that the distance arm carries; everything is withheld and the session reports an instrument that did not work.
  2. F2 above 0.20 → the seats are filling in a scheme rather than reading line-ends; the primary is withheld and only the false-alarm figure is published.
  3. Fewer than 8 of the 14 links carrying a tool-STRICT English rhyme — i.e. R44 failed to place the rhyme often enough to measure — → P1 is withheld; the reachability grades from stage A are reported alone, as a translator's finding about what could not be done.
  4. Stage F catches fewer than 4 of 6 plants → the R44 grades are reported self-certified and P1 is reported with a stated fidelity limit, not withheld (the seats' rhyme reading does not depend on the renderings being faithful).
  5. Any seat returning fewer than 49 of its 58 stage-S bodies parsed → that seat is dropped whole and the result reports the remaining seats.
  6. Saturation — if the target-pair hit rate exceeds 0.95 in both NEAR and FAR, the task is too easy to discriminate and P1 is reported as uninformative rather than as a null.

7. What the result may say, and what it may not

8. Budget

Declared experiment ceiling $1.60. Runner ceiling $1.35, stop-loss $1.20. UTC day 2026-08-23 opens with the full $5.00 unspent.

The ceiling rose from $1.20 to $1.50 and then to $1.60 during design, before any data call, both times on a critic finding: A1 added an arm, and A6 replaced that arm with a 28-body matched-pair stage while S2 kept the whole-passage arms. The reasons are in critic-response.md and the arithmetic is here.

stage calls worst case basis
pre-run critic, up to 3 rounds 3 $0.30 rounds 1 and 2 billed $0.054955 and $0.079273
stage S1 — matched pairs 84 $0.40 max_tokens 400; stimuli are 4 lines; P3 priced from its own first 12 calls
stage S2 — whole passages 90 $0.45 max_tokens 400; passages ≤ 8 lines
stage F 48 $0.35 max_tokens 500; prompts carry Persian and English
headroom — $0.10 re-dispatches
total 225 $1.60

The worst case is built from max_tokens, not from an assumed output length — note (abc). Costs are the API-returned billed figures with "usage": {"include": true}; the key-usage delta is not a cross-check and is not reported as one, note (bof).