Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260826-balanced-period/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260826-balanced-period
statusfrozen
created2026-08-26
updated2026-08-26
sensesstyle-correspondence
linkswiki/arms/ARM-balanced-period.md, wiki/base/anchors/A-hariri-hands/README.md, workshop/regimes/R50-balanced-periods.md, workshop/translations/maqamat-hulwan/R50-v1/translation.md, workshop/translations/maqamat-sanaa/R50-v1/translation.md, workshop/experiments/E-20260826-balanced-period/metrics.py, workshop/experiments/E-20260826-balanced-period/operationalize.py, framework/v0.2/README.md

E-20260826-balanced-period — does the man who printed the rule obey it?

Frozen before any comparator file was opened. The two published English texts this design measures were located in Internet Archive OCR by running-header line number, sliced to files by a script that printed only line and word counts, and not read. The lead's own R50 rendering of Assembly II was made and frozen from the Arabic alone before this page was written.

1. The question

Theodore Preston, in the introduction to his 1850 Makamat, or Rhetorical Anecdotes of Al Hariri of Basra, refused to carry al-Ḥarīrī's rhymed prose into English and in the same sentence printed what he would put in its place: clauses "though not rhyming together, are arranged as far as possible in evenly balanced periods, and never exceed a certain length."

A-hariri-hands records that he said this. Nobody has checked whether his own printed English does it.

Does a translator's declared compositional policy leave a measurable trace on his own page — and how large is that trace beside the trace left by a translator who is deliberately executing the same policy as a rule?

The second half is what makes the first answerable. A difference between Preston and another Victorian translator is uninterpretable on its own: two hands differ. The lead's arms supply the scale. R50 is Preston's sentence read as a rule and executed on purpose; R43 and R48 are the same hand on the same Arabic with no balance rule at all. R50 − R43 is what the policy is worth when someone is actually trying. Preston − Chenery is what it is worth in the historical record. The ratio of the two is the finding.

2. Why this is not method work

The subject rule (wiki/tracks.md) asks what the unit teaches about translating literature or evaluating translations. It teaches whether a translator's stated method is a description of his page or a description of his intentions — on the one nineteenth-century English hand that stated a positive method precisely enough to be checked, and against a modern rendering made under that method. Nothing here is about the project's instruments, raters, or published figures.

3. The measure, and why it is not the lead's choice

The obvious threat is that the lead picks the metric that produces the finding. So the metric was fixed first, by seats shown Preston's sentence with the author, the work, the language of the original and the existence of this study all withheld, and asked only what they would count (operationalize.py, raw/operationalize.json, three seats, $0.0641).

Registered before dispatch: the seats will name (i) a ceiling on clause length and (ii) some equality of length between neighbouring clauses.

Outcome: 3 of 3, and both, and in that order. All three put the clause-length ceiling first; all three named the absolute syllable difference between neighbouring clauses; all three independently added the third thing Preston's sentence says, zero end-rhyme. Two of the three (P1, P2) also named the dispersion of clause length — standard deviation or coefficient of variation — which R50's own rule does not mention. It is therefore measured here as a seat-named quantity, not a lead-named one.

The four measures, and nothing else:

# measure unit "the sentence is true" looks like
M1 syllables between one clause boundary and the next each clause a low maximum and 95th percentile
M2 |Δ syllables| between consecutive clauses each adjacent pair small, concentrated near 0
M3 SD and coefficient of variation of clause length the whole text low
M4 rhyme relation of the two clause-end bearers each adjacent pair NONE throughout

Syllables from CMUdict; rhyme from tools/rhyme_pairs.py under its 2026-08-22 rule, unaltered. Both are in metrics.py, written before any comparator was measured.

4. Materials

Panel A — «المقامة الصنعانية», Assembly I. Seven English texts of one Arabic text: Preston 1850, Chenery 1867, Chappelow 1767, and the lead's R43, R48, R48D, R50. The three published extracts are the ones already stored and verified at E-20260824b-hariri-hands/raw/comparators/.

Panel B — «المقامة الحلوانية», Assembly II. Three English texts: Preston 1850, Chenery 1867, and the lead's R50-v1, made this session. Panel B is out-of-sample for the published hands and for the regime alike: neither R50 nor this design existed when Assembly II was chosen, and the comparators were sliced blind.

Segmentation is each hand's own printed marking, never the lead's judgement.

5. The three controls

C1 — the shuffle null, and it is the one that matters. A hand whose clauses are all short gets a small M2 for free: |Δ| cannot be large if nothing is long. So for each text, M2 is also computed over 10,000 random permutations of that same text's own clause lengths, and the reported statistic is the ratio M2_observed / M2_shuffled. The ratio is pairing with the length distribution divided out. A hand that balances neighbouring clauses on purpose scores below 1; a hand whose low M2 is an artefact of short clauses scores at 1. Registered: the raw M2 and the ratio are both reported whatever they say, and the ratio is the primary.

C2 — the source-tracking control. "Evenly balanced periods" may not be a property the translator imposes at all: al-Ḥarīrī's own cola may already be balanced, and a hand that follows the Arabic closely would inherit the balance without doing anything. So the same M2 is computed on the Arabic (|Δ tokens| between consecutive cola of the copy-text) and the same shuffle ratio is computed for it. If the Arabic's ratio is itself well below 1, balance is in the source; if the English hands' ratios differ from each other and from the Arabic's, it is in the hands.

C3 — the count check on the extraction. Preston's clause count, recovered from the OCR, is compared with the copy-text's colon count. If the recovered count differs from the Arabic colon count by more than 20% on a panel, that panel's Preston row is withheld and the extraction failure is the reported result for it. Chenery's em-dash segmentation gets the same check. This is stated before the files are opened.

6. Registered predictions

Written before any comparator was measured. Each names what would falsify it.

What would make the whole thing uninterpretable, stated in advance: if C3 fails on both panels, there is no measurement here and the result page says so.

7. Analysis and verification

Exact one-sided permutation tests on the paired-by-position differences where the count check licenses pairing, and on the unpaired clause distributions where it does not; the shuffle null of C1 is exhaustive-by-simulation at 10,000 draws with a fixed seed recorded in the output. Every reported number is recomputed by verify.py, which imports metrics.py but re-derives the summaries and the nulls independently.

8. What this cannot show

It cannot show that Preston's English reads as balanced, only that it is. It cannot show that balance is worth having — RS-20260825c-worth-paying is the arm that asks that, and its own answer is that the preference instrument it used could not be trusted. And a difference between two Victorian translators, even a large one, is two hands: Preston is the only nineteenth-century English translator of al-Ḥarīrī who stated this policy, so the design has one declarer and no replication of the declaring. That limit is structural and no amount of measurement removes it.

9. Cost

Gate stage operationalize.py: $0.0640929 actual (three seats, two re-dispatched at a higher cap after P1 returned an empty body and P2 a truncated one at 1500). Pre-run critic: two seats, estimated ceiling $0.12. Everything else is arithmetic and costs nothing. Declared session ceiling $0.40.


10. Amendments after the pre-run critic — made before any comparator was measured

critic-response.md records the round in full: two non-Anthropic seats, NEEDS-REDESIGN, six BLOCKING findings, six SERIOUS, two MINOR, nine accepted in full and two in part, none overruled. What binds the run:

  1. P4's "working rule / mostly a preface" classification is deleted (B1, B5). Both contrasts are reported with direction and both units; neither is treated as a calibrated effect size.
  2. Three segmentations, not one (B2). OWN (each hand's printed marking), U-COMMA and U-STRONG (one mechanical rule applied identically to every text). A result counts only if it survives all three; a result that appears under OWN alone is reported as a property of the printed page.
  3. Two units of length, not one (M2). Syllables, the registered one, and words, added here because the two OCR sources differ in quality and a corrupted letter changes a syllable count and not a word count. Out-of-dictionary rates are reported per text.
  4. C2's inference is deleted (B3, B6). What is kept is the one comparison that does not cross scales: the dimensionless coefficient of variation of unit length, English against Arabic.
  5. Every paired-by-position test is removed (B4). Nothing pairs a Preston unit with a Chenery unit or with an Arabic colon. All statistics are within-text or distributional.
  6. C1 claims only unusually smooth adjacency relative to a random reordering of that text's own unit lengths (S3). No intention is inferred from it.
  7. P1–P3 now have three outcomes (S4): supported only if Preston is lower on both panels under all three segmentations; contradicted if equal or higher on either; the margin is reported either way.
  8. M3 is descriptive, not evidence for the ceiling (S2), and §3's claim is demoted: the seat probe shows that the obvious operationalisation is the obvious one, not that metric choice was foreclosed.
  9. P5 is an extraction check with an exact criterion (S5): the pipeline must reproduce RS-20260825c's four published rhyme figures for the lead's Assembly I arms exactly.
  10. The lineation claim is measured, not asserted (M1): the share of recovered Preston lines that end in punctuation and begin with a capital is reported.