Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260825b-flippancy/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260825b-flippancy
statusfrozen
created2026-08-25
updated2026-08-25
sensesstyle-correspondence, voice, literary-quality
provisionaltrue
internal-judgment-onlyfalse
linkswiki/arms/ARM-declared-function.md, wiki/base/anchors/A-hariri-hands/README.md, workshop/translations/maqamat-sanaa/R48-v1/translation.md, workshop/translations/maqamat-sanaa/R48D-v1/translation.md, workshop/translations/maqamat-sanaa/R43-v1/translation.md, workshop/translations/maqamat-sanaa/cola.tsv, workshop/regimes/R48-rhyme-first.md, workshop/regimes/R49-dechime.md, config/models.md, config/budget.md, framework/v0.2/README.md

E-20260825b — two printed reasons for not rhyming English prose, put to a blind panel

Design v2, 2026-08-25. v1 was frozen and sent to two independent adversarial critics before any data call; both returned NEEDS-REDESIGN and both, independently, found the same fatal defect in the primary item. Every change v2 makes is listed in critic-response.md and marked [v2] below. The frozen v1 text is recoverable at commit ecb9b08’s successor in this file’s history.

Frozen 2026-08-25 before any scored call was dispatched. The pre-run critic's findings and the response to them are in critic-response.md.

1. The question, and why it is not a question about this project

Two of the three published English hands on al-Ḥarīrī's Maqāmāt print a refusal to carry the Arabic's rhymed prose, and both give a reason about the reader (wiki/base/anchors/A-hariri-hands/README.md §3):

Chappelow, 1767 — "nor shall I imitate the author in my translation. To attempt it might be looked upon as a piece of pedantry: and indeed our English tongue will not admit of it."

Preston, 1850 — "Rhyming prose is extremely ungraceful in English, and introduces an air of flippancy, unless the subject be of the most light and frivolous description."

These are declared accounts of what a device does to a reader, made in print by working translators, 176 and 259 years before anything in this repository. They are the reason the English Maqāmāt have no rhyme, and — through framework/v0.2 §7.12, which "declines to say compensate and declines to say the opposite" — they are part of why this project's handbook has nothing to say about the oldest positive move in the craft. Nobody has tested them, because nobody made the object they are a reason about. T-maqamat-sanaa-R48-v1 is that object.

The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5), written before the design: this unit teaches whether the reason two published translators give in print for not carrying a source's central sound figure describes what a reader of the carried version actually receives, and whether the exemption one of them grants light subject matter is real. That is a sentence about translating literature. It is not a sentence about this project's raters.

2. What this replaces, and why the arm's step 1 was re-planned

ARM-declared-function step 1 as constituted was "build the rubric and the held-out split, and run the labelling blind", with the instruction: write the subject-rule sentence first, and if it comes out about this project's raters, close the arm retired. The honest sentence for that step is do independently blinded labellers agree on function categories? — which is about the raters. The arm's page anticipated this exactly.

Retiring the arm on that ground would have been an error, because the arm's question survives the failure of its designed route. The declaration does not have to be manufactured by labellers. Four printed declarations of policy are already on the shelf, and Preston's carries a falsifiable prediction about readers together with a stated moderator. So step 1 is re-planned, not abandoned: the reason is written on the arm page, and the arm keeps its budget and its completion criterion.

3. Materials

One work, one hand, three arms, differing only where the design says.

arm text what it is
ORD T-maqamat-sanaa-R43-v1 the ordinary restrained rendering, frozen 2026-08-24, before this design existed. 2 STRICT adjacent colon-end rhymes of 138. Context arm; in no primary.
RHY T-maqamat-sanaa-R48-v1 the rhyme-forward rendering. 33 STRICT, 17 NEAR, 4 IDENTICAL.
DRH T-maqamat-sanaa-R48D-v1 RHY with the colon-end word of one colon changed at each graded chime and nothing else. 0 STRICT, 3 NEAR, 0 IDENTICAL [v2: four further substitutions on C__P2 finding 5], and none of RHY's chimes surviving.

Segments: 15, cut from the Arabic's own structure before any English was scored, verse excluded (all three arms print the same verse, and it rhymes in all three).

The opening frame p001–p033 is excluded: it is neither grave nor comic, and a third stratum would have cost 45 calls to measure a moderator with no prediction attached.

The source's figure density is constant across the strata: 137 of the 139 cola of this maqāma stand inside a rhyme figure (loci-frozen.md §2). Preston's exemption therefore cannot be a fact about the Arabic. It is a claim about English.

4. The estimand, stated plainly because it is not the obvious one

The effect of a translator adopting a rhyme-forward policy on this work — not the effect of the sound in isolation. Reaching a chime forces diction: perambulation ⁄ ministration, at bay, replete. R48D removes the chime but keeps that diction, which is the tightest available control and still not a pure one, because 41 of the 65 de-chiming substitutions make the control the more faithful text (R48D-v1/translation.md §The limit this page exists to declare).

[v2, on C__P1 finding 2 — v1 said "a null under this design is clean" and that was wrong.] Neither direction is clean. A positive is confounded with the exactness the rhyme cost. A null may be cancellation: the chime pushing an item one way and the lost exactness pushing it back. A translator choosing Preston's question does not get to choose the sound without the diction it forces, so the confounded quantity is the one he would actually face — but §10 forbids the result page from reporting it as an effect of rhyme alone, in either direction.

5. Seats

P1 openai/gpt-5.6-terra · P2 google/gemini-3.6-flash · QR qwen/qwen3.7-max (config/models.md; P3 is a cost problem, P4/P5 are out on notes (bps)/(bne), GL is out on long prompts). Every call independent, no context shared, temperature 1.

No seat is told that the passage is a translation, that it is from Arabic, that there are other versions, that rhyme is at issue, or that the passage has anything to do with sound. The prompt says "a passage of English prose".

6. The instrument — five items per call, in this order [v2]

Both critics found the same defect in v1's primary item and it was the finding that paid for the critic stage. v1 asked whether the writing is arch or knowing — enjoying itself. Preston's flippancy is disrespectful levity: a lightness of manner improper to a grave subject. Ornate, self-conscious, wholly serious prose scores high on the first and low on the second. v2 puts Preston's construct first and keeps v1's wording as a separate item, so that neither reading can be chosen after the numbers are in.

The prompt now opens with an instruction the seats did not have in v1, on C__P2 finding 2: questions 1 to 3 are about the manner of the writing and not the subject it treats; a passage may be about the gravest matter in the world and still be written in a light manner.

  1. LEVITY, 0–10 — the primary. Does the manner of this writing treat its subject with due seriousness, or is there a levity in the manner — a lightness that sits oddly with what is being said? 0 = the manner is entirely serious; 10 = markedly light. This is Preston's "air of flippancy", and the word flippancy is not used.*
  2. ARCH, 0–10. v1's item, kept: arch or knowing … enjoying itself. 0 = straight-faced.
  3. GRACE, 0–10. How graceful is the English? Preston: "extremely ungraceful in English".
  4. MANNER, 0–10. How much does the writing draw attention to its own manner rather than to what it says? Chappelow: "a piece of pedantry". His second clause — "our English tongue will not admit of it" — is a claim about the language's capacity and is declared out of scope (C__P1 finding 14).
  5. PLACE, forced choice, exploratory. If you met this passage with no context, where would you expect it to have come from? A a book of devotion or a sermon · B a serious literary romance, history or scripture · C a comic or picaresque tale · D a humorous magazine piece · E a parody or burlesque.

Strict JSON out, no working in the body (note (bnk)). max_tokens 1200 on P1/P2, 2000 on QR, reasoning.max_tokens 600. A ```json fence is stripped before parsing.

6a. The pilot, run before the 135 were dispatched [v2, on C__P1 finding 7]

Six calls: segment G4, arms RHY and DRH, all three seats, $0.014797. All six parsed.

item DRH (P1/P2/QR) RHY (P1/P2/QR)
LEVITY 0 · 0 · 0 1 · 1 · 0
ARCH 0 · 0 · 1 1 · 2 · 1
GRACE 8 · 8 · 8 7 · 5 · 8
MANNER 4 · 7 · 7 7 · 9 · 7
PLACE A · A · A A · A · A

What the pilot establishes, and it is registered here rather than discovered later. MANNER and GRACE discriminate. LEVITY sits at or near the floor on a grave segment in both arms — exactly what C__P2 predicted. The primary is not changed on that evidence, for two reasons: the primary was chosen to be Preston's claim and not the most sensitive item, and a floor at 0 in both arms is an answer to Preston on a grave subject rather than an instrument failure. What is registered is the consequence: if LEVITY floors, the informative content is in the registered secondaries, and the result page will say the primary floored rather than quietly promoting one of them. PLACE is at ceiling A on this segment and may carry nothing on GRAVE.

The six pilot cells are re-dispatched under their stage-P tags and their pilot draws are folded into the stage-R repeatability count.

7. Stages and call count [v2]

stage what calls
C pre-run adversarial critic, P1 and P2 on the frozen v1 2 (spent)
pilot §6a — G4 × RHY/DRH × 3 seats 6 (spent)
P scored: 15 segments × 3 arms × 3 seats 135
A audibility: 15 × RHY/DRH × all three seats (v1 had one seat; C__P2 finding 6) 90
G gravity: 15 × RHY/DRH × P2 (v1 tested ORD, the wrong arm; C__P2 finding 3) 30
R repeatability: 5 segments × 3 arms × P2, a second draw of stage P 15
total 278

Stages A, G and R run after P, and their prompts are printed below and frozen here, so no choice in them can be made in the light of P.

Stage A, verbatim — "does this prose rhyme or chime at the ends of its clauses — that is, do the words that end successive clauses or sentences repeatedly echo one another in sound?" Answer {"chimes": "yes|no"}. Coding: the string is lowercased and must begin yes or no; anything else is a dropped cell. A segment counts as heard for an arm if 2 of its 3 seats say yes.

Stage G, verbatim — "judging by its subject matter alone — what it is about, not how it is written — is this passage grave, or is it light?" Answer {"subject": "grave|light"}, coded the same way. A segment counts as grave if both its arms are called grave.

C__P1 finding 9 is right that a manipulation check which names the feature directs attention to it. That is inherent to every manipulation check; what is claimed from stage A is only the device is findable when looked for, which is the weakest premise the primary needs.

Dispatch order is shuffled with a fixed seed (20260825) so that no arm sits in a block of the run; provider, finish reason and cost are logged per call (C__P1 finding 18).

8. Predictions and tests, registered [v2]

Scores are z-standardised within seat before pooling (C__P2 finding 4: three model families do not share a scale). Raw means are reported beside the standardised ones. Each seat also gets its own sign test over its own 15 paired differences — this project's standing practice of reporting jurors separately.

The test statistic, named honestly (C__P1 finding 3, accepted verbatim). The arms are fixed texts, not randomly assigned treatments. Enumerating all 2^15 sign assignments of the 15 paired differences yields a reference distribution under exchangeability of the arm label, and the tail probability from it is reported as such. It is not a Type-I error guarantee and the result page will not call it one.

Power, stated because v1 did not (C__P1 finding 20). With 15 paired differences a one-sided sign test needs 12 of 15 to reach 0.0176 and 11 of 15 gives 0.0592. This design can detect a large and consistent effect and nothing smaller. No minimally interesting effect size is claimed.

9. Failure criteria, registered [v2]

10. What the result page may not say, whatever the numbers are

  1. Not "rhyme makes English prose flippant". The arms differ in the diction the rhyme forced and in fidelity (§4). The largest claim available from a positive is a rhyme-forward policy on this work, at this dose, moved these seats on this item.
  2. Not that the panel's reading is a human reader's. No sense here is Tier-D calibrated (config/models.md); every figure carries provisional: true, and Preston's readers were Victorian.
  3. Not that Preston is refuted by a null. Three model seats failing to register an effect is not a demonstration that readers do not, and F2 is the only thing standing between a null and "the instrument cannot see this".
  4. One work, one hand, one language pair, one dose. Every figure is of this maqāma in this hand's English.

11. Budget

Declared ceiling raised from $2.20 to $2.60 at v2, the critic having added 78 calls (stage A tripled, stage G doubled, the pilot) — the same mechanism as S218 and S220. UTC-day headroom $4.404928625 after S220. Worst case built from max_tokens as note (abc) requires: 278 calls at per-call worst $0.0072 (P1), $0.0045 (P2), $0.0088 (QR at 2000) plus ~500 input tokens — ≈$1.6 — plus a margin for provider routing (the S022 caution: routing can price a call 4× its list rate) and for F5 re-dispatches. Spent before stage P: $0.113350 (critic $0.098553, pilot $0.014797).

The translation limb cost $0 and is not ledgered (charter §3, A4).