Repository path: workshop/experiments/E-20260825b-flippancy/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260825b-flippancy |
| status | frozen |
| created | 2026-08-25 |
| updated | 2026-08-25 |
| senses | style-correspondence, voice, literary-quality |
| provisional | true |
| internal-judgment-only | false |
| links | wiki/arms/ARM-declared-function.md, wiki/base/anchors/A-hariri-hands/README.md, workshop/translations/maqamat-sanaa/R48-v1/translation.md, workshop/translations/maqamat-sanaa/R48D-v1/translation.md, workshop/translations/maqamat-sanaa/R43-v1/translation.md, workshop/translations/maqamat-sanaa/cola.tsv, workshop/regimes/R48-rhyme-first.md, workshop/regimes/R49-dechime.md, config/models.md, config/budget.md, framework/v0.2/README.md |
E-20260825b — two printed reasons for not rhyming English prose, put to a blind panel
Design v2, 2026-08-25. v1 was frozen and sent to two independent adversarial critics before any
data call; both returned NEEDS-REDESIGN and both, independently, found the same fatal defect in
the primary item. Every change v2 makes is listed in critic-response.md and marked [v2] below.
The frozen v1 text is recoverable at commit ecb9b08’s successor in this file’s history.
Frozen 2026-08-25 before any scored call was dispatched. The pre-run critic's findings and the
response to them are in critic-response.md.
1. The question, and why it is not a question about this project
Two of the three published English hands on al-Ḥarīrī's Maqāmāt print a refusal to carry the
Arabic's rhymed prose, and both give a reason about the reader
(wiki/base/anchors/A-hariri-hands/README.md §3):
Chappelow, 1767 — "nor shall I imitate the author in my translation. To attempt it might be looked upon as a piece of pedantry: and indeed our English tongue will not admit of it."
Preston, 1850 — "Rhyming prose is extremely ungraceful in English, and introduces an air of flippancy, unless the subject be of the most light and frivolous description."
These are declared accounts of what a device does to a reader, made in print by working
translators, 176 and 259 years before anything in this repository. They are the reason the English
Maqāmāt have no rhyme, and — through framework/v0.2 §7.12, which "declines to say compensate
and declines to say the opposite" — they are part of why this project's handbook has nothing to say
about the oldest positive move in the craft. Nobody has tested them, because nobody made the
object they are a reason about. T-maqamat-sanaa-R48-v1 is that object.
The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5), written before the
design: this unit teaches whether the reason two published translators give in print for not
carrying a source's central sound figure describes what a reader of the carried version actually
receives, and whether the exemption one of them grants light subject matter is real. That is a
sentence about translating literature. It is not a sentence about this project's raters.
2. What this replaces, and why the arm's step 1 was re-planned
ARM-declared-function step 1 as constituted was "build the rubric and the held-out split, and run
the labelling blind", with the instruction: write the subject-rule sentence first, and if it
comes out about this project's raters, close the arm retired. The honest sentence for that step
is do independently blinded labellers agree on function categories? — which is about the raters.
The arm's page anticipated this exactly.
Retiring the arm on that ground would have been an error, because the arm's question survives the failure of its designed route. The declaration does not have to be manufactured by labellers. Four printed declarations of policy are already on the shelf, and Preston's carries a falsifiable prediction about readers together with a stated moderator. So step 1 is re-planned, not abandoned: the reason is written on the arm page, and the arm keeps its budget and its completion criterion.
3. Materials
One work, one hand, three arms, differing only where the design says.
| arm | text | what it is |
|---|---|---|
ORD |
T-maqamat-sanaa-R43-v1 |
the ordinary restrained rendering, frozen 2026-08-24, before this design existed. 2 STRICT adjacent colon-end rhymes of 138. Context arm; in no primary. |
RHY |
T-maqamat-sanaa-R48-v1 |
the rhyme-forward rendering. 33 STRICT, 17 NEAR, 4 IDENTICAL. |
DRH |
T-maqamat-sanaa-R48D-v1 |
RHY with the colon-end word of one colon changed at each graded chime and nothing else. 0 STRICT, 3 NEAR, 0 IDENTICAL [v2: four further substitutions on C__P2 finding 5], and none of RHY's chimes surviving. |
Segments: 15, cut from the Arabic's own structure before any English was scored, verse excluded (all three arms print the same verse, and it rhymes in all three).
GRAVE, 8 segments — the sermon,p034–p094. Abū Zayd on death, the grave, the Muster and the reckoning. 44–85 words per segment.LIGHT, 7 segments — the swindle,p096–p139. The collection, the slipping away, the cave, the roast kid and the wine-jar, the reveal. 23–63 words per segment.
The opening frame p001–p033 is excluded: it is neither grave nor comic, and a third stratum
would have cost 45 calls to measure a moderator with no prediction attached.
The source's figure density is constant across the strata: 137 of the 139 cola of this maqāma
stand inside a rhyme figure (loci-frozen.md §2). Preston's exemption therefore cannot be a fact
about the Arabic. It is a claim about English.
4. The estimand, stated plainly because it is not the obvious one
The effect of a translator adopting a rhyme-forward policy on this work — not the effect of the
sound in isolation. Reaching a chime forces diction: perambulation ⁄ ministration, at bay,
replete. R48D removes the chime but keeps that diction, which is the tightest available control
and still not a pure one, because 41 of the 65 de-chiming substitutions make the control the
more faithful text (R48D-v1/translation.md §The limit this page exists to declare).
[v2, on C__P1 finding 2 — v1 said "a null under this design is clean" and that was wrong.]
Neither direction is clean. A positive is confounded with the exactness the rhyme cost. A
null may be cancellation: the chime pushing an item one way and the lost exactness pushing it
back. A translator choosing Preston's question does not get to choose the sound without the diction
it forces, so the confounded quantity is the one he would actually face — but §10 forbids the result
page from reporting it as an effect of rhyme alone, in either direction.
5. Seats
P1 openai/gpt-5.6-terra · P2 google/gemini-3.6-flash · QR qwen/qwen3.7-max
(config/models.md; P3 is a cost problem, P4/P5 are out on notes (bps)/(bne), GL is out on
long prompts). Every call independent, no context shared, temperature 1.
No seat is told that the passage is a translation, that it is from Arabic, that there are other versions, that rhyme is at issue, or that the passage has anything to do with sound. The prompt says "a passage of English prose".
6. The instrument — five items per call, in this order [v2]
Both critics found the same defect in v1's primary item and it was the finding that paid for the critic stage. v1 asked whether the writing is arch or knowing — enjoying itself. Preston's flippancy is disrespectful levity: a lightness of manner improper to a grave subject. Ornate, self-conscious, wholly serious prose scores high on the first and low on the second. v2 puts Preston's construct first and keeps v1's wording as a separate item, so that neither reading can be chosen after the numbers are in.
The prompt now opens with an instruction the seats did not have in v1, on C__P2 finding 2:
questions 1 to 3 are about the manner of the writing and not the subject it treats; a passage may
be about the gravest matter in the world and still be written in a light manner.
LEVITY, 0–10 — the primary. Does the manner of this writing treat its subject with due seriousness, or is there a levity in the manner — a lightness that sits oddly with what is being said? 0 = the manner is entirely serious; 10 = markedly light. This is Preston's "air of flippancy", and the word flippancy is not used.*ARCH, 0–10. v1's item, kept: arch or knowing … enjoying itself. 0 = straight-faced.GRACE, 0–10. How graceful is the English? Preston: "extremely ungraceful in English".MANNER, 0–10. How much does the writing draw attention to its own manner rather than to what it says? Chappelow: "a piece of pedantry". His second clause — "our English tongue will not admit of it" — is a claim about the language's capacity and is declared out of scope (C__P1finding 14).PLACE, forced choice, exploratory. If you met this passage with no context, where would you expect it to have come from? A a book of devotion or a sermon · B a serious literary romance, history or scripture · C a comic or picaresque tale · D a humorous magazine piece · E a parody or burlesque.
Strict JSON out, no working in the body (note (bnk)). max_tokens 1200 on P1/P2, 2000 on
QR, reasoning.max_tokens 600. A ```json fence is stripped before parsing.
6a. The pilot, run before the 135 were dispatched [v2, on C__P1 finding 7]
Six calls: segment G4, arms RHY and DRH, all three seats, $0.014797. All six parsed.
| item | DRH (P1/P2/QR) |
RHY (P1/P2/QR) |
|---|---|---|
LEVITY |
0 · 0 · 0 | 1 · 1 · 0 |
ARCH |
0 · 0 · 1 | 1 · 2 · 1 |
GRACE |
8 · 8 · 8 | 7 · 5 · 8 |
MANNER |
4 · 7 · 7 | 7 · 9 · 7 |
PLACE |
A · A · A | A · A · A |
What the pilot establishes, and it is registered here rather than discovered later.
MANNER and GRACE discriminate. LEVITY sits at or near the floor on a grave segment in both
arms — exactly what C__P2 predicted. The primary is not changed on that evidence, for two
reasons: the primary was chosen to be Preston's claim and not the most sensitive item, and a floor
at 0 in both arms is an answer to Preston on a grave subject rather than an instrument failure.
What is registered is the consequence: if LEVITY floors, the informative content is in the
registered secondaries, and the result page will say the primary floored rather than quietly
promoting one of them. PLACE is at ceiling A on this segment and may carry nothing on GRAVE.
The six pilot cells are re-dispatched under their stage-P tags and their pilot draws are folded
into the stage-R repeatability count.
7. Stages and call count [v2]
| stage | what | calls |
|---|---|---|
C |
pre-run adversarial critic, P1 and P2 on the frozen v1 |
2 (spent) |
pilot |
§6a — G4 × RHY/DRH × 3 seats |
6 (spent) |
P |
scored: 15 segments × 3 arms × 3 seats | 135 |
A |
audibility: 15 × RHY/DRH × all three seats (v1 had one seat; C__P2 finding 6) |
90 |
G |
gravity: 15 × RHY/DRH × P2 (v1 tested ORD, the wrong arm; C__P2 finding 3) |
30 |
R |
repeatability: 5 segments × 3 arms × P2, a second draw of stage P |
15 |
| total | 278 |
Stages A, G and R run after P, and their prompts are printed below and frozen here, so
no choice in them can be made in the light of P.
Stage A, verbatim — "does this prose rhyme or chime at the ends of its clauses — that is, do
the words that end successive clauses or sentences repeatedly echo one another in sound?" Answer
{"chimes": "yes|no"}. Coding: the string is lowercased and must begin yes or no; anything
else is a dropped cell. A segment counts as heard for an arm if 2 of its 3 seats say yes.
Stage G, verbatim — "judging by its subject matter alone — what it is about, not how it is
written — is this passage grave, or is it light?" Answer {"subject": "grave|light"}, coded the
same way. A segment counts as grave if both its arms are called grave.
C__P1 finding 9 is right that a manipulation check which names the feature directs attention to
it. That is inherent to every manipulation check; what is claimed from stage A is only the device
is findable when looked for, which is the weakest premise the primary needs.
Dispatch order is shuffled with a fixed seed (20260825) so that no arm sits in a block of the
run; provider, finish reason and cost are logged per call (C__P1 finding 18).
8. Predictions and tests, registered [v2]
Scores are z-standardised within seat before pooling (C__P2 finding 4: three model families do
not share a scale). Raw means are reported beside the standardised ones. Each seat also gets its
own sign test over its own 15 paired differences — this project's standing practice of reporting
jurors separately.
Q1— the sole primary.LEVITY(RHY) > LEVITY(DRH), on the 15 per-segment means of within-seat z-scores. One-sided, α = 0.05.Q3,Q4— secondary, Holm-adjusted across the two.GRACE(RHY) < GRACE(DRH);MANNER(RHY) > MANNER(DRH).Q2— exploratory, demoted from confirmatory (C__P1findings 11 and 12). TheRHY − DRHdifference is larger onGRAVEthanLIGHT, forLEVITYand forGRACE(C__P1finding 13: Preston's exemption grammatically qualifies both). It is confounded and will not be reported as a test of Preston's exemption: the two strata differ in rhetorical mode as well as subject, and a picaresque swindle is not obviously "the most light and frivolous description".Q5, exploratory.PLACEshifts from {A, B} toward {C, D, E}.ARCHis reported besideLEVITY, untested, as the record of what v1 would have measured.ORDenters no test at all (C__P1finding 19). Its columns are printed as scale context.
The test statistic, named honestly (C__P1 finding 3, accepted verbatim). The arms are fixed
texts, not randomly assigned treatments. Enumerating all 2^15 sign assignments of the 15 paired
differences yields a reference distribution under exchangeability of the arm label, and the
tail probability from it is reported as such. It is not a Type-I error guarantee and the result
page will not call it one.
Power, stated because v1 did not (C__P1 finding 20). With 15 paired differences a one-sided
sign test needs 12 of 15 to reach 0.0176 and 11 of 15 gives 0.0592. This design can detect a
large and consistent effect and nothing smaller. No minimally interesting effect size is claimed.
9. Failure criteria, registered [v2]
F1— objective manipulation. Discharged before freezing and recorded, not gated.RHY33STRICTadjacent colon-end rhymes;DRH0, with none ofRHY's chimes surviving.F2— reader-side audibility (withholdsQ1,Q3,Q4). StageA, majority of three seats. IfRHYis not heard on at least 8 more of the 15 segments thanDRH, every primary is withheld: a panel that cannot find the device when it is looking cannot be reacting to it.F3— repeatability. A diagnostic that withholds nothing (C__P1finding 8; v1's claim that it "bounds the P values" is withdrawn). StageRplus the six pilot cells: the count of repeated cells returning an identical five-item vector is reported, and if it is high the result page says the instrument is close to deterministic on this task — which is a fact about the seats, not a correction to any number.F4— gravity manipulation (withholdsQ2only). StageGonRHYandDRH.Q2is withheld unless ≥ 6 of 8GRAVEsegments are called grave and ≥ 5 of 7LIGHTsegments are called light.F5— truncation and missingness (C__P1finding 17). Any call returningfinish_reason: length, an error, or unparseable JSON is re-dispatched once; a second failure is a dropped cell. A segment enters a paired test only if all six of itsRHY/DRHcells parsed; segments dropped for this reason are named and the reference distribution recomputed on the reduced n. More than 10 dropped cells of 135 voids stageP.
10. What the result page may not say, whatever the numbers are
- Not "rhyme makes English prose flippant". The arms differ in the diction the rhyme forced and in fidelity (§4). The largest claim available from a positive is a rhyme-forward policy on this work, at this dose, moved these seats on this item.
- Not that the panel's reading is a human reader's. No sense here is Tier-D calibrated
(
config/models.md); every figure carriesprovisional: true, and Preston's readers were Victorian. - Not that Preston is refuted by a null. Three model seats failing to register an effect is not
a demonstration that readers do not, and
F2is the only thing standing between a null and "the instrument cannot see this". - One work, one hand, one language pair, one dose. Every figure is of this maqāma in this hand's English.
11. Budget
Declared ceiling raised from $2.20 to $2.60 at v2, the critic having added 78 calls (stage A
tripled, stage G doubled, the pilot) — the same mechanism as S218 and S220. UTC-day headroom
$4.404928625 after S220. Worst case built from max_tokens as note (abc) requires: 278
calls at per-call worst $0.0072 (P1), $0.0045 (P2), $0.0088 (QR at 2000) plus ~500 input
tokens — ≈$1.6 — plus a margin for provider routing (the S022 caution: routing can price a call
4× its list rate) and for F5 re-dispatches. Spent before stage P: $0.113350 (critic
$0.098553, pilot $0.014797).
The translation limb cost $0 and is not ledgered (charter §3, A4).