Repository path: workshop/experiments/E-20260816c-checked-ornament/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260816c-checked-ornament |
| status | frozen |
| created | 2026-08-16 |
| updated | 2026-08-16 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-invented-ornament.md, wiki/findings/results/RS-20260816b-invented-figure.md, wiki/findings/results/RS-20260816-answering-figure.md, workshop/translations/kalila-nasik/R34-v1/translation.md, workshop/translations/kalila-ibn-urs/R36-v1/translation.md, workshop/regimes/R34-ornamentalist.md, workshop/regimes/R36-answering-hand.md, framework/v0.2/README.md, config/models.md, config/budget.md |
E-20260816c — the ornament, checked: does source access repair a reader's judgement of a supplied device, and does the checking reader call the invention a defect?
ARM-invented-ornament step 1. v2, rebuilt on a pre-run critic pass that returned
NEEDS REDESIGN with 28 findings, 8 of them BLOCKING, before any body was dispatched (§12).
Nothing changes after §12 is written.
1. What is at issue, in three sentences
RS-20260816b put two doctored arms of one English translation to three blind seats and found that
a supplied sound device makes a reader believe the Arabic had a figure there — 9 of 13 at places
it has none — statistically indistinguishable from the 8 of 13 that a genuine compensation scores
at places it does have one. So for a reader who cannot see the source, the judgement "this device
answers something in the original" carries no information about the original.
framework/v0.2 §7.14 carries that as a warning, and the warning is addressed to a reader who cannot
check. This design asks what changes when the reader can check, and it asks two things about
that reader: whether the source repairs the discrimination, and whether the invention is then counted
a defect.
2. The subject-rule sentence, and the wire
What this unit teaches about translating literature (wiki/tracks.md, continue-prompt.md
§4.5): whether the ornament a translator supplies survives being checked — whether a reader with the
original in front of them separates a device that answers a figure from a device that invents one,
and whether they hold the invention against the translation. The subject is what a translation
transmits and what a reader can hold it to. Source access is the condition varied, not the object
measured.
The wire between the limbs, in one sentence. The translation limb rendered a fresh chapter of the
same work under R36, a regime that answers all nineteen of that chapter's sound figures and forbids
invention anywhere else, and measured that the line can be held for +35 words; the study limb asks
whether holding it buys anything — whether any reader, even one with the Arabic open and the locus
marked, can tell the two hands apart, and whether they mind.
3. Materials
The site set is inherited, not chosen here. T-kalila-nasik-R34-v1 §4 censuses 22 ornamented
sites in a whole chapter, frozen at cf2dbb8f before this design existed, with an at column
scored against an Arabic figure inventory frozen before any English was written. This design drops
O22 (a chapter-wide repetition spanning sentences 2 and 19, with no single Arabic locus) and uses
the remaining 21 sites.
One truth label is changed from the census, on the critic's finding 1, and it is a correction to the
census rather than a judgement of this design's. O12 («فاستحسن الضيف كلامه وأعجبه») was scored
at: —, i.e. nothing in the Arabic. But T-kalila-nasik-R34-v1 §3's exclusion (a) names
كلامه وأعجبه by name as a doublet whose only rhyme is a grammatical enclitic — which is the
definition of the MEM stratum. The census contradicted its own inventory; O12 moves NONE →
MEM.
| stratum | n | what the Arabic has there | what the English device is |
|---|---|---|---|
FIG |
10 | an admitted sound figure (G1–G10) |
a compensation |
MEM |
5 | matched members whose only rhyme is a grammatical ending | intermediate |
NONE |
6 | neither sound nor matched members | an invention |
MEM is never pooled into a registered comparison and is reported on its own throughout.
3.1 The five arms
build_materials.py builds them; each item is the sentence of the frozen English containing the
site (median 41 words — the critic's finding 4: the whole paragraph carried other unmarked devices),
with the site's span marked ⟪…⟫, and the corresponding Arabic stretch marked ⟪…⟫ too wherever
Arabic is shown (finding 6.3: an unmarked Arabic sentence makes Q2 a scanning task).
| arm | Arabic | English | prompts |
|---|---|---|---|
AR |
the marked sentence, alone | — | Q2AR |
ORN.B |
— | the frozen R34 sentence |
Q1, Q2 |
PLN.B |
— | the same, with the census's own plain wording refused at the span | Q1, Q2 |
ORN.S |
marked | as ORN.B |
Q2, Q4 |
PLN.S |
marked | as PLN.B |
Q2, Q4 |
The AR arm is new in v2 and it is the critic's finding 22: nothing in this project has ever
measured whether these seats can identify a classical Arabic sound figure at all, and MC1 in v1
tried to infer it from a condition that still showed them English. AR shows Arabic and nothing else.
The plain arm is derived from the ornate one by one substitution, never written independently —
note (bpn). build_materials.py's check F asserts the two arms are character-identical
outside the marked span on all 21 items, and it passes. The substituted wording is not invented
here: it is the census's plain wording refused column, written by the translator at the time of
rendering and frozen with it.
3.2 What PLN is and is not — the critic's finding 2, accepted in full
v1 called PLN "English carrying no device". That was false and the claim is withdrawn. The
critic read the spans and named seven where the plain wording is itself patterned: O3 (devout and
diligent), O6 (there … there), O11 (get … get … get), O12, O15 (stepping and walking),
O19 (tongue … tongue), O21 (fathers … grandfathers). Several are unavoidable, because the
source figure at O19 and O21 is lexical repetition and a plain English rendering repeats the
word too.
So PLN is not a no-device control and is never described as one. It is the translator's own
plain alternative — what this hand would have written without the licence — and how much sound it
still carries is measured, not assumed: Q1 on PLN.B is exactly that measurement, and every
registered number below is reported a second time on the subset of items where PLN.B's Q1 rate is
≤ 1/3. That split is declared here, before the data exist, so it cannot be chosen afterwards.
3.3 Sites are not independent — the critic's finding 5, accepted
Six Arabic sentences carry more than one site; the 21 sites fall into 14 clusters. Every registered comparison is therefore computed twice: site-level, and cluster-level with each cluster contributing the mean of its sites. A finding is called only if both agree in sign and both clear the bar; where they disagree, the result is reported as unresolved.
3.4 The limits the critic named that this design does not fix
Written here rather than in a limits section after the fact.
- The ornate and plain spans differ in more than sound. Finding 3 lists ten places where they
differ in emphasis, explicitness, register or idiom — never tasted asserts more than
ليطرفه به; grew greedy moralises where set his heart on does not; forefathers is broader
than grandfathers. This is irreducible, it is the same limit
RS-20260816b§10.2 recorded, and no number here is presented as isolating sound. O15andO20remainFIGagainst the critic's advice, which held that a pair of matched imperfect verbs (تدرج وتمشي, ألا تدركه وتنسى وترجع) is conjugational morphology, not a figure. Overruled, with the reason: the frozen inventory rule admits muwāzana — the same morphological pattern in matched position — and its exclusion (a) is about rhyme produced by a grammatical ending, not about shared pattern. Changing the rule after seeing which sites it admits is the move the freeze exists to stop. But the objection is registered: every number is reported a second time under the critic's labelling (O15,O20→MEM), as a pre-declared sensitivity.- Three fixed models on 21 authored sites of one chapter of one work. Nothing here licenses a claim about human bilingual critics. Every claim on the result page is about these seats on these materials; the phrase "the only reader a translator is answerable to" is struck from v1 (finding 26.7).
4. Seats
P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the seats of
RS-20260816 and RS-20260816b. P5 is out on any task shape, note (bne). Temperature 0,
stateless calls, one body per (cell, prompt, seat), no seat told that arms exist.
The jury is not calibrated (config/models.md, Tier D NOT PASSED), and it has never been
screened on Arabic — the S015 competence probe covered Russian, French and Japanese. The AR arm
is that screen, and it is a gate, not a footnote.
5. The four questions, verbatim
Q2AR — arm AR.
Here is a sentence of classical Arabic prose, with one stretch of it marked ⟪ ⟫: … Look only at the marked stretch. Is the Arabic doing something with sound there — saj' (rhyme at the ends of matched cola), jinās, or muwāzana (matched word-shapes in matched position)? Answer Y or N, then a semicolon, then at most twelve words of reason.
Q1 — arms ORN.B, PLN.B. Verbatim from E-20260816 and E-20260816b.
Here is a passage from an English translation of a classical Arabic prose work. Look only at the span marked ⟪ ⟫. Is there conspicuous sound-patterning in the English there — alliteration, rhyme, matched word-shapes, or a word repeated for effect? Answer Y or N, then a semicolon, then at most twelve words of reason.
Q2 — all four English arms. Blind form verbatim from E-20260816b; the source form adds the
marked Arabic above it and points the question at the Arabic stretch.
[blind] … At the place in the original that this span translates, was the Arabic doing something with sound — a rhyme, a chime, or matched word-shapes? …
[source] Here is a sentence of classical Arabic, with one stretch of it marked ⟪ ⟫: … Here is a passage from an English translation of it, with the corresponding stretch marked ⟪ ⟫: … Look only at the marked Arabic stretch. Is the Arabic doing something with sound there …?
Q4 — arms ORN.S and PLN.S. New in v2, replacing v1's Q3 (finding 8: Q3 asked about
unwarranted impression and P4 claimed it measured fault; Q3 also presupposed that a device was
present, and was asked only where one was).
Here is a sentence of classical Arabic, with one stretch of it marked ⟪ ⟫: … Here is a passage from an English translation of it, with the corresponding stretch marked ⟪ ⟫: … Considering the Arabic, is the marked English a defect in the translation — something a reviser ought to change? Answer Y or N, then a semicolon, then at most twelve words of reason.
Q4 presupposes nothing about sound, and asking it in both arms turns P4 into a difference of
differences rather than a question whose answer is built into its wording.
6. Instrument controls
The four third-party Q1 controls of E-20260815, carried verbatim: CTL.POS (Dickens's fog
anaphora, expect Y), CTL.NEG (Anderson, expect N), CTL.ARCH (Genesis 22, expect N — archaic
register with no sound device), CTL.ARCH+ (Morris, expect Y — archaic register with one).
CTL.POS or CTL.NEG failing aborts the run; CTL.ARCH/CTL.ARCH+ not separating withholds
every Q1 number. All four were unanimous and correct at RS-20260816b.
7. Cells and bodies
| prompt | arms | items | seats | bodies |
|---|---|---|---|---|
Q2AR |
AR |
21 | 3 | 63 |
Q1 |
ORN.B, PLN.B |
21 | 3 | 126 |
Q2 |
ORN.B, PLN.B, ORN.S, PLN.S |
21 | 3 | 252 |
Q4 |
ORN.S, PLN.S |
21 | 3 | 126 |
Q1 controls |
— | 4 | 3 | 12 |
| total | 579 |
8. Pre-registered gate, predictions and failure criteria
Fixed before any body is dispatched. Rates are votes over (item × seat): denominators 30 (FIG), 15
(MEM), 18 (NONE). Every figure is computed site-level and cluster-level (§3.3) and under both
labellings (§3.4.2).
MC1 — THE GATE, and it is now a direct measurement. In AR — Arabic alone, locus marked, no
English anywhere — Q2AR(FIG) − Q2AR(NONE) ≥ 0.40.
If
MC1fails,P2,P3andP4are reported as WITHHELD, and the one licensed statement is that these three seats do not identify this inventory's figures in the Arabic — a fact about the panel, reported as such.
P1 — the replication, source-blind. In ORN.B: |Q2(FIG) − Q2(NONE)| ≤ 0.20, and
Q2(NONE) ≥ 0.40.
The second clause is the critic's finding 9 and it matters: a zero gap at 0.05/0.05 is not the
RS-20260816bresult, it is a jury answering N to everything.P1passes only if both clauses hold; a small gap with a lowNONElevel is reported as not a replication.
P2 — THE PRIMARY: does source access repair the discrimination? The interaction, not a bare
threshold (finding 7):
[Q2(FIG) − Q2(NONE)]_ORN.S − [Q2(FIG) − Q2(NONE)]_ORN.B ≥ 0.40.
v1 registered the source-visible gap alone, which the critic showed could pass trivially because the answer is in the prompt. The interaction is the quantity the question is actually about.
P3 — does the invented device still mislead a reader who is checking? At NONE sites,
Q2(ORN.S) − Q2(PLN.S) ≥ 0.20, paired by site.
A null here is the good news. It would mean source access is a complete remedy.
P3holding would mean the supplied device pulls a false belief about the original out of a reader with the original marked in front of them.
P4 — does the checking reader call the invention a defect? The difference of differences:
[Q4(ORN.S) − Q4(PLN.S)]_NONE − [Q4(ORN.S) − Q4(PLN.S)]_FIG ≥ 0.30.
If
P2holds andP4fails, the finding is that the reader sees the invention and does not mind it, and the framework's warning is about the uninformed reader only. Both readings are written here before the run because both are publishable.
Reported, not registered: every MEM figure; Q1 in both blind arms, including the PLN.B
residual-device rate that defines §3.2's subset; Q4 levels as well as differences; per-seat
breakdowns; free-text reasons.
9. Statistics, and what they are not
Randomisation scores are exact enumerations — over the relabelings of which sites are FIG for a
between-stratum gap, over 2^n sign flips for a paired difference — computed by verify.py
independently of analyse.py. They are descriptive: the loci are authored, the seats are three
fixed models at temperature 0, the sites were selected by one hand's ornamenting reflex, and the
sites are clustered by Arabic sentence. No number here is a frequentist P-value for a causal
claim (finding 21). RS-20260816b §9's measured sitting-to-sitting spread — 0 to 2 loci out of
13 on this instrument — is the scale at which to read every difference.
10. Procedure
build_materials.py→materials/items.json, with check F and the Arabic-locus assertion.cells.py→cells.json, asserting: 21 items; strata 10/5/6; 14 clusters; the four controls; everyPLNitem character-identical to itsORNtwin outside the marked span; exactly one⟪…⟫pair in every English and Arabic string; every Arabic locus present verbatim in its sentence; 579 bodies.- Independent pre-run critic pass — done,
NEEDS REDESIGN, §12 — and the rebuild it forced. run.pydispatches in the gate orderAR→ORN.S→PLN.S→ blind arms →Q1→Q4, one body at a time, appending torun.jsonlwith provider, finish_reason and the cost of every attempt including discarded ones (note (boe), fixed in this runner).analyse.py→analysis.json;verify.pyrecomputes every reported number fromrun.jsonlby an independent path, with mutation tests.
11. Budget
Pre-flight built from max_tokens, not from an assumed output length (note (abc)): 579 bodies,
caps P1 400 / P2 1400 / P3 800 completion tokens, prompts 120–700 tokens. Worst case
$2.40. Expectation from RS-20260816b's realised $0.002714 per body, plus Arabic on 315 of the 579
prompts: ≈ $1.70. The pre-run critic cost $0.121977. Today's UTC ledger stood at $1.187735
of $5.00 before this session, so the worst case leaves $1.29 unspent. If the run passes $2.40 it
halts and the completed arms are reported.
12. The pre-run critic, and what was done about it
One seat, openai/gpt-5.6-terra (P1), non-Anthropic, shown v1 of this design and all 21 items with
their Arabic, their truth label and both wordings. Verdict NEEDS REDESIGN; 28 findings, 8
BLOCKING. critic.json, $0.121977, two calls (the first hit the 9,000-token cap and a
continuation was budgeted, note (bnr)).
Accepted, and the design rebuilt:
| # | finding | what changed |
|---|---|---|
| 1 | a truth label contradicts the census's own inventory | O12 NONE → MEM; strata 10/5/6, all bars recomputed |
| 2 | PLN is not device-free and v1 said it was |
the claim withdrawn (§3.2); Q1(PLN.B) becomes a measurement, and every number is reported again on the low-residual subset |
| 4 | the paragraph window leaks other unmarked devices | window narrowed to the sentence |
| 5 | sites sharing an Arabic sentence are not independent | cluster-level analysis registered alongside site-level, and a finding is called only if both agree |
| 6 | the Arabic locus was not marked, so Q2 was a scanning task |
the Arabic stretch is now marked in every source arm |
| 7 | P2 could pass trivially; no source-access interaction was registered |
P2 is now the interaction [FIG−NONE]_source − [FIG−NONE]_blind |
| 8 | Q3 did not measure fault and presupposed a device |
Q3 replaced by Q4, presupposing nothing, asked in both source arms; P4 becomes a difference of differences |
| 9 | P1 could pass by the jury answering N to everything |
P1 gains a level clause, Q2(NONE) ≥ 0.40 |
| 21 | randomisation scores given more weight than they carry | §9 restated: descriptive, not inferential |
| 22 | the panel has never been shown able to read Arabic figures | the AR arm added and made the gate |
| 25, 26 | several self-descriptions false; do not dispatch v1 | v1 not dispatched; this is v2 |
| 28 | the theoretical conclusion outruns the evidence | all claims narrowed to these seats and these materials; the normative sentence struck |
Overruled, with reasons:
- Finding 1 on
O15andO20— that matched imperfect verbs are conjugational morphology and not a figure. The frozen inventory rule admits muwāzana by matched morphological pattern, and its exclusion is about rhyme produced by a grammatical ending. Re-reading the rule after seeing which sites it admits is what the freeze exists to prevent. Registered as a sensitivity instead (§3.4.2): every number reported under the critic's labelling too. - Finding 3, that the ORN/PLN pairs differ semantically — accepted as true and irreducible, and written into §3.4.1 as a limit rather than repaired, because repairing it means abandoning the translator's own frozen alternatives and writing new ones for the experiment, which is the defect note (bpn) exists to name.
- Finding 23, run every cell twice — refused on cost. The sitting-to-sitting spread is measured
(
RS-20260816b§9) and is used as the noise scale instead. - Finding 24, randomise item order within seat — refused. Calls are stateless and independent, so order cannot leak between them; dispatch order is fixed and published in §10.4 because it is a gate order, which is worth more here than randomisation buys.
- Finding 28's strongest form, that nothing about translation can be concluded — the narrowing is
accepted; the position that a model panel's behaviour on these materials is uninformative about
translation is not, and is the standing position of
config/models.md's calibration state, which every page here already carries.