Repository path: workshop/experiments/E-20260816h-target-set/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260816h-target-set |
| status | frozen |
| created | 2026-08-16 |
| updated | 2026-08-16 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-invented-ornament.md, wiki/findings/results/RS-20260816c-checked-ornament.md, wiki/findings/results/RS-20260816b-invented-figure.md, workshop/translations/kalila-labwa/R37-v1/translation.md, workshop/regimes/R37-declared-rule.md, framework/v0.2/README.md, config/models.md, wiki/goodness-senses.md |
E-20260816h — which rule fixes the target set: what a reader of the Arabic says a translator is obliged to answer
ARM-invented-ornament step 2, first half. Study limb of the paired unit whose translation limb
is T-kalila-labwa-R37-v1, frozen at 3ee5001 before this design existed.
v2, frozen after a pre-run adversarial critic returned NEEDS REDESIGN, 11 findings, 3
BLOCKING. v1 was never dispatched. Nine findings were accepted and implemented in the six
amendments of §12; two were overruled in writing. The critic also read the frozen figure
inventories and named three misclassifications in them; two were right and the third forced an
amendment that is right, and they are corrected on the translation page at §3.0 rather than
here.
1. The question
RS-20260816c failed its own gate and the failure was the finding: the instruction "where the
source has a sound figure you cannot reproduce, put a device of your own at that place"
presupposes that the places where the source has a sound figure is a determinate set, and on
one chapter of «كليلة ودمنة» the project's inventory and three readers of the Arabic picked out
sets that agreed on 5 of 10 figures. The disagreement was three nameable rules, not noise.
The translation limb has since written both rules out as rules and applied each to a fresh chapter of the same work: after the erratum this design's critic forced (translation §3.0), Rule A admits 18 loci, Rule B admits 11, the union is 24 and the intersection is 5.
This experiment asks which of the two rules a reader of the Arabic is applying when the question is put in the translator's own terms: is this a place the source obliges my English to answer?
What the unit teaches about translating literature (continue-prompt.md §4.5): it tells a
translator, on a language pair where the source's characteristic ornament cannot be reproduced,
which places in the source carry an obligation — and if the answer is that two defensible rules
disagree at four fifths of the licensed places, that a declared rule is part of the translation, not a
preliminary to it. The object measured is the Arabic and what it obliges; the project's inventory
is one of the two rules under test, not the instrument doing the testing.
2. Materials
../../translations/kalila-labwa/copy-text-adopted.txt, «باب اللبؤة والإسوار والشغبر», 599 words
in 37 sentences, collated whole against the Amīriyya 1937 printing.
Every locus comes from the frozen inventories of T-kalila-labwa-R37-v1 §3, built before this
design existed. cells.json is generated by build_materials.py, which asserts that every fully
marked sentence reproduces the adopted copy-text character for character when the markers are
stripped, that every locus carries at least two marked members, and that every dispatched string
carries exactly that locus's members and no brace. The class is never in the dispatched string.
28 study loci in 20 sentences, after the critic pass. Class sizes, with the count of loci
carrying exactly two marked members in brackets — that subset is the member-count-matched
contrast registered at PR2 (critic BLOCKING 1):
| class | what it is | n | (2-member) |
|---|---|---|---|
BOTH |
both rules admit it | 5 | 4 |
A-REPEAT |
Rule A only — a word repeated unchanged in matched position | 7 | 6 |
A-SHAPE |
Rule A only — a matched morphological pattern that does not rhyme | 5 | 3 |
B-RHYME |
Rule B only — members ending in the same sound by a grammatical ending or enclitic | 3 | 2 |
B-ROOT |
Rule B only — a root in two or more different shapes | 3 | 2 |
NONE |
neither rule admits it — a doublet of matched members with no sound relation | 5 | 5 |
Five loci of the inventory are DROPPED from the study set, each under a rule stated before any
dispatch: a locus whose class cannot be assigned unambiguously, or whose span carries a second
figure of another class that marking cannot separate, is dropped. They are F20 (its three cola
also rhyme on ر at two of three ends, so a verdict would not identify a class), F8, F17, F23
(the three the critic named — translation §3.0) and N6 (جوراً / وظلماً, an accusative pair with
the same consonantal skeleton, which the critic's MAJOR 7 correctly refused as a hard negative).
cells.json carries them as DROPPED and no body is bought for them.
Three third-party controls, none from this chapter — the critic's MAJOR 9 asked for a
negative matched to the study's own structure, and C3 is it:
C1CTL.POS— al-Ḥarīrī, al-Maqāma al-Ṣanʿāniyya, «خاويَ ⟪الوِفاضِ⟫ … باديَ ⟪الإنْفاضِ⟫», overt saj'. Expectsoundworktrue,owedtrue,relationRHYME.C2CTL.NEG— «كليلة ودمنة/باب القرد والغيلم», ⟪كبر⟫ … ⟪المملكة⟫, an arbitrary pair with no relation of sound, root or pattern. Expect false, false,NONE.C3CTL.NEG, structurally matched — the same sentence, ⟪ماهر⟫ … ⟪القردة⟫: two nouns in a parallel construction, the shape of the study's ownNONEstratum, with no sound relation. Expect false, false,NONEorPARALLEL.
3. Procedure
One body per (seat, LOCUS) — the critic's BLOCKING 2, taken in full. v1 asked every group of
a sentence in one body, and the critic was right that displaying an A-only locus beside a
B-only one in the same clause manufactures exactly the contrast PR measures. Each dispatched
string is the whole sentence with one locus's members marked ⟪ ⟫ and every other marker stripped,
so no seat ever sees two loci together, and group letter, group count and display position carry no
information. build_materials.py asserts it. The seats see the Arabic sentence alone: no English,
no translation, no mention of this project's inventories.
Seats P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the same
three as RS-20260816c, so this is the same jury asked a sharper question, not a new one.
temperature 0.4, max_tokens 500, reasoning.max_tokens 150 — the content cap comes
down because one locus per body needs one small object, and the body count goes up. Note (bpx): the temperature
is above 0 and the seat, not the body, is the unit — one body per seat per locus, three
seats, no replicates. Note (bph): the reasoning cap is set well below the content cap.
Each body asks three things, in this order (critic MAJOR 8):
soundwork— is what holds between the marked stretches an audible effect the author has made, something a reader of the Arabic would hear as deliberate work with sound?owed— a translator is putting this chapter into English and cannot reproduce Arabic sound patterns directly; would he be right to treat this place as a sound effect of the original that his English ought to answer with a device of its own?relation— one label fromRHYME·ROOT·REPEAT·SHAPE·PARALLEL·OTHER·NONE.
v1 put the label list first and named exactly the five categories the classes are built from. The
critic's MAJOR 8 was that this makes owed a prompted consequence of recognising a supplied
category. v2 asks the two judgements first in their own words, puts the labels last, and adds two
options that are not inventory classes — PARALLEL (parallel in grammar or sense, not in
sound) and OTHER — which also answers the forced-choice half of MAJOR 7. The residual leak is
that a model sees the whole prompt at once; it is stated at §6.3 and it is why relation carries
no prediction.
owed is the dependent variable of every registered prediction. soundwork is registered as a
secondary with the same directions, and relation is descriptive.
3a. Frozen parse rule
A body must return JSON with a groups object containing every group letter of its sentence, each
with a relation in the closed list, a boolean soundwork and a boolean owed. Anything else is
invalid and is dropped from its denominators, reported as dropped, and never re-rolled.
3b. Dispatch order, and it is a gate
- Stage 1 — the three controls, all three seats, 9 bodies. This is the probe and the gate at
once (note (bps)), which is why no study locus is spent on probing; the critic's
MAJOR10 objected that v1 probed onS15, a sentence carrying both primary classes. - Parse/liveness: a seat returningfinish_reason: length, an empty body or an unparseable body on any control is dropped for the whole run, its bodies kept and reported, never re-rolled. Below two surviving seats the run aborts. - The control gate, widened by the critic'sMAJOR9 to coverowedand not onlysoundwork. Over the surviving seats:CTL.POS(C1, al-Ḥarīrī's saj') must besoundworktrue andowedtrue at ≥ 2 of 3; both negatives (C2an arbitrary pair,C3a structurally matched noun pair) must besoundworkfalse andowedfalse at ≥ 2 of 3. Any failure aborts the run before a single study body is bought, and the failure is the reported result. - Stage 2, the 28 study loci, surviving seats, one body each.
Append-and-resume: every body is written to run.jsonl as it returns.
4. Registered predictions
All rates are owe rates: over the loci of a class and over the surviving seats, the proportion
of owed = true among valid bodies. A-only = A-REPEAT ∪ A-SHAPE (12 loci);
B-only = B-RHYME ∪ B-ROOT (6 loci); the disputed set is their union, 18 loci.
-
P0— the level clause. It is not a withholding gate, and §5 says exactly what its failure costs (criticMAJOR11). owe(BOTH) ≥ 0.70 and owe(NONE) ≤ 0.35. -
PR— THE PRIMARY. owe(B-only) − owe(A-only) ≥ +0.40. Tested by exact one-sided enumeration over all C(18,6) = 18,564 relabelings of which 6 of the 18 disputed loci areB-only, on the per-locus mean owe value. Reported with its exact score.
What a positive PR licenses, written before the number exists (critic BLOCKING 1). It
licenses: the loci Rule B admits and Rule A does not are called owed more often, by these seats,
than the loci Rule A admits and Rule B does not, by more than a relabelling of the same eighteen
loci produces. It does not license: that the seats are "applying Rule B", that the effect
would survive independently sampled loci, or that any mechanism has been identified. The critic's
proposed rival — that enclitic chains and root echoes are simply more conspicuous than
repetitions and matched patterns — is not a rival to the hypothesis but a statement of it,
since the two rules differ precisely in whether audibility or matched form is what counts; the
page will say so and will not claim more.
-
PR2— the member-count-matched primary (criticBLOCKING1, the part that is a real confound). The same contrast restricted to loci with exactly two marked members —A-only9,B-only4 — so that aB-onlyadvantage cannot be bought by longer chains. Exact enumeration over C(13,4) = 715 relabelings. Registered as a condition onPR's reading: ifPRholds andPR2does not, the page reports that the contrast is confounded with member count and claims nothing about the classes. -
P1owe(A-REPEAT) ≤ 0.35.RS-20260816cO19: 0 of 3. P2owe(A-SHAPE) ≤ 0.35.RS-20260816cO15,O20: 0 of 3 each.P3owe(B-RHYME) ≥ 0.60.RS-20260816cO6,O10,O16: 7 of 8 valid.P4owe(B-ROOT) ≥ 0.60.RS-20260816cO9: 3 of 3.P5— the target-set verdict. Let R be the set of loci where a majority of surviving seats sayowed— and with two surviving seats "majority" means both, which is a stricter rule and is declared here rather than discovered (criticMAJOR10). Then J(R, RuleB) > J(R, RuleA), J the Jaccard index over the 28 study loci, with Rule A =BOTH∪A-REPEAT∪A-SHAPE(17) and Rule B =BOTH∪B-RHYME∪B-ROOT(11).
soundwork carries the same five class predictions as a registered secondary, reported beside
owed at every class.
Registered as descriptive, not predictions, because the design cannot power them:
- the class-map corroboration (the free half of the critic's
BLOCKING3): the proportion of study loci where the majorityrelationlabel falls in the set that corroborates the inventory's class (cells.py'sCLASS_RELATION). This is the only check in the design on whether one informed hand's class map describes anything outside itself, and a low value is reported as undermining every class rate on the page. It is not a gate, because the labels are prompted (§6.3) and a gate on a leaky measure is decoration. soundworkagainstowedat every locus: how often a reader says the sound-work is there and the translator need not answer it, or the reverse.- per-seat rates for every class, so that a class difference carried by one seat is visible
(critic
MAJOR10). - the two sentences that carry an
A-onlyand aB-onlylocus —S15(F10againstF11) andS32(F18againstF19) — as paired observations, now that they were never shown together.
5. Failure criteria
- Any control fails its gate (§3b.1) → abort before any study body, and the failure is the reported result.
- Fewer than two seats survive stage 1 → abort.
- A seat's invalid rate above 1/3 across stage 2 → that seat's bodies are reported and its rates printed separately; the seat is not dropped retrospectively, because dropping on an outcome is not a frozen rule. Invalid bodies are dropped from their denominators, counted, and never re-rolled (§3a).
PRfails → reported as a failure.PRat or near zero withdraws this arm's premise: the disagreement between the two written rules would then be one the readers do not reproduce, andframework/v0.2§7.15's indeterminacy would be a fact about two inventories rather than about what a source obliges.P0fails → every number is still computed and printed, and no class-based conclusion about which rule readers apply may be drawn,PR,PR2andP5included (criticMAJOR11). v1 said the class rates could not be read while leavingPRandP5standing, which was inconsistent.PRholds andPR2fails → the contrast is reported as confounded with member count.
6. What this design cannot establish
- Three fixed models, one chapter, one work, one language pair. Nothing here is a claim about
human readers of Arabic (charter §4). The jury is not calibrated (
config/models.md, Tier D NOT PASSED); everything isprovisionalandinternal-judgment-only. The three seats are not a sample from a population and no population inference is made:PR's exact score is a randomization score over locus labels, not a p-value about readers. - One hand wrote both rules, both inventories, the class map and the marking, knowing what the
second pass was for —
T-kalila-labwa-R37-v1D1. This is the critic'sBLOCKING3 and the design does not repair it. Two things stand against it and neither is enough: the two rules are written out as rules a second hand could apply, which is what makes the successor possible; and the class-map corroboration in §4 gives one free, weak, outside reading of the map. The repair is a second hand applying the two written rules to the whole chapter, blind, and it is this arm's named successor, not this design. - The label list is still in the prompt, after the last of the three questions. A model reads
the whole prompt before answering any of it, so
owedis not fully insulated from the vocabulary; the ordering and the two non-inventory options reduce the leak and do not remove it. This is whyrelationcarries no prediction and why the corroboration check is called weak. - This measures what a reader says is owed, not what a reader notices in a translation. The arm's constituting question — whether a source-visible reader can tell an invented ornament from a compensation — is the arm's second half and is untouched here.
NONEis not a no-figure control in the strong sense. All fiveNONEloci are doublets of matched members, chosen because that is whereT-kalila-nasik-R34-v1§4 found an ornamentalist inventing. They are the hard negatives, not the easy ones, so a lowNONErate is a stronger result than it looks and a high one is weaker.C3is the third-party version of the same shape.- Class sizes of 3 are small and the page will not hide it.
B-RHYMEandB-ROOThave three loci each, nine seat-locus verdicts apiece before any failure.P3andP4are read as descriptions of nine verdicts; the primary isPR, which pools them against twelve.
7. Budget
Worst case built from max_tokens, not from an expected output length (note (abc)).
Input ≤ 750 tokens, output capped at 500. Prices config/models.md: P1 $1.00/$6.00,
P2 $0.75/$3.75, P3 $2.00/$6.00 per M.
| worst case per body | |
|---|---|
P1 |
$0.00375 |
P2 |
$0.00244 |
P3 |
$0.00450 |
| per item, three seats | $0.01069 |
| × 31 items (28 study + 3 control) | $0.331 |
pre-run critic, one P1 body, billed |
$0.033371 |
| worst case, whole experiment | $0.365 |
Stop-loss $0.39, enforced inside run.py. Declared ceiling $0.42. Note (bpq): worst case <
stop-loss < ceiling. The UTC day 2026-08-16 stood at $4.388966 of $5.00 before this session, so the
ceiling leaves $0.19 of the day unspent.
8. Verification
verify.py recomputes every number on the result page from run.jsonl and cells.json,
including both exact enumerations, and re-checks that every dispatched string reproduces the
adopted copy-text when its markers are stripped. Mutation tests: perturb the class map, the parse
rule, the member-count filter and each enumeration, and confirm every one is caught.
12. The critic pass, and what was done with it
One P1 body, openai/gpt-5.6-terra, $0.033371, critic.json. NEEDS REDESIGN, 11
findings, 3 BLOCKING. P1 also sits on the jury; that is the arrangement E-20260816c used
and it is declared, not hidden — a critic that will later answer the items has seen the design, and
P1's own rates are reported separately at §4 for that reason.
Six amendments, implementing nine findings.
| # | from | what changed |
|---|---|---|
| A1 | BLOCKING 2, MAJOR 10 (part) |
One locus per body. Every other marker is stripped from the dispatched string; group letter, count and position carry no information. Body count 60 → 93; max_tokens 900 → 500 to pay for it. |
| A2 | MAJOR 4 |
Member-end defined: the last word of the colon, a colon delimited by a coordinating conjunction, a pause mark or the sentence end; in a bare doublet «X وY» the members are X and Y. F5, F7, F21 re-marked to member-final stretches only. |
| A3 | MAJOR 4, 6, 7 |
F17, F23, N6 dropped, F8 dropped, under the declared drop rule; the translation page carries the erratum at §3.0 and its headline counts move from 20/13/27/6 to 18/11/24/5. |
| A4 | MAJOR 8, MAJOR 7 (part) |
Prompt re-ordered: soundwork and owed first, in their own words; relation last, with PARALLEL and OTHER added to the label list. |
| A5 | BLOCKING 1 (the real part), MAJOR 10, 11 |
PR2 registered, the member-count-matched contrast; PR's licence written out before the number exists; P5's "majority" defined for two seats; per-seat rates required; P0 failure made consistent — it invalidates PR, PR2 and P5 too. |
| A6 | MAJOR 9, MAJOR 10 (probe) |
Third control added (C3, structurally matched negative); the control gate widened to owed; the controls are the probe, so no study locus is spent probing and no primary class is exposed to probe attrition. |
Two findings overruled, in writing.
BLOCKING1's remedy — "recastPRas descriptive, or redesign with independently sampled loci and a feature-conditioned model." Refused. The loci are the chapter's, not a sample, and the design says so; an exact enumeration over locus labels is a randomization score and is reported as one, not as a population inference. The remedy's own last clause concedes that even the redesign "would estimate associations, not identify a mechanism", which is exactly what §4 now saysPRlicenses. The finding's real content — that aB-onlyadvantage could be bought by chain length — is taken in full atPR2.MAJOR5 — "F6(وصاحت / وضجّت) is dubiously saj', because final ‑ت is inflectional." Refused. Rule B is explicitly indifferent to what produces the ending ("by a rhyming word, by a shared morphological ending, or by an identical enclitic pronoun") and Rule A's exclusion (a) is explicitly about enclitic pronouns only. That asymmetry is not an inconsistency in the design; it is the substance of the disagreement the whole arm is about. Recorded as a sensitivity:F6is one of fiveBOTHloci, and the page reportsBOTHwith and without it.