Repository path: workshop/experiments/E-20260816-answering-figure/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260816-answering-figure |
| status | frozen |
| created | 2026-08-16 |
| updated | 2026-08-16 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-answering-figure.md, workshop/translations/kalila-saih/R30-v1/translation.md, workshop/regimes/R30-sound-plain.md, workshop/regimes/R31-answered-figure.md, workshop/regimes/R32-supplied-figure.md, wiki/findings/results/RS-20260815-supplied-sound-confirm.md, framework/v0.2/README.md, config/models.md, config/budget.md |
E-20260816 — can a reader tell a compensation from an invention?
ARM-answering-figure step 1. Frozen before any body is dispatched. Amendments, if the pre-run
critic forces any, are appended to §11 with the commit that made them.
1. The question
The handbook's oldest positive instruction is compensate: where the source has a figure the
target language cannot reproduce in kind, put a device of the target's own at that place. This
project has just measured a published hand doing it — RS-20260815 §6: at the four loci where the
Arabic repeats a sound, Lane answers with an English device at 3 of 4 and Burton at 1 of 4 —
and framework/v0.2 §7.8 carries the opposite instruction for supplied ornament: a mark the
reader can see that the source's reader could not is a defect.
Nothing in the framework says which of those two a compensation is. The defence of compensation has a testable consequence: a reader of the English alone should be readier to believe the original was doing something with sound where it actually was. If the same device planted where the source is plain produces the same belief, the compensation has transmitted texture and no information.
The measured question. Does a sound device placed where the source has a figure raise a reader's inference that the source had a figure there, by more than the same device placed where the source has none?
2. The wire between the limbs
Translating «باب السائح والصائغ» under R30 produced (i) the exhaustive inventory of the Arabic's
sound figures, made from the Arabic alone, and (ii) an English that carries none of them — so the
study limb can put the devices back at the right places and at the wrong places and change nothing
else. The translation is not illustration here; it is the only way to obtain a text that differs
from itself in exactly one property.
3. Materials
- Base arm
R30—T-kalila-saih-R30-v1, committed atbd87990before this design existed. 998 Arabic words → 1,862 English.contamination: noneagainst Knatchbull 1819 (21 shared 7-grams, 0 twelve-grams, longest run 11), measured before any locus was selected. SITEloci, 13. Every manipulable figure in the inventory, with no selection among them. The inventory has fourteen;F1is excluded by the translator's logD4as free-crossing (English reproduces comparative anaphora natively, so no arm could differ there), and that exclusion is on the frozen translation page.PLAINloci, 13. Chosen mechanically, not by the translator. Numbering warning, and it caught the critic out (§11 finding 1): the figure inventory numbers ARABIC sentences and the decoy pool numbers ENGLISH sentences of theR30rendering. They are different sequences. Algorithm, now stated so it can be re-run: unit = one English sentence of theR30rendering; candidate pool = every English sentence that (i) contains no enumerated figure — English sentences 5, 8, 9, 12, 15, 16, 28, 39, 47, 50, 62, 63 removed — (ii) contains no within-sentence figure of an excluded class — English 6, 7, 10, 13, 32 removed, being place/place, kinsman/kinship, knows/known, the Arabic cognate at 13, went off/went off — and (iii) is 15–40 words; pool size 29; then, taking theSITEloci in copy-text order, greedily and without replacement, the candidate minimising |Δwords|, ties broken by the lower English sentence index. Realised match: 29/29, 24/24, 32/32, 26/26, 30/28, 25/25, 35/35, 24/24, 31/32, 33/34, 24/24, 21/22, 18/18. Selection recorded inmaterials/loci-base.json.- The displayed unit is the locus span itself, not its sentence, and
cells.pyasserts that no two loci overlap and that noPLAINspan occurs inside anySITEspan. FourSITEspans are sentence fragments, because two figures share one English sentence in each case (F2/F3andF13/F14); each is displayed alone and carries one device only. - Arm
R31(answered), one device at eachSITElocus. ArmR32(supplied), one device at eachPLAINlocus. Every edit inmaterials/arms.jsonwith its device named. - Device classes are matched between the arms as a multiset: 9
MAT(matched pattern / alliterating pair), 2REP(a word or phrase repeated in matched position), 2RHY(rhyme).cells.pyfails if the two multisets differ. Per-index pairing is not claimed; the twoRHYpairs are matched on proximity instead —F11andP12adjacent,F10andP10a clause apart. - Every edit is a substitution inside the existing syntax; no edit adds a new member, and no edit
is a pure addition (critic BLOCKING 4). Word deltas:
SITE−1 over 13 loci,PLAIN+13, the largest single delta being +3. Every edit is written out inmaterials/arms.jsonwith the words it replaced. - Two further labels are carried on every
SITElocus and reported as subsets:stratum— frame discourse (9) or narrative (4) — andmorph—LEX(11), where the Arabic figure survives without its case endings, againstINF(2,F3andF4), where the figure is carried by inflectional endings in parallel syntax and might be obligatory morphology rather than chosen sound-work (critic MAJOR 10). - External panel (secondary). Lane 1839 and Burton 1885 at
E-20260815's fourSOUNDloci and its fourPLAINlociPL5–PL8, reused byte-for-byte from that experiment's stored cells. ErratumE1, entered 2026-08-16 before stage 4 was dispatched: this line first called all fourPLAINloci narration. Three are (PL6–PL8);PL5is direct speech, andRS-20260815§5 turns on exactly that distinction. Corrected here rather than in the reading.
Counts, stated separately (critic MINOR 13): 26 textual loci; 52 locus-by-version variants; 4
Q1 controls; 16 external cells; therefore 52 Q1 + 52 Q2 + 4 Q1 + 16 Q2 = 124
instrument-specific cells, and 372 bodies at three seats.
4. The two instruments
Q1 — the validated one, reused verbatim from E-20260814g / E-20260815, where it came back
unanimous on four third-party controls including the ARCH/ARCH+ pair built to break its
archaic-register confound. It asks about the English: does this passage use conspicuous
sound-patterning? Here it is the manipulation check: it decides whether the devices are audible
and whether the base arm is plain.
Q2 — the primary instrument, new. It asks about the source: judging from this English
alone, was the ORIGINAL doing something conspicuous with sound at this point? This is the inference
framework/v0.2 §7.8's diagnostic is about, put as a closed question.
Both are one body per (cell, prompt, seat), temperature: 0, blind to the Arabic, to the arm, to
the locus class and to every other cell. Seats P1 P2 P3 (config/models.md); P5 is out on
any task shape, note (bne). Majority of three.
5. Quantities
For a locus set K ∈ {SITE, PLAIN} and version v ∈ {base, dev}, q2(K,v) is the
proportion of the 13 loci whose Q2 majority is Y, and q1(K,v) the same for Q1.
lift_SITE = q2(SITE,dev) − q2(SITE,base)— what a compensation buys.lift_PLAIN = q2(PLAIN,dev) − q2(PLAIN,base)— what an invention buys.Δ = lift_SITE − lift_PLAIN— the primary.
Δ is reported three ways, all registered here (critic MAJOR 11: on 13 loci a majority rate moves in steps of 1/13, so a thresholded point estimate alone cannot carry an equivalence claim):
Δon majorities, as above — the headline number.Δ_bodyon the 156 individualQ2ratings (13 loci × 3 seats × 2 versions × 2 arms), which does not throw away seat-level information and moves in steps of 1/39.- An exact permutation test: the 13 paired locus differences
d_i = Y(dev) − Y(base)are formed per arm from the seat means; the observedΔis compared with the distribution ofΔunder all 2¹³ = 8,192 sign-flips of the arm label on the paired differences. Exhaustive, deterministic, no random numbers.
Subset reports, all pre-declared: the narrative stratum (SITE 4, PLAIN 11), the
frame stratum (SITE 9, PLAIN 2), the LEX subset of SITE (11 of 13), and the
length-matched subset (loci whose word delta is within ±1 in both arms). Each is underpowered
and each is reported with its n.
6. Predictions, registered
| id | prediction | bar |
|---|---|---|
M1 |
the devices are audible | Q1 Y at ≥ 9 of 13 SITE dev and ≥ 9 of 13 PLAIN dev |
M2 |
the base arm is plain | Q1 Y at ≤ 3 of 26 base cells |
M3 |
the two manipulated arms are equally audible (critic BLOCKING 5) | \|q1(SITE,dev) − q1(PLAIN,dev)\| ≤ 0.155 (≤ 2 of 13) and no device class differing by more than one locus |
P1 |
a compensation raises the source inference | lift_SITE ≥ +0.30 |
P2 |
an invention raises it too | lift_PLAIN ≥ +0.30 |
P3 |
PRIMARY — placement carries no information | \|Δ\| ≤ 0.15 and permutation P > 0.20 |
P4 |
structure alone transmits nothing | q2(SITE,base) − q2(PLAIN,base) ≤ +0.15 |
P5 |
secondary, underpowered — Lane's compensations read as source figures | Lane Q2 Y rate higher at EXT-SOUND than at EXT-PLAIN |
P3 is the prediction this design exists to test and the lead's prediction is the null. Two
readings are registered in advance, and which one is reported is decided by the numbers, not after
seeing them: |Δ| ≤ 0.15 and permutation P > 0.20 → not distinguishable on this
instrument; Δ ≥ +0.25 and P ≤ 0.05 → placement carries. Anything else is reported as
measured with no rule firing.
What the null branch may and may not be said to license (critic BLOCKING 6). If Δ ≈ 0, the
licensed sentence is: on this instrument, the inference a reader draws about the source tracks
the presence of English sound-patterning and not its placement. It is not licensed to say
that compensation is worthless, that readers ignore location in general, or that Q2 is a valid
instrument for locating source figures — Q2 has no third-party positive control, and building one
(a comparative two-alternative task on passages with verified source status) is named here as the
successor design, not claimed as done.
7. Failure criteria
F1—Q1controlsPOSorNEGnot unanimous on the expected code → abort, nothing is reported.ARCHorARCH+not unanimous → every rate withheld, table still printed.F2—M2fails → the base arm is not plain,R30did not do its job, and every lift is withheld.F3—M1fails on either arm → that arm's devices were not heard, so its lift measures nothing, andP3is withheld.F3b—M3fails →P3is withheld (critic BLOCKING 5). An arm difference in audibility confounds placement with device strength, and the run says so rather than noting it in a limits section.F4— seat unanimity across allQ2cells below 0.60 → the primary is withheld as instrument noise (the floorE-20260815used, where the realised value was 0.800).- No gate is moved after it fires. A gate that fires is reported as having fired.
8. Known confounds, declared before the run
- The decoy arm is longer.
R32adds +13 words over its 13 loci;R31−1. The invention arm therefore has marginally more material to be heard in, which makesP3's null easier to obtain. A length-matched subset is reported beside the full figure. - One hand wrote both manipulated arms, knowing the hypothesis. This is unavoidable: the
manipulation is the arm. It is why the class multiset was fixed before any device was written,
why every edit is a substitution rather than an addition, and why
M3is a binding gate. - The
SITEloci are not randomly located, and are not matched to thePLAINloci on discourse mode (critic MAJOR 8): 9 of 13SITEloci are in the philosopher's frame discourse against 2 of 13PLAIN. The lifts are within-locus differences, so a constant discourse effect cancels in each lift and cannot by itself produceΔ; what survives is a possible discourse × device interaction, and the stratified figures in §5 are reported for it. - Seats are models, not readers. No sentence in the result page will say otherwise.
Q2has never been validated; see §6's licensing paragraph. Its floor is the base arm, its only external evidence is the 16-cell published panel atn= 4 per cell, and that panel is secondary by declaration.- Two
SITEloci (F3,F4) rest on Arabic inflectional endings in parallel syntax and may be obligatory morphology rather than chosen sound-work; theLEXsubset excluding them is reported (critic MAJOR 10). Q1measures conspicuity only, not naturalness or intrusiveness (critic MINOR 12, accepted in part): a device can be conspicuous because it is intrusive. SplittingQ1into three items would triple the manipulation-check cost and is refused for this run; the consequence is that no sentence here may say a compensation is good English, only that it was heard.
9. Procedure and dispatch order
python3 cells.py --dump— verifies every base span against the frozenR30text and the class multisets, then writescells.json. Non-zero exit stops the run.- Stage 1 — the four
Q1controls, 12 bodies. Scored before anything else is bought (F1). - Stage 2 —
Q1on all 52 lead cells, 156 bodies.M1/M2scored. - Stage 3 —
Q2on all 52 lead cells, 156 bodies. - Stage 4 —
Q2on the 16 external cells, 48 bodies. Last, because it is secondary. - Append-and-resume (note (bnx)): every body appended to
run.jsonlas it returns; a restart buys nothing twice. - Judgment is not parallelised.
10. Budget
Pre-flight, built from max_tokens and not from an assumed output length (note (abc)):
caps P1 400, P2 1400, P3 800; list prices config/models.md.
| per body worst case | bodies | worst case | |
|---|---|---|---|
P1 |
$0.0028 | 124 | $0.35 |
P2 |
$0.0056 | 124 | $0.69 |
P3 |
$0.0056 | 124 | $0.69 |
pre-run critic (P1) |
1 | $0.06 | |
| ceiling declared | 373 | $1.79 |
Day 2026-08-15 stands at $1.516321 of $5.00; headroom $3.483679. The ceiling fits with
$1.69 to spare. Stop-loss: if the running total passes $1.79 the run halts and reports what it
bought. Expected actual, from E-20260815's realised $0.00267 per body: ≈ $1.00.
11. Amendments — the pre-run critic pass
openai/gpt-5.6-terra (P1), one call, 15,461 characters, finish_reason: stop, $0.046800.
Verdict NEEDS REDESIGN, thirteen findings, six of them BLOCKING. The critic was shown the
frozen design, the frozen Arabic figure inventory, and every edit of both manipulated arms —
which is what let it attack the arms' matching, the design's central claim. Every finding is
recorded below with what was done. critic.json holds the text verbatim.
| # | severity | finding | disposition |
|---|---|---|---|
| 1 | BLOCKING | three PLAIN loci allegedly contain enumerated figures (P8↔F8, P4↔F9, P5↔F11) |
ACCEPTED AS A DEFECT OF THE DESIGN, REFUSED ON THE FACTS. The collision is between two numbering systems the design used and did not distinguish: F8 F9 F11 are at Arabic sentences 24, 35, 45 and P8 P4 P5 are at English sentences 24, 35, 45. P8 renders Arabic 19–20, P4 Arabic 31, P5 Arabic 41 — no overlap. §3 now warns about the two sequences, states the algorithm in full, and cells.py asserts mechanically that no PLAIN span lies inside any SITE span and that no two loci overlap. A finding that is wrong because the design was unreadable is a finding |
| 2 | BLOCKING | multi-figure sentences (F2/F3, F13/F14) put two devices in one displayed unit |
ACCEPTED IN PART. The displayed unit is the locus span, not the sentence, and the two spans are disjoint; this is now said in §3 and asserted in cells.py. The residue — four SITE spans are sentence fragments — is declared |
| 3 | BLOCKING | the edits change meaning; eleven examples given | ACCEPTED IN FULL, AND IT IS THE MOST USEFUL FINDING. Every edit was re-screened for denotation. Withdrawn: F5 loss→rot, F10 told me→made me know, P5 took→stole, P12's loss of food, P13's change of population, P4's fell flat. Rewritten to substitutions of near-synonyms: F3 (4 substitutions → 2), F6 (2 → 1), F7, F8, F13. The devices are weaker for it and the arms are cleaner, and M1 may now fail — which is the honest exposure |
| 4 | BLOCKING | class labels do not match device quality; several PLAIN edits merely add redundant material while SITE edits substitute |
ACCEPTED IN FULL. P2, P7, P11, P14 rebuilt as substitutions; no edit in either arm is now a pure addition, and the PLAIN word delta falls +21 → +13 against SITE −1. The two RHY pairs are additionally matched on proximity |
| 5 | BLOCKING | M1 cannot show the arms are matched, yet P3 depends on it |
ACCEPTED IN FULL. New binding gate M3 and new failure criterion F3b: \|q1(SITE,dev) − q1(PLAIN,dev)\| ≤ 0.155 with class-level agreement, or P3 is withheld |
| 6 | BLOCKING | the null can be produced by the trivial heuristic English sound ⇒ source sound, so Δ ≈ 0 does not establish that placement carries nothing |
ACCEPTED IN FULL. §6 now states exactly what the null branch licenses and what it does not, and names the comparative two-alternative validation as the successor design rather than pretending it is done |
| 7 | MAJOR | the decoy rule is under-specified and not reconstructable | ACCEPTED. Full algorithm in §3: unit, pool, three filters, greedy without replacement, tie-break, pool size |
| 8 | MAJOR | SITE and PLAIN confounded by discourse mode |
ACCEPTED IN PART. Stratified reports added to §5 and the confound to §8, with the reason a constant discourse effect cannot produce Δ (both lifts are within-locus). Rebuilding the decoy set to match on discourse would have cost the mechanical selection rule, which is a worse trade |
| 9 | MAJOR | several device descriptions are wrong or unstable | ACCEPTED. Every device string rewritten to describe what the revised edit actually does, with the replaced words listed in an edits field |
| 10 | MAJOR | the inventory may count obligatory inflection as sound-work | ACCEPTED. F3 and F4 labelled INF; the LEX subset (11 of 13) is a pre-declared report |
| 11 | MAJOR | 13 loci give Δ a granularity of 0.077, so \|Δ\| ≤ 0.15 is a one-locus band and cannot carry an equivalence claim |
ACCEPTED IN FULL. §5 adds the body-level Δ on 156 ratings and an exact 8,192-fold sign-flip permutation; P3's null branch now requires P > 0.20 as well as the band |
| 12 | MINOR | Q1 should be split into conspicuity, naturalness and intrusiveness |
REFUSED, with the reason on the record. It triples the manipulation-check cost for a property no registered quantity uses. The consequence is written into §8.7: nothing here may say a compensation is good English |
| 13 | MINOR | "72 cells" conflates variants with instrument cells | ACCEPTED. §3 now states all five counts separately |
Amended and re-frozen before any body was dispatched. The critic saw no ratings, because none existed.