Repository path: workshop/experiments/E-20260821-matched-shape/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260821-matched-shape |
| status | frozen |
| created | 2026-08-21 |
| updated | 2026-08-21 |
| version | v2-post-critic |
| senses | style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-published-figure.md, wiki/base/anchors/A-knatchbull-kalila/README.md, wiki/findings/results/RS-20260816j-published-figure.md, workshop/translations/kalila-fanza/R39-v1/translation.md, workshop/regimes/R39-matched-shape.md, workshop/experiments/E-20260821-matched-shape/critic-response.md, framework/v0.2/README.md, config/models.md |
E-20260821-matched-shape — a second coding hand on the published translator's matched-shape figures, and a third chapter
One-sentence design. Give three blind model seats the whole of a Knatchbull 1819 chapter, one Arabic figure-locus at a time, and have each seat find where that material is rendered, quote the English, and name which formal relation — from a menu fixed before dispatch — holds between the English members; on the fourteen matched-shape loci of an already-censused chapter and on forty loci of a third chapter rendered and frozen this session.
Revised v2-post-critic after a pre-run adversarial pass returned NEEDS REDESIGN with 3 BLOCKING
findings, all accepted: critic-response.md. The window is abolished, the seat's threshold is
replaced by a reported feature, and the falsification-by-vote bar is struck.
1. Question and predictions
Question. framework/v0.2 §7.18 tells a practitioner that the one published English hand from
the Arabic answers 0 of 18 matched-shape figures — muwāzana, the class Knatchbull's own 1818
preface calls the sententious brevity of the Arabic and says he could not express. All
seventy-two of the anchor's three-way calls were the lead's, unchecked, made by the party holding
the hypothesis (RS-20260816j §7.2). What a second coding hand finds at those loci, and what it
finds on a third chapter, is unknown.
The subject rule (continue-prompt.md §4.5), in one sentence: it teaches what a published
translator does at the place where the source's prose makes a shape English has no morphology for —
and, paired with the translation limb, whether that shape can be carried at all. The materials
include this project's own inventories; the question is about a translator and about English.
Predictions, registered before any census body was dispatched
| # | prediction |
|---|---|
P1 |
Seat-majority strict match on the fourteen Set A matched-shape loci is 0 to 3 of 14 |
P2 |
Among loci the seats find rendered, seat-majority match on Set B's matched-shape loci is below that on Set B's non-matched-shape loci. Descriptive; 16 non-SHAPE items give little precision and no test is run |
P3 |
On the six non-PLAIN cells of the lead's Set B coding (F9 F15 F26 ABSENT; F59 F91 F94 ANSWERED — six of these seven, F94 being the sixth-and-seventh borderline), the seat majority agrees on at least 4. Overall agreement is reported against the trivial all-PLAIN baseline of 34/40 = 85%, which any constant answerer would reach |
P4 |
F59 and F91 — the two matched shapes the lead's own rendering flagged as free, where English's ordinary wording already matches — are the two Set B SHAPE loci most likely to come back matched from the seats as well |
P5 |
Seats are less unanimous on matched-shape loci than on the other three classes |
No prediction appears in any prompt. No seat is told the work, the translator, the century, the existence of a hypothesis, any count, or which class the design cares about.
2. Materials
materials/items.json, built by materials/build_items.py; 54 items.
- Set A — 14 items: every locus of «باب القرد والغيلم» whose frozen inventory ground is
SHAPE(T-kalila-qird-R38-v1§3:G2G5G8G11G19G23G26G28G29G31G48G50G52G54). - The crosswalk, stated because the critic was right that it is an inference.
T-kalila-qird-R38-v1§3.1 reports SHAPE is a ground at 14 loci;A-knatchbull-kalila§3.1 reports 18 pooled across the two chapters, all 54 qird loci being codable and six labwa loci lost to a scan defect. Those 14 are therefore the qird contribution, and the published 0 of 18 entails that each is on record as not answered. The assumption is that the anchor's ground-counts were computed from the two inventories' own class labels; it is auditable from the two pages and from nothing else, because the per-locus calls were not stored. - «باب اللبؤة والإسوار»'s four
SHAPEloci are excluded, because its inventory is recorded under Rule A / Rule B rather than under the four grounds, and identifying which four the anchor counted would require a fresh classification by the hand that already knows the aggregate. Set A therefore covers 14 of the published 18 and is not reported as validating the 18. - Set B — 40 items from «باب ابن الملك والطائر فنزة», rendered whole this session
(
T-kalila-fanza-R39-v1), inventory frozen at commit1c749ff0before any English existed: all 24SHAPEloci, and 16 of the 73 non-SHAPEloci drawn bysha256('E-20260821-matched-shape|' + id)—F1F4F6F9F14F15F21F22F24F31F37F40F45F64F86F94. - Each item carries: the Arabic span, the Arabic sentence it sits in, a flat content-only gloss, the figure type with a neutral definition, the relation menu for that type, and the whole comparator chapter. Presentation is uniform across the two sets.
- The lead's Set B coding is committed before dispatch at
d6a4e14f, inmaterials/lead-coding.json, with its five close calls written out. Set A's lead coding is the published one and was made at S203.
2.1 The relation menu — what replaces the seat's threshold
The seat does not decide whether a figure is answered. It reports what it sees, from a menu fixed
before dispatch (build_items.py RELATIONS), and the analyser computes the readings:
| class | strict | loose only | not a match |
|---|---|---|---|
SHAPE |
identical-suffix · identical-inflection | same-class-and-length | frame-only · none |
RHYME |
rhyming-words · identical-ending | same-pronoun-or-particle | none |
REPEAT |
same-content-word · same-phrase | same-frame-varied-words | none |
ROOT |
same-stem-two-forms | etymological-pair | none |
Every reported rate is printed twice, strict and loose. The lead's own three close calls fall
out of this menu mechanically: F59 (‑tion/‑tion) is strict, F38 (‑ed/‑en) is not identical
inflection, F34 (‑ness/‑ment) is not identical suffix.
3. Procedure
- Stage A. 54 items × 3 seats = 162 bodies. Seats
P1(openai/gpt-5.6-terra),P2(google/gemini-3.6-flash),P3(x-ai/grok-4.5) — the panel lessP4(note bps) andP5(note bne), both out. Deterministic shuffle.temperature: 0. One call per item per seat; judgement is never batched. - Each body returns
{"status": "rendered"|"omitted"|"unlocated", "english": "...", "members": "...", "relation": "<from the menu>", "reason": "..."}.unlocatedis a distinct state fromomitted(critic finding 1) and is counted separately, never folded into either. - Re-dispatch each dead body once at the doubled cap (note bgk); still dead ⇒
dead: trueand that cell is void. - The seat verdict for a locus is the majority of three on
status, and separately on strict-match / loose-match / no-match. No majority ⇒ counted as such, never broken by the lead. verify.pyrecomputes every reported number fromrun.jsonlby a path sharing no code with the runner, including the confusion matrix, the trivial baseline, and the exact binomial bounds.
4. Gates and failure criteria
F1— majority availability. If fewer than 36 of 54 items have a 2-of-3 majority onstatus, no rate is reported as a census; the disagreement is the finding.F2— contamination. Discharged before this design was written:dependence.jsonforkalila-fanzareturns 1 shared 7-gram, 0 twelve-grams, longest run 7 tokens (the worst of wives is she who) on 2,551 / 2,348 tokens. Had it returned a 12-gram or a run above 10, every lead-versus-Knatchbull comparison would have been withheld. It did not fire.F3— body loss. Fewer than 90% of 162 bodies parsing after one re-dispatch ⇒ the loss goes in the headline.F4— variance floor. If every seat returns the samerelationfor every item, nothing is claimed.F5—unlocatedceiling. Ifunlocatedis the majoritystatuson more than 8 of 54 items, the whole-chapter presentation is failing as an instrument and the affected items are reported as unmeasured rather than as omissions.
What this run may and may not conclude about the published figure (critic finding 3, accepted):
No published figure is revised by a vote of model seats, and none is confirmed by one. A 0-of-14 seat result does not confirm §7.18's zero; it is reported with its one-sided 95% upper bound of ≈0.19. The single route by which this run may amend §7.18 is not a vote: if seats quote English at a locus and the quoted English exhibits a strict formal match, the quotation is the evidence — checkable by any reader against the printed 1819 page — and the handbook is amended on the quotation. Everything else this run produces is a second, independent, non-lead coding, reported as a coding and not as an adjudication.
5. Stopping rule against critic regress (note bqp, imported verbatim)
One round of critique per version; findings addressed on the page; the next round is bought only if the last one killed a numbered primary, and never to widen the primary. One round was bought; all eleven findings were accepted; no second round is bought, because every remedy narrows.
6. Cost pre-flight
Ceiling for this experiment: $2.50, against the day's $5.00 (UTC 2026-08-21). Worst case built
from max_tokens, not from expected length (note abc).
| stage | calls | worst case |
|---|---|---|
pre-run critic (P1) |
1 | $0.15 — actual $0.109895 |
| stage A census | 162, whole chapter in each prompt | $2.20 |
Stage A per call: input ≈ 3,900 tokens (chapter 3,000–3,400 + frame), max_tokens 600, reasoning
cap 300, both doubled once on re-dispatch. At the dearest seat (P3, $2.00 / $6.00 per M) with
caps doubled: 3,900 × 2e-6 + 1,200 × 6e-6 = $0.0150; at P1 $0.0111; at P2 $0.0074. Sum over
54 items × the three seats, every call re-dispatched: $1.80. Allowing for provider routing above
list (note from config/models.md, S022), $2.20 is the declared worst case.
Stop-loss in the runner: $2.20.
Note (bof) binds the report: per-request billed cost is the ledger and is exact; the key-usage delta is not a cross-check and is not reported as one.
7. What this run cannot answer
- The seats are models, not readers of Arabic and not readers of 1819 English. The relation menu takes the threshold out of their hands; it does not make them right about what stands in the English. Where lead and seats diverge the honest report is a divergence.
- The menu itself is unpiloted. The critic asked for it to be tested on held-out loci before dispatch; that round was not bought. Its inter-rater reliability is therefore measured only by this run's own seat agreement, which is a weaker thing.
- The type label is shown on every item, so a seat can see which items are matched-shape items, and a model may recognise Kalīla wa-Dimna from the prose. Neither is a leak of the hypothesis, and neither is controlled.
- The comparator is OCR of a 1819 scan and is corrupt in places —
G5's passage badly. Both sets now get the same reflow of the same scan, so the differential the critic named is gone; the corruption is not, and no diplomatic transcription was made. - Set B's inventory is the lead's, frozen before any English but by the hand that renders and
codes.
RS-20260816cmeasured an analogous inventory against three readers of the Arabic at 13 of 21 sites. This design buys no such check and claims none. - One translator, one work, one recension, three chapters, 1819 English.
- Tier D is NOT PASSED. No jury is involved; nothing here is a quality claim about any rendering, the lead's least of all.
- The translation limb's
24 MATCHED / 0 IMPOSSIBLEis one hand's, under resources that hand declared in advance and called generous; the strict subset (16) is printed beside it. The seats code the published hand, not the lead's, so neither number is checked here.
8. Contamination declaration
The rendering was written from the Arabic alone; the comparator was extracted by a script printing
only line counts, and was not read until T-kalila-fanza-R39-v1 was committed at cbcc3179. An
earlier candidate chapter was abandoned because part of its comparator had been displayed while
locating chapter boundaries (T-kalila-fanza-R39-v1 §6). Measured: F2.