Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260821-matched-shape/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260821-matched-shape
statusfrozen
created2026-08-21
updated2026-08-21
versionv2-post-critic
sensesstyle-correspondence
provisionaltrue
internal-judgment-onlytrue
linkswiki/arms/ARM-published-figure.md, wiki/base/anchors/A-knatchbull-kalila/README.md, wiki/findings/results/RS-20260816j-published-figure.md, workshop/translations/kalila-fanza/R39-v1/translation.md, workshop/regimes/R39-matched-shape.md, workshop/experiments/E-20260821-matched-shape/critic-response.md, framework/v0.2/README.md, config/models.md

E-20260821-matched-shape — a second coding hand on the published translator's matched-shape figures, and a third chapter

One-sentence design. Give three blind model seats the whole of a Knatchbull 1819 chapter, one Arabic figure-locus at a time, and have each seat find where that material is rendered, quote the English, and name which formal relation — from a menu fixed before dispatch — holds between the English members; on the fourteen matched-shape loci of an already-censused chapter and on forty loci of a third chapter rendered and frozen this session.

Revised v2-post-critic after a pre-run adversarial pass returned NEEDS REDESIGN with 3 BLOCKING findings, all accepted: critic-response.md. The window is abolished, the seat's threshold is replaced by a reported feature, and the falsification-by-vote bar is struck.

1. Question and predictions

Question. framework/v0.2 §7.18 tells a practitioner that the one published English hand from the Arabic answers 0 of 18 matched-shape figures — muwāzana, the class Knatchbull's own 1818 preface calls the sententious brevity of the Arabic and says he could not express. All seventy-two of the anchor's three-way calls were the lead's, unchecked, made by the party holding the hypothesis (RS-20260816j §7.2). What a second coding hand finds at those loci, and what it finds on a third chapter, is unknown.

The subject rule (continue-prompt.md §4.5), in one sentence: it teaches what a published translator does at the place where the source's prose makes a shape English has no morphology for — and, paired with the translation limb, whether that shape can be carried at all. The materials include this project's own inventories; the question is about a translator and about English.

Predictions, registered before any census body was dispatched

# prediction
P1 Seat-majority strict match on the fourteen Set A matched-shape loci is 0 to 3 of 14
P2 Among loci the seats find rendered, seat-majority match on Set B's matched-shape loci is below that on Set B's non-matched-shape loci. Descriptive; 16 non-SHAPE items give little precision and no test is run
P3 On the six non-PLAIN cells of the lead's Set B coding (F9 F15 F26 ABSENT; F59 F91 F94 ANSWERED — six of these seven, F94 being the sixth-and-seventh borderline), the seat majority agrees on at least 4. Overall agreement is reported against the trivial all-PLAIN baseline of 34/40 = 85%, which any constant answerer would reach
P4 F59 and F91 — the two matched shapes the lead's own rendering flagged as free, where English's ordinary wording already matches — are the two Set B SHAPE loci most likely to come back matched from the seats as well
P5 Seats are less unanimous on matched-shape loci than on the other three classes

No prediction appears in any prompt. No seat is told the work, the translator, the century, the existence of a hypothesis, any count, or which class the design cares about.

2. Materials

materials/items.json, built by materials/build_items.py; 54 items.

2.1 The relation menu — what replaces the seat's threshold

The seat does not decide whether a figure is answered. It reports what it sees, from a menu fixed before dispatch (build_items.py RELATIONS), and the analyser computes the readings:

class strict loose only not a match
SHAPE identical-suffix · identical-inflection same-class-and-length frame-only · none
RHYME rhyming-words · identical-ending same-pronoun-or-particle none
REPEAT same-content-word · same-phrase same-frame-varied-words none
ROOT same-stem-two-forms etymological-pair none

Every reported rate is printed twice, strict and loose. The lead's own three close calls fall out of this menu mechanically: F59 (‑tion/‑tion) is strict, F38 (‑ed/‑en) is not identical inflection, F34 (‑ness/‑ment) is not identical suffix.

3. Procedure

  1. Stage A. 54 items × 3 seats = 162 bodies. Seats P1 (openai/gpt-5.6-terra), P2 (google/gemini-3.6-flash), P3 (x-ai/grok-4.5) — the panel less P4 (note bps) and P5 (note bne), both out. Deterministic shuffle. temperature: 0. One call per item per seat; judgement is never batched.
  2. Each body returns {"status": "rendered"|"omitted"|"unlocated", "english": "...", "members": "...", "relation": "<from the menu>", "reason": "..."}. unlocated is a distinct state from omitted (critic finding 1) and is counted separately, never folded into either.
  3. Re-dispatch each dead body once at the doubled cap (note bgk); still dead ⇒ dead: true and that cell is void.
  4. The seat verdict for a locus is the majority of three on status, and separately on strict-match / loose-match / no-match. No majority ⇒ counted as such, never broken by the lead.
  5. verify.py recomputes every reported number from run.jsonl by a path sharing no code with the runner, including the confusion matrix, the trivial baseline, and the exact binomial bounds.

4. Gates and failure criteria

What this run may and may not conclude about the published figure (critic finding 3, accepted):

No published figure is revised by a vote of model seats, and none is confirmed by one. A 0-of-14 seat result does not confirm §7.18's zero; it is reported with its one-sided 95% upper bound of ≈0.19. The single route by which this run may amend §7.18 is not a vote: if seats quote English at a locus and the quoted English exhibits a strict formal match, the quotation is the evidence — checkable by any reader against the printed 1819 page — and the handbook is amended on the quotation. Everything else this run produces is a second, independent, non-lead coding, reported as a coding and not as an adjudication.

5. Stopping rule against critic regress (note bqp, imported verbatim)

One round of critique per version; findings addressed on the page; the next round is bought only if the last one killed a numbered primary, and never to widen the primary. One round was bought; all eleven findings were accepted; no second round is bought, because every remedy narrows.

6. Cost pre-flight

Ceiling for this experiment: $2.50, against the day's $5.00 (UTC 2026-08-21). Worst case built from max_tokens, not from expected length (note abc).

stage calls worst case
pre-run critic (P1) 1 $0.15 — actual $0.109895
stage A census 162, whole chapter in each prompt $2.20

Stage A per call: input ≈ 3,900 tokens (chapter 3,000–3,400 + frame), max_tokens 600, reasoning cap 300, both doubled once on re-dispatch. At the dearest seat (P3, $2.00 / $6.00 per M) with caps doubled: 3,900 × 2e-6 + 1,200 × 6e-6 = $0.0150; at P1 $0.0111; at P2 $0.0074. Sum over 54 items × the three seats, every call re-dispatched: $1.80. Allowing for provider routing above list (note from config/models.md, S022), $2.20 is the declared worst case. Stop-loss in the runner: $2.20.

Note (bof) binds the report: per-request billed cost is the ledger and is exact; the key-usage delta is not a cross-check and is not reported as one.

7. What this run cannot answer

  1. The seats are models, not readers of Arabic and not readers of 1819 English. The relation menu takes the threshold out of their hands; it does not make them right about what stands in the English. Where lead and seats diverge the honest report is a divergence.
  2. The menu itself is unpiloted. The critic asked for it to be tested on held-out loci before dispatch; that round was not bought. Its inter-rater reliability is therefore measured only by this run's own seat agreement, which is a weaker thing.
  3. The type label is shown on every item, so a seat can see which items are matched-shape items, and a model may recognise Kalīla wa-Dimna from the prose. Neither is a leak of the hypothesis, and neither is controlled.
  4. The comparator is OCR of a 1819 scan and is corrupt in places — G5's passage badly. Both sets now get the same reflow of the same scan, so the differential the critic named is gone; the corruption is not, and no diplomatic transcription was made.
  5. Set B's inventory is the lead's, frozen before any English but by the hand that renders and codes. RS-20260816c measured an analogous inventory against three readers of the Arabic at 13 of 21 sites. This design buys no such check and claims none.
  6. One translator, one work, one recension, three chapters, 1819 English.
  7. Tier D is NOT PASSED. No jury is involved; nothing here is a quality claim about any rendering, the lead's least of all.
  8. The translation limb's 24 MATCHED / 0 IMPOSSIBLE is one hand's, under resources that hand declared in advance and called generous; the strict subset (16) is printed beside it. The seats code the published hand, not the lead's, so neither number is checked here.

8. Contamination declaration

The rendering was written from the Arabic alone; the comparator was extracted by a script printing only line counts, and was not read until T-kalila-fanza-R39-v1 was committed at cbcc3179. An earlier candidate chapter was abandoned because part of its comparator had been displayed while locating chapter boundaries (T-kalila-fanza-R39-v1 §6). Measured: F2.