Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260816h-target-set/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260816h-target-set
statusfrozen
created2026-08-16
updated2026-08-16
sensesstyle-correspondence
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-invented-ornament.md, wiki/findings/results/RS-20260816c-checked-ornament.md, wiki/findings/results/RS-20260816b-invented-figure.md, workshop/translations/kalila-labwa/R37-v1/translation.md, workshop/regimes/R37-declared-rule.md, framework/v0.2/README.md, config/models.md, wiki/goodness-senses.md

E-20260816h — which rule fixes the target set: what a reader of the Arabic says a translator is obliged to answer

ARM-invented-ornament step 2, first half. Study limb of the paired unit whose translation limb is T-kalila-labwa-R37-v1, frozen at 3ee5001 before this design existed.

v2, frozen after a pre-run adversarial critic returned NEEDS REDESIGN, 11 findings, 3 BLOCKING. v1 was never dispatched. Nine findings were accepted and implemented in the six amendments of §12; two were overruled in writing. The critic also read the frozen figure inventories and named three misclassifications in them; two were right and the third forced an amendment that is right, and they are corrected on the translation page at §3.0 rather than here.

1. The question

RS-20260816c failed its own gate and the failure was the finding: the instruction "where the source has a sound figure you cannot reproduce, put a device of your own at that place" presupposes that the places where the source has a sound figure is a determinate set, and on one chapter of «كليلة ودمنة» the project's inventory and three readers of the Arabic picked out sets that agreed on 5 of 10 figures. The disagreement was three nameable rules, not noise.

The translation limb has since written both rules out as rules and applied each to a fresh chapter of the same work: after the erratum this design's critic forced (translation §3.0), Rule A admits 18 loci, Rule B admits 11, the union is 24 and the intersection is 5.

This experiment asks which of the two rules a reader of the Arabic is applying when the question is put in the translator's own terms: is this a place the source obliges my English to answer?

What the unit teaches about translating literature (continue-prompt.md §4.5): it tells a translator, on a language pair where the source's characteristic ornament cannot be reproduced, which places in the source carry an obligation — and if the answer is that two defensible rules disagree at four fifths of the licensed places, that a declared rule is part of the translation, not a preliminary to it. The object measured is the Arabic and what it obliges; the project's inventory is one of the two rules under test, not the instrument doing the testing.

2. Materials

../../translations/kalila-labwa/copy-text-adopted.txt, «باب اللبؤة والإسوار والشغبر», 599 words in 37 sentences, collated whole against the Amīriyya 1937 printing.

Every locus comes from the frozen inventories of T-kalila-labwa-R37-v1 §3, built before this design existed. cells.json is generated by build_materials.py, which asserts that every fully marked sentence reproduces the adopted copy-text character for character when the markers are stripped, that every locus carries at least two marked members, and that every dispatched string carries exactly that locus's members and no brace. The class is never in the dispatched string.

28 study loci in 20 sentences, after the critic pass. Class sizes, with the count of loci carrying exactly two marked members in brackets — that subset is the member-count-matched contrast registered at PR2 (critic BLOCKING 1):

class what it is n (2-member)
BOTH both rules admit it 5 4
A-REPEAT Rule A only — a word repeated unchanged in matched position 7 6
A-SHAPE Rule A only — a matched morphological pattern that does not rhyme 5 3
B-RHYME Rule B only — members ending in the same sound by a grammatical ending or enclitic 3 2
B-ROOT Rule B only — a root in two or more different shapes 3 2
NONE neither rule admits it — a doublet of matched members with no sound relation 5 5

Five loci of the inventory are DROPPED from the study set, each under a rule stated before any dispatch: a locus whose class cannot be assigned unambiguously, or whose span carries a second figure of another class that marking cannot separate, is dropped. They are F20 (its three cola also rhyme on ر at two of three ends, so a verdict would not identify a class), F8, F17, F23 (the three the critic named — translation §3.0) and N6 (جوراً / وظلماً, an accusative pair with the same consonantal skeleton, which the critic's MAJOR 7 correctly refused as a hard negative). cells.json carries them as DROPPED and no body is bought for them.

Three third-party controls, none from this chapter — the critic's MAJOR 9 asked for a negative matched to the study's own structure, and C3 is it:

3. Procedure

One body per (seat, LOCUS) — the critic's BLOCKING 2, taken in full. v1 asked every group of a sentence in one body, and the critic was right that displaying an A-only locus beside a B-only one in the same clause manufactures exactly the contrast PR measures. Each dispatched string is the whole sentence with one locus's members marked ⟪ ⟫ and every other marker stripped, so no seat ever sees two loci together, and group letter, group count and display position carry no information. build_materials.py asserts it. The seats see the Arabic sentence alone: no English, no translation, no mention of this project's inventories.

Seats P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the same three as RS-20260816c, so this is the same jury asked a sharper question, not a new one. temperature 0.4, max_tokens 500, reasoning.max_tokens 150 — the content cap comes down because one locus per body needs one small object, and the body count goes up. Note (bpx): the temperature is above 0 and the seat, not the body, is the unit — one body per seat per locus, three seats, no replicates. Note (bph): the reasoning cap is set well below the content cap.

Each body asks three things, in this order (critic MAJOR 8):

  1. soundwork — is what holds between the marked stretches an audible effect the author has made, something a reader of the Arabic would hear as deliberate work with sound?
  2. owed — a translator is putting this chapter into English and cannot reproduce Arabic sound patterns directly; would he be right to treat this place as a sound effect of the original that his English ought to answer with a device of its own?
  3. relation — one label from RHYME · ROOT · REPEAT · SHAPE · PARALLEL · OTHER · NONE.

v1 put the label list first and named exactly the five categories the classes are built from. The critic's MAJOR 8 was that this makes owed a prompted consequence of recognising a supplied category. v2 asks the two judgements first in their own words, puts the labels last, and adds two options that are not inventory classes — PARALLEL (parallel in grammar or sense, not in sound) and OTHER — which also answers the forced-choice half of MAJOR 7. The residual leak is that a model sees the whole prompt at once; it is stated at §6.3 and it is why relation carries no prediction.

owed is the dependent variable of every registered prediction. soundwork is registered as a secondary with the same directions, and relation is descriptive.

3a. Frozen parse rule

A body must return JSON with a groups object containing every group letter of its sentence, each with a relation in the closed list, a boolean soundwork and a boolean owed. Anything else is invalid and is dropped from its denominators, reported as dropped, and never re-rolled.

3b. Dispatch order, and it is a gate

  1. Stage 1 — the three controls, all three seats, 9 bodies. This is the probe and the gate at once (note (bps)), which is why no study locus is spent on probing; the critic's MAJOR 10 objected that v1 probed on S15, a sentence carrying both primary classes. - Parse/liveness: a seat returning finish_reason: length, an empty body or an unparseable body on any control is dropped for the whole run, its bodies kept and reported, never re-rolled. Below two surviving seats the run aborts. - The control gate, widened by the critic's MAJOR 9 to cover owed and not only soundwork. Over the surviving seats: CTL.POS (C1, al-Ḥarīrī's saj') must be soundwork true and owed true at ≥ 2 of 3; both negatives (C2 an arbitrary pair, C3 a structurally matched noun pair) must be soundwork false and owed false at ≥ 2 of 3. Any failure aborts the run before a single study body is bought, and the failure is the reported result.
  2. Stage 2, the 28 study loci, surviving seats, one body each.

Append-and-resume: every body is written to run.jsonl as it returns.

4. Registered predictions

All rates are owe rates: over the loci of a class and over the surviving seats, the proportion of owed = true among valid bodies. A-only = A-REPEAT ∪ A-SHAPE (12 loci); B-only = B-RHYME ∪ B-ROOT (6 loci); the disputed set is their union, 18 loci.

What a positive PR licenses, written before the number exists (critic BLOCKING 1). It licenses: the loci Rule B admits and Rule A does not are called owed more often, by these seats, than the loci Rule A admits and Rule B does not, by more than a relabelling of the same eighteen loci produces. It does not license: that the seats are "applying Rule B", that the effect would survive independently sampled loci, or that any mechanism has been identified. The critic's proposed rival — that enclitic chains and root echoes are simply more conspicuous than repetitions and matched patterns — is not a rival to the hypothesis but a statement of it, since the two rules differ precisely in whether audibility or matched form is what counts; the page will say so and will not claim more.

soundwork carries the same five class predictions as a registered secondary, reported beside owed at every class.

Registered as descriptive, not predictions, because the design cannot power them:

5. Failure criteria

6. What this design cannot establish

  1. Three fixed models, one chapter, one work, one language pair. Nothing here is a claim about human readers of Arabic (charter §4). The jury is not calibrated (config/models.md, Tier D NOT PASSED); everything is provisional and internal-judgment-only. The three seats are not a sample from a population and no population inference is made: PR's exact score is a randomization score over locus labels, not a p-value about readers.
  2. One hand wrote both rules, both inventories, the class map and the marking, knowing what the second pass was for — T-kalila-labwa-R37-v1 D1. This is the critic's BLOCKING 3 and the design does not repair it. Two things stand against it and neither is enough: the two rules are written out as rules a second hand could apply, which is what makes the successor possible; and the class-map corroboration in §4 gives one free, weak, outside reading of the map. The repair is a second hand applying the two written rules to the whole chapter, blind, and it is this arm's named successor, not this design.
  3. The label list is still in the prompt, after the last of the three questions. A model reads the whole prompt before answering any of it, so owed is not fully insulated from the vocabulary; the ordering and the two non-inventory options reduce the leak and do not remove it. This is why relation carries no prediction and why the corroboration check is called weak.
  4. This measures what a reader says is owed, not what a reader notices in a translation. The arm's constituting question — whether a source-visible reader can tell an invented ornament from a compensation — is the arm's second half and is untouched here.
  5. NONE is not a no-figure control in the strong sense. All five NONE loci are doublets of matched members, chosen because that is where T-kalila-nasik-R34-v1 §4 found an ornamentalist inventing. They are the hard negatives, not the easy ones, so a low NONE rate is a stronger result than it looks and a high one is weaker. C3 is the third-party version of the same shape.
  6. Class sizes of 3 are small and the page will not hide it. B-RHYME and B-ROOT have three loci each, nine seat-locus verdicts apiece before any failure. P3 and P4 are read as descriptions of nine verdicts; the primary is PR, which pools them against twelve.

7. Budget

Worst case built from max_tokens, not from an expected output length (note (abc)). Input ≤ 750 tokens, output capped at 500. Prices config/models.md: P1 $1.00/$6.00, P2 $0.75/$3.75, P3 $2.00/$6.00 per M.

worst case per body
P1 $0.00375
P2 $0.00244
P3 $0.00450
per item, three seats $0.01069
× 31 items (28 study + 3 control) $0.331
pre-run critic, one P1 body, billed $0.033371
worst case, whole experiment $0.365

Stop-loss $0.39, enforced inside run.py. Declared ceiling $0.42. Note (bpq): worst case < stop-loss < ceiling. The UTC day 2026-08-16 stood at $4.388966 of $5.00 before this session, so the ceiling leaves $0.19 of the day unspent.

8. Verification

verify.py recomputes every number on the result page from run.jsonl and cells.json, including both exact enumerations, and re-checks that every dispatched string reproduces the adopted copy-text when its markers are stripped. Mutation tests: perturb the class map, the parse rule, the member-count filter and each enumeration, and confirm every one is caught.

12. The critic pass, and what was done with it

One P1 body, openai/gpt-5.6-terra, $0.033371, critic.json. NEEDS REDESIGN, 11 findings, 3 BLOCKING. P1 also sits on the jury; that is the arrangement E-20260816c used and it is declared, not hidden — a critic that will later answer the items has seen the design, and P1's own rates are reported separately at §4 for that reason.

Six amendments, implementing nine findings.

# from what changed
A1 BLOCKING 2, MAJOR 10 (part) One locus per body. Every other marker is stripped from the dispatched string; group letter, count and position carry no information. Body count 60 → 93; max_tokens 900 → 500 to pay for it.
A2 MAJOR 4 Member-end defined: the last word of the colon, a colon delimited by a coordinating conjunction, a pause mark or the sentence end; in a bare doublet «X وY» the members are X and Y. F5, F7, F21 re-marked to member-final stretches only.
A3 MAJOR 4, 6, 7 F17, F23, N6 dropped, F8 dropped, under the declared drop rule; the translation page carries the erratum at §3.0 and its headline counts move from 20/13/27/6 to 18/11/24/5.
A4 MAJOR 8, MAJOR 7 (part) Prompt re-ordered: soundwork and owed first, in their own words; relation last, with PARALLEL and OTHER added to the label list.
A5 BLOCKING 1 (the real part), MAJOR 10, 11 PR2 registered, the member-count-matched contrast; PR's licence written out before the number exists; P5's "majority" defined for two seats; per-seat rates required; P0 failure made consistent — it invalidates PR, PR2 and P5 too.
A6 MAJOR 9, MAJOR 10 (probe) Third control added (C3, structurally matched negative); the control gate widened to owed; the controls are the probe, so no study locus is spent probing and no primary class is exposed to probe attrition.

Two findings overruled, in writing.