Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260816h-target-set.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260816h-target-set
statusfrozen
created2026-08-16
updated2026-08-16
sensesstyle-correspondence
internal-judgment-onlytrue
provisionaltrue
linksworkshop/experiments/E-20260816h-target-set/design.md, wiki/findings/results/RS-20260816c-checked-ornament.md, wiki/findings/results/RS-20260816b-invented-figure.md, workshop/translations/kalila-labwa/R37-v1/translation.md, workshop/regimes/R37-declared-rule.md, workshop/translations/kalila-nasik/R34-v1/translation.md, wiki/arms/ARM-invented-ornament.md, framework/v0.2/README.md, wiki/goodness-senses.md, config/models.md

RS-20260816h — the primary fails, and what fails with it is this project's own account of why readers disagreed: obligation follows the audible ending, and the exclusion in the standing inventory rule is the one place the rule is flatly wrong

ARM-invented-ornament step 2, on E-20260816h (v2, frozen at 0273b40 after a pre-run critic returned NEEDS REDESIGN, 11 findings, 3 BLOCKING, and the design was rebuilt rather than dispatched). Translation limb: «باب اللبؤة والإسوار والشغبر» from «كليلة ودمنة», the chapter whole, 599 Arabic words → 1,078 English, frozen at 3ee5001 under the new R37 declared-rule regime. 93 bodies, $0.352836 billed against a declared ceiling of $0.42; verifier 323 checks, 0 failures, 5 mutation tests, 5 caught.

1. The one-sentence answer

Asked, locus by locus, whether the Arabic obliges an English answer, three readers of the Arabic say yes at every one of the 11 places Rule B admits and at 9 of the 12 places only Rule A admits, and no at every one of the 5 places neither admits — so the registered primary, which predicted that the readers would side with Rule B against Rule A by at least +0.40, FAILS at +0.2778 (exact score 0.0924 over 18,564 relabelings), and the reason it fails is that these readers decline almost nothing that is a figure at all.

The interesting number is not the primary. It is what the readers said the places were:

what the seats themselves called it loci owe rate
ROOT — a root in two shapes 4 1.0000
SHAPE — matched pattern, no rhyme 2 1.0000
RHYME — cola ending in the same sound 10 0.9667
REPEAT — a word repeated unchanged 6 0.6667
PARALLEL — parallel in grammar or sense, not in sound 6 0.0000

PARALLEL was not in the design's first draft. The pre-run critic put it there, on the ground that a five-label list built from the two rules forces a sound answer where none belongs. Six loci took it — the five NONE doublets and one figure — and not one of them was called owed by any seat. The option the critic added is the option that carries the page's cleanest result.

2. What was at issue, and what this withdraws

RS-20260816c §5 read three unanimous disagreements between its seats and this project's frozen inventory as three rules, and this project wrote them up as one: rhyme is admitted however it is produced; a root in a changed shape is admitted; a word repeated unchanged is not; a matched pattern that does not rhyme is not. That is Rule B, and the translation limb applied it and Rule A to a fresh chapter.

Two of those three inferences do not survive contact with fresh loci, and this page withdraws them as stated.

RS-20260816c §5 what it was built on here, on fresh loci
1. a word repeated unchanged is not sound work O19, one locus, 0 of 3 7 loci, owe 0.7143 — withdrawn as a rule; the split is by distance, §5
2. rhyme by a grammatical ending or enclitic is admitted O6, O10, O16, 7 of 8 3 loci, owe 1.0000, 9 of 9 bodies — replicates
3. matched shape that does not rhyme is not admitted O15, O20, 0 of 3 each 5 loci, owe 0.7333 — narrower than stated, §6

RS-20260816c reported those three as "rules and not noise". On this evidence, one of the three is a rule. The other two were single loci or a small class read as a class, and the successor design is what caught it. That is the correction, and it is this page's first business.

3. The registered map

prediction value verdict
P0 owe(BOTH) ≥ 0.70 and owe(NONE) ≤ 0.35 1.0000 and 0.0000 HOLDS, at both ceilings
PR owe(B-only) − owe(A-only) ≥ +0.40 +0.2778, exact score 0.0924 / 18,564 FAILS
PR2 the same on 2-member loci only +0.2593, exact score 0.1762 / 715 FAILS
P1 owe(A-REPEAT) ≤ 0.35 0.7143 FAILS
P2 owe(A-SHAPE) ≤ 0.35 0.7333 FAILS
P3 owe(B-RHYME) ≥ 0.60 1.0000 HOLDS
P4 owe(B-ROOT) ≥ 0.60 1.0000 HOLDS
P5 J(readers, RuleB) > J(readers, RuleA) 0.5789 against 0.5652 holds, and means nothing — §4

P0 holding at both ceilings is what makes the failures readable. The design registered a level clause precisely so that a null could not be a jury saying yes to everything or no to everything. This jury separates 1.0000 against 0.0000 on the two strata where the two rules agree, with every one of 24 bodies on the right side. Nothing on this page is a floor or ceiling artefact of an undiscriminating jury; the jury discriminates perfectly on the classes it was given as anchors.

PR failed in the predicted direction and short of its bar, +0.2778 against +0.40, and it did not clear the registered significance level either (0.0924 > 0.05). PR2, the member-count-matched contrast the critic's BLOCKING 1 required, is the same story at +0.2593. The direction is right and the size is not, and by the design's own §5 that is a failure, not a partial success.

4. P5 holds and the page will not lean on it

J(readers, Rule B) = 11/19 = 0.5789; J(readers, Rule A) = 13/23 = 0.5652. The registered criterion is met by 0.0137, which is a difference of set arithmetic and not a finding. The two numbers say something the criterion does not:

Neither rule names the readers' set. Rule B is sound and incomplete; Rule A is nearly complete and slightly over-inclusive. The union of the two contains every one of the 19 places these readers owe, and 4 places they do not — so a translator who answers the union misses nothing and over-answers at about one place in six.

That is the first thing this arm has produced that a translator can act on without a jury standing by, and it is stated as what it is: a result on 28 loci of one chapter, from three uncalibrated models.

5. Where the readers decline a repetition, and it is distance

The A-REPEAT class is the one that splits, and it does not split at random.

locus the Arabic how far apart owe
F26 ذلك العدل : وفي العدل adjacent clauses, 2 words 1.000
F27 رضا الله تعالى ورضا الناس adjacent, 3 words 1.000
F11 فعل غيرك … على فعلك one clause 1.000
F15 لا أرى ولا أسمع … ما أرى وأسمع one sentence, chiastic 1.000
F19 رزقك … رزق غيرك across a long clause 0.500
F22 أكل الثمار … أكل الحشيش across a clause boundary 0.500
F24 أكل اللحم … أكل الثمار across the widest span in the chapter 0.000

The seats see the repetition in every case — the majority relation label is REPEAT at all seven, including all three of the low ones — and decline to call it owed when it is far. P1 on F24: "the unchanged word أكل occurs in both marked phrases", soundwork false, owed false. P1 on F26: "the word al-ʿadl is repeated at the end and beginning of successive clauses", both true.

This is unregistered and it is a hypothesis, not a result. Distance was not a declared variable, the ordering is read off seven loci, and no threshold is estimated. It is recorded here because the next design should register it, and because it explains RS-20260816c §5's O19 without needing a rule about repetition at all: O19's two لسان are the width of a sentence apart.

6. Why the A-SHAPE places are owed — and it is not the shape

Five loci, owe 0.7333. Read the seats' own labels and the class dissolves:

locus inventory says seats' majority label their reason owe
F9 SHAPE — matched 2fs imperfects RHYME ×3 "both clauses close with the matching feminine imperfect ending ‑īna" 1.000
F14 SHAPE — matched فِعْلة nouns RHYME ×2 "both words end in the same rhyming sound with tāʾ marbūṭah" 0.667
F1 SHAPE — two فاعل participles SHAPE ×3 1.000
F4 SHAPE — three form-VIII perfects SHAPE ×2 1.000
F3 SHAPE — three negated jussives PARALLEL ×2 "parallel negated verbs in grammar and sense" 0.000

And the same happens inside A-REPEAT: at F11 («فعل غيرك … على فعلك») two of three seats call it RHYME, not REPEAT — "both end in ‑ka forming sajʿ" — because the two repeated words happen to sit at colon ends.

Three of the twelve A-only loci are owed because the seats hear a rhyme at them, and every one of those rhymes is carried by a grammatical ending or an enclitic — which is exactly what this project's standing inventory rule excludes by name.

F3 is the one A-SHAPE locus the seats decline, and it is a matched-verb locus with no end rhyme — which is RS-20260816c §5's disagreement 3, replicating on one locus of three and not on the class. F4, three matched perfect verbs with no rhyme, is owed at 1.000 by both valid seats. So the decline is not "matched verbs" and not "matched shape". Of the two loci in this chapter that the seats declined outright, both are places where they reported hearing no sound at all.

7. The correction this forces to a standing rule

T-kalila-saih-R30-v1 §3's inventory rule has been imported verbatim into four chapters of this work. Its exclusion (a) reads: rhyme produced solely by an identical enclitic pronoun (‑هُ، ‑ها) is not admitted.

On this chapter, every place excluded by exclusion (a) that was put to the seats is owed at 1.0000, 9 of 9 bodies, with the reason given as saj' at 3 of 3 seats at two of the three loci.

P3 on F5: "all four end in ‑humā, deliberate sajʿ rhyme." P2 on F7: "both clauses end with the same rhyming suffix (‑humā)." P1 on F21: "both end in the possessive plural cadence ‑هم, creating internal prose rhyme."

This is RS-20260816c §5's disagreement 2, on fresh loci, in a different enclitic, unanimous again — the only one of the three that replicated. Together the two runs put it at 3 of 3, 2 of 3 and 2 of 2 valid at kalila-nasik, and 3 of 3 three times over here.

The exclusion is a defensible position in classical rhetoric and it is not what a reader translating from the Arabic behaves as if he believes. The consequence for the framework is at §8, and it is the concrete thing this arm produced.

8. What the framework can now say, and what it must take back

framework/v0.2 §7.15 currently says, on RS-20260816c's authority, that the diagnostic in §7.14 — count the devices the English contains against the figures the source licensed — has a rule-dependent denominator, because two defensible rules pick out different sets.

That is now too strong and is replaced. On this chapter the readers' set is not rule-dependent at all: it is determinate (19 of 23 figure loci, 0 of 5 non-figure loci, unanimous at 24 of the 28), and the two rules differ from it in opposite and nameable ways. What was actually rule-dependent was one clause of this project's own inventory rule, and the readers are against it 9 of 9.

Three things go into v0.2 §7.16 (written this session):

  1. Strike exclusion (a). Rhyme carried by a shared grammatical ending or an identical enclitic pronoun counts as a figure the source is making, and a translator should expect to answer it. It is the commonest sound effect in this prose and the rule was throwing it away.
  2. Answer the union, not one rule. On this chapter the union of a form-based and a sound-based rule contains every place these readers say is owed. The cost of the union is over-answering at 4 of 23 places and about 4.5% in length.
  3. A repetition close by is owed; a repetition far apart may not be — flagged untested, from §5, seven loci, unregistered.

9. Instrument controls, and the four dead bodies

The three third-party controls, all three seats, 9 of 9 on both judgements and 9 of 9 on the label: C1 (al-Ḥarīrī's saj') soundwork true, owed true, RHYME ×3; C2 and C3 (Kalīla, «باب القرد والغيلم») false, false, NONE ×3 each. The gate the critic widened to cover owed passed on every clause. C3, the structurally matched negative the critic asked for, behaves exactly like the study's own NONE stratum.

Four bodies of 93 returned no verdict, and every one is finish_reason: length — P1 ×2 (F4, F19), P2 ×2 (F3, F22). They are dropped from their denominators, counted, and never re-rolled, per design §3a. Method note (bph) fires again: the reasoning cap was set to 150 against a 500-token content cap, which is the shape the note prescribes, and four bodies still spent the content cap on reasoning. Effect on the page: F3, F4, F19 and F22 are measured on two seats rather than three, and two of those four are among the loci that carry §5's and §6's readings. Both readings are re-stated as resting on two-seat loci where they do.

Class-map corroboration, the one free check on the critic's BLOCKING 3: the seats' majority relation label falls in the class's own relation set at 24 of 28 loci (0.857). The four failures are F3, F9, F11, F14 — and all four are §5's and §6's material, so the check did not merely pass, it located the disagreement.

10. Limits

  1. The primary failed. No claim is made that these readers apply Rule B. The design registered what a positive PR would license and it did not get one.
  2. Three fixed models, 28 loci, one chapter, one work, one language pair. Nothing here is a claim about human readers of Arabic (charter §4). The jury is not calibrated (config/models.md, Tier D NOT PASSED); everything is provisional and internal-judgment-only. No anchor is cited and none is claimed. PR's exact score is a randomization score over locus labels, not a population p-value, and §3 says so.
  3. One hand wrote both rules, both inventories, the class map and the marking, knowing what the second pass was for (T-kalila-labwa-R37-v1 D1). The corroboration in §9 is a weak outside reading of that map and not a repair. The repair is a second hand applying the two written rules to the whole chapter, blind, and it is this arm's named successor.
  4. The label list is in the prompt, after the two judgement questions. owed is not fully insulated from the vocabulary. This is why §1's by-label table, striking as it is, is descriptive.
  5. §5's distance reading and §6's rhyme reading are unregistered, read off seven and five loci, and are hypotheses for a successor to register, not results.
  6. The arm's constituting question is still unanswered. Whether a reader who can see the source tells an invented ornament from a compensation, and whether that reader counts the invention against the translation, is untouched by this design.
  7. P5 is met by 0.0137 and the page does not lean on it; §4 reports the containment structure instead, which is what the two Jaccard numbers actually encode.

11. What it hands forward