Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260816b-invented-figure/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260816b-invented-figure
statusfrozen
created2026-08-16
updated2026-08-16
sensesstyle-correspondence
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-answering-figure.md, wiki/findings/results/RS-20260816-answering-figure.md, workshop/experiments/E-20260816-answering-figure/design.md, workshop/translations/kalila-saih/R30-v1/translation.md, workshop/translations/kalila-nasik/R34-v1/translation.md, workshop/regimes/R34-ornamentalist.md, workshop/regimes/R35-membered-plain.md, framework/v0.2/README.md, config/models.md, config/budget.md

E-20260816b — is it the sound or the members? and does an ornamentalist with words to spend become audible?

ARM-answering-figure step 2, the arm's last declared step. First frozen at fe1202b6; amended at §4, §7, §8, §9, §11 and §12 on the pre-run critic's findings, before any body was dispatched. Nothing changes after §11 is written.

1. What RS-20260816 left open, in two sentences

One. Thirteen invented sound devices, written at constant length, were heard by three blind seats at 0 of 13, and every one of them was necessarily a two-word pair, because a plain sentence with no spare words offers no matched members to hang a larger device on. §10.4 of that page named the consequence as a limit: a translator inventing ornament in the wild adds words, and this run therefore says nothing about supplied ornament as it actually occurs.

Two. Every affirmative source-inference in that run was justified by the seats in terms of parallel structure, not sound — and three loci moved the inference without being heard as conspicuous English sound at all. That reading is post hoc, is marked untested in framework/v0.2 §7.12 item 3, and has never been given a chance to be wrong.

This design answers both on the same materials, in one sitting, with the same instrument.

2. The subject-rule sentence, and the wire

What this unit teaches about translating literature (wiki/tracks.md, continue-prompt.md §4.5): it tells a translator whether the reader's sense that the original was doing something comes from the sound the translator supplies or from the matched members the source already gave — and whether an ornamentalist free to spend words can make invented ornament audible at all. The subject is craft and what a translation transmits; no instrument, statistic or past figure of this project is its subject. The instrument is reused verbatim precisely so that none of the session goes into it.

The wire between the limbs, in one sentence. The translation limb rendered a fresh Arabic chapter whole under a regime that licenses invented sound ornament and the words to carry it — the licence this design's decoy arm was denied last time — and measured what that hand costs and where it lands; the study limb takes the same licence back to the frozen base text and asks whether ornament written under it becomes audible, and whether the same members with the sound taken out do the work the post-hoc account credits them with.

3. Materials

what provenance
base T-kalila-saih-R30-v1, «باب السائح والصائغ» rendered whole under R30 sound-plain frozen bd87990, S191
R31 the compensation arm: one device at each of the 13 SITE loci frozen d002893, S191
R32 the constant-length invention arm: one device at each of the 13 PLAIN loci frozen d002893, S191
R34 new — the invention arm rebuilt with permission to add words and members materials/invented.json
R35 new — R34's de-sounded twin: the same additions, the same members, no chime materials/invented.json
controls four third-party Q1 controls frozen, E-20260815

The locus set is inherited, not re-chosen. The 13 PLAIN loci are S191's, fixed there by a mechanical length-matching rule over the sentences the Arabic figure inventory does not touch. This design selects nothing. The 13 SITE loci are likewise S191's, fixed by the Arabic inventory frozen on the translation page before any English existed.

3.1 What the fresh chapter contributes, and what it does not

T-kalila-nasik-R34-v1 — «باب الناسك والضيف» rendered whole under R34, 287 Arabic words → 559 English, frozen at cf2dbb8f, contamination: none (0 shared 7-grams, longest run 5 tokens against Knatchbull 1819, measured after the freeze). Its §4 census is the profile this design's R34 arm is written against: 22 ornamented sites in 559 words, class mix 13 MAT / 8 REP / 1 RHY, word cost +34 (6.1%, +1.55 per site), and the hand landed on some structure the source actually has at 14 of 22 sites and invented freely at 8.

It contributes no data to any number in this run. It is cited for one thing: it establishes that the licence in R34 is a licence a hand actually takes, at a measurable rate, on continuous prose, before that licence is exercised at declared loci here.

The arm is more generously funded than the wild hand was, deliberately. R34's cost here is +52 words over 13 loci, +4.00 per locus, against the chapter's +1.55. Every locus here must be manufactured, where the chapter's hand could take the free chances and skip the rest. The direction is the conservative one for M1: if invented ornament is still inaudible when it is given more than the natural allowance, the finding is stronger.

4. The arms, and the repair the pre-run critic forced

Four versions of each of the 13 PLAIN loci, differing at that locus and nowhere else:

arm net word cost over 13 loci matched members added sound
base R30 0 — —
R32 (S191) +13 — S191's rule was substitution only, no new member, not no word may change no yes — necessarily two-word pairs
R34 new +52 yes yes
R35 new +52 yes, the same ones no

R35 is derived from R34, word for word, and is not written independently. That is the whole repair. The critic's BLOCKING 1 on the first draft of this design was that two independently written arms differ in propositions, member count and structure as well as in sound, so their difference cannot be attributed to sound. It was right. In the frozen arms:

R34 − R35 therefore isolates the sound, which is what the primary needs. R35 − base isolates members plus the words that make them, which is a genuine confound and is declared in §7 and §12 rather than denied: you cannot add a matched member without adding the words it is made of.

What cells.py asserts before a single body is dispatched, exiting non-zero on any failure:

Two deliberate limits of check C, both of which the critic named and both of which are stated here rather than discovered later. (i) It screens on digraph vowel nuclei only; a shared single vowel letter is not evidence of a shared vowel sound in English orthography, and screening on it produced false positives (rope/line, cured/borne). Single-vowel assonance is not machine screened. (ii) It screens derivational suffixes only; -ed and -er are inflectional and sit on almost every verb in past-tense narrative, so screening them would reject any two past-tense members. The empirical backstop for both is q1(R35) measured by three blind seats, which is registered in §7 and gated in §8.

The class multiset is preserved: 9 MAT, 2 REP, 2 RHY over the 13 PLAIN loci, identical to R31 and R32. What changes in R34 is that a device may now be a multi-member parallel rather than a two-word pair — the mechanism §6 of RS-20260816 identified as separating its audible four from its inaudible nine.

Rule 1 of R34 binds both arms: no fact, event, referent or evaluation added, nothing dropped. Because the added material is common to both arms, a residual violation of rule 1 damages the base − R35 and base − R34 contrasts and cannot produce a R34 − R35 difference. The common_addition field records what each locus adds, in words, so the residual can be judged.

5. The instrument, reused verbatim

Q1 and Q2 are imported from E-20260816/cells.py by import, not retyped, so a silent divergence is impossible. Q1 was itself validated at E-20260815 and reused verbatim there.

Four third-party Q1 controls, reused verbatim and unanimous on the expected code in both prior sittings: POS (Dickens, Bleak House ch. 1, Y), NEG (Anderson, Winesburg, Ohio, N), ARCH (KJV Genesis 22:3, N), ARCH+ (Morris, The House of the Wolfings, Y).

No new instrument is built in this session and no instrument question is this session's subject.

6. Seats, blinding, dispatch

Three seats, config/models.md: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. Temperature 0. One body per (cell, prompt, seat); judgment is not parallelised. Every seat is blind to the Arabic, to the arm, to the locus class and to every other cell.

Dispatch order is a gate:

  1. the four Q1 instrument controls — a POS or NEG failure aborts the run;
  2. Q1 on all 52 PLAIN cells — M1, F2, F3 scored here;
  3. Q2 on all 52 PLAIN cells — the primary;
  4. Q2 on the 26 SITE cells — the anchor, bought last.

402 bodies. Append-and-resume to run.jsonl (note (bnx)). A body whose finish_reason is not stop, or whose first character is not Y/N, is re-dispatched once; a body failing twice is re-dispatched with an explicit reasoning cap (note (bnl)) and flagged on its stored row.

7. Registered quantities and their bars

Every quantity below is computed from THIS run's run.jsonl and from nothing else. S191 figures appear only in the right-hand column and are historical context, not a contemporaneous control; no bar is defined against them. q1(R31) at the PLAIN loci does not exist in either sitting and is never used.

Notation: q1(arm) and q2(arm) are the number of the 13 PLAIN loci at which the majority of three seats answers Y; q2site(arm) is the same over the 13 SITE loci.

id question registered bar S191, context only
P1 THE PRIMARY — does the sound do anything, given the members? Δsound = (q2(R34) − q2(R35))/13, on two arms that are token-identical but for the ornamented slots \|Δsound\| ≤ 0.15 → the supplied sound carries nothing the members were not already carrying, and §7.12 item 3's structural account survives its first real test; Δsound ≥ +0.25 → the sound carries beyond the structure and item 3 must be rewritten; between → indeterminate, printed and licensing nothing —
P2 do the added matched members raise the source-inference at all? lift35 = (q2(R35) − q2(base))/13 lift35 ≥ +0.30 → members without sound move it; ≤ +0.10 → they do not, and the structural account is refuted at this dose; between → indeterminate. Confounded by construction: the members bring their own words (§12.4) q2(R32) = 1 of 13
M1 the arm's declared step — does the licence to add words and members make invented ornament audible as sound? q1(R34) against q1(base) in the same sitting q1(R34) ≥ 7 → audible; ≤ 3 → not audible; 4–6 → indeterminate. M1 does not isolate the licence — R34 differs from R32 in words, members and diction at once — and no causal attribution to word-licence alone is registered or will be made q1(R32) = 0 of 13
M2 is the sound manipulation audible as sound at all? q1(R34) − q1(R35) q1(R34) > q1(R35) required for P1 to be readable (§8 F3) —
P3 does invention with full licence reach the compensation arm, measured in this sitting? q2(R34) against q2site(R31), both from stage 3 and stage 4 of this run \|q2(R34) − q2site(R31)\|/13 ≤ 0.15 → indistinguishable; (q2site(R31) − q2(R34))/13 ≥ 0.25 → placement still separates them. The conclusion is stated as a Q2 comparison and nothing broader q2site(R31) = 7 of 13
P4 in-sitting replication of S191's cells that are re-dispatched here q1(R32) ≤ 1, q2(R32) ≤ 2, q2site(base) ≤ 3. Reported as a replication, descriptively; no bar here gates anything 0, 1, 2

The registered expectation, written down before the run so it can fail. The structural account predicts P1 near zero, P2 high, M1 low. A device-quality account — the explanation RS-20260816 could not exclude — predicts the opposite on both of the first two: what the seats respond to is the sound a hand supplies, so P1 large and P2 near zero. The two accounts are opposed on the primary, which is why this design is worth its money.

8. Manipulation checks and failure criteria

No gate below is moved after the run.

id check consequence of failure
F1 Q1 controls POS and NEG unanimous on the expected code abort before stage 2
F1b ARCH and ARCH+ unanimous every rate withheld; the census still printed, as at S191
F2 q1(base) over the 13 PLAIN loci ≤ 2 every Q1-dependent claim withheld; M1 unreadable
F3 q1(R34) > q1(R35) — the sound manipulation landed as heard P1 is withheld. If the two arms are equally audible as sound, the slots did not carry the manipulation and their Q2 difference cannot be read as a sound effect
F4 Q2 seat unanimity ≥ 0.60 over its cells every Q2 rate withheld
F5 invalid bodies ≤ 5% after the re-dispatch ladder the run is reported unusable
F6 missing cells — if the stop-loss halts the run, or any cell is unfilled for an operational reason, every rate computed over an incomplete arm is withheld and the arm is reported as partial. Operational failure is never reported as a substantive result —

q1(R35) is a measurement, not a gate on its own. Q1 asks about conspicuous sound-patterning, and a seat may reasonably code a repeated syntactic frame as patterning. If q1(R35) comes back high, that is a finding about what the instrument hears, reported as such; what it must not do is exceed q1(R34), which is F3.

9. Analysis

Fixed here, computed by verify.py, which recomputes every reported figure from run.jsonl independently of analyse.py.

  1. Cell code = majority of the three seats' first characters.
  2. Rates as in §7; body-level rates (39 ratings per arm) printed alongside every locus-level rate.
  3. Exact sign-flip enumeration over the 13 matched pairs, 2^13 = 8,192 arrangements, enumerated exhaustively, for P1, P2 and M1. This number is reported as a descriptive randomisation score and is not called a frequentist P-value for a causal claim. The critic's objection is accepted in full: the loci are authored, not randomly assigned; the seats are three fixed models at temperature 0, not a sample of readers; arm labels are not exchangeable under a null in the way a permutation test assumes. The score says how extreme the observed split is among the 8,192 relabellings of these thirteen pairs, and nothing about a population.
  4. Seat agreement: mean pairwise agreement and unanimity fraction, per prompt.
  5. Pre-declared subsets, each printed with its n and each underpowered: by device class (MAT 9 / REP 2 / RHY 2) and stratum (frame 2 / narrative 11). No claim rests on a subset.
  6. Free-text reasons stored and read; any reading of them labelled POST HOC.

10. Cost

Pre-flight built from the max_tokens cap the request permits — note (abc).

stage seat bodies cap worst case
pre-run critic, two calls, both spent P1 2 9,000 / 5,000 $0.096722 actual
the census P1 134 400 $0.369
P2 134 1,400 $0.738
P3 134 800 $0.737
re-dispatch ladder — ≤ 10% — +$0.184 (note (boe))
declared ceiling 404 $2.40

Expected, from S191's realised $0.00262 per body on the same three seats: ≈ $1.13. run.py carries a stop-loss at $2.40. Budget position: UTC day 2026-08-16 opens at $0.00 of $5.00. Per-request billed cost is the ledger and is exact; no key-usage delta is reported as a cross-check — note (bof).

11. Pre-run critic — verdict NEEDS REDESIGN, 22 findings, 13 BLOCKING, all 22 accepted

One call, P1 openai/gpt-5.6-terra, $0.065025, finish_reason: stop, shown the frozen design and every locus of both new arms beside the frozen base wording and beside S191's R32. Raw output in critic.json. A second, focused call was then made on the rebuilt arms alone, recorded in §11.2.

11.1 The findings that changed the design

# severity finding disposition
1 BLOCKING the R34–R35 contrast does not isolate sound: the arms differ in semantic material, member count, rhetorical form and, at several loci, word count. "Same class multiset" is not equivalence of the manipulation ACCEPTED, and it is the reason this design was rebuilt. R35 is now derived from R34 word for word: same additions, same members, same word count at every locus, ≤ 4 differing token positions. §4
2 BLOCKING rule 1 is violated repeatedly — the additions add actions, states, attributes, temporal and legal facts, not restatements; and the added_material field records additions without establishing equivalence ACCEPTED. Every added proposition is now common to both arms and recorded in common_addition, so a residual violation cannot generate a primary effect. §4 last paragraph. The critic's demand for an independent proposition audit is not met and is carried as a limit, §12.7
3–14 BLOCKING ×11, MAJOR ×1 locus-by-locus audit: P2 different facts and a register-marked simian in one arm only · P3 befit/is not owed are three different claims · P4 two falls versus a lying-down · P5 holds them yet is a new fact · P6 two different intensifications, and weighed/gave assonate · P7 fair fruit and fine is elliptical and changes good · P8 an incomplete the chief assertion · P9 quick to climb is a new attribute and assonates with light · P10 both arms add different experiential material · P11 no hearing adds a legal fact · P12 the best of food adds an evaluation, wait a while alliterates internally · P13 a man genders the referent, confidence/close chime · P14 laced is a new action, down/wound chime ACCEPTED at every locus. All thirteen were rewritten. The rebuilt arms are in materials/invented.json; the specific chimes the critic heard (wait a while, weighed/gave, light/climb, confidence/close, down/wound) are gone with the wordings that carried them
15 MAJOR the "no sound" claim is false as a claim about English reading even when cells.py passes: the screen missed internal alliteration, medial vowel echo, contextual echo and repeated function-word rhythm ACCEPTED. Check C now screens every content word of every member, adds member-internal alliteration, member-final chime and a digraph-nucleus test, and the two limits that remain are printed in §4 rather than left implicit. The critic's proposed blind phonological panel is not run; the registered q1(R35) measurement and F3 are the empirical substitute, and §12.6 says so
16 BLOCKING F3 could fail because the structural manipulation is salient, not because R35 contains sound ACCEPTED, and the gate is inverted. q1(R35) is now a measurement; the gate is q1(R34) > q1(R35). Because the two arms carry identical parallelism, any Q1 difference between them is sound and not structure. §8
17 BLOCKING M1's absolute bar cannot establish that the word-licence makes ornament audible: R34 varies words, members and diction at once ACCEPTED. M1's row now states in the design that it does not isolate the licence and that no causal attribution to word-licence alone is registered. §7
18 BLOCKING P1 as first drafted did not test "structure alone", because R35 carried its own semantic and rhetorical changes ACCEPTED, and the primary was replaced. The primary is now R34 − R35, the one contrast in which everything but sound is held identical. The old primary survives as P2 with its confound declared
19 MAJOR the statistical language overstates what a sign-flip calculation supports ACCEPTED in full. §9.3 now reports it as a descriptive randomisation score and refuses the frequentist reading in the design's own words
20 MAJOR P2/P3 contain unsupported equivalence and interpretation leaps; no outcome assigned to Δ below −0.15; P3's conclusion overreaches ACCEPTED. Bars are now symmetric with a named indeterminate band, and P3's conclusion is narrowed to the Q2 comparison
21 MAJOR historical and current measurements are ambiguously mixed ACCEPTED. §7's header now states that every registered quantity comes from this run's run.jsonl; S191 figures are labelled context; P3 uses only stage-4 cells; q1(R31) is named as nonexistent and unused
22 MAJOR gates can fire for operational rather than substantive reasons, with no missing-data rule ACCEPTED in part. F6 is new and separates operational failure from instrument failure. A same-sitting Q2 positive control is NOT added — building one is instrument work and the subject rule forbids it as a unit — and §12.6 carries the consequence
— MINOR ×2 §4/§12 inconsistency on the negative rule's scope; R32 mislabelled "0 words added by construction (+13)" ACCEPTED. §4's table now states R32's rule correctly (substitution only, no new member) with its realised +13

11.2 The second, focused call — verdict RUN WITH AMENDMENTS

P1, $0.031697, finish_reason: stop, shown only the rebuilt arms beside the base wording and asked three questions: does any locus differ in meaning, member count or structure rather than sound; is any R35 locus audibly patterned to an English ear; is any R34 device absent or too faint. Raw output in critic2.json. Fourteen findings. Six accepted with rewrites, six refused on a mechanical ground, one accepted as an irreducible limit, one accepted as a description error.

Accepted, and the locus rewritten — the arms now differ at 15 token positions in total across thirteen loci, down from 22:

locus finding what changed
P2 sank it low is not a normal way to say dug it deep → sank it deep; now one word differs between the arms, and sink a shaft deep is ordinary English
P5 murderer / thief adds criminal and intentional implication that slayer / seizer does not both arms rebuilt on the killer and the keeper / the killer and the **holder; one word differs
P7 fresh and fair ≠ ripe and sound, and R35's sound … set alliterates → fresh and fair / fresh and **choice; one word differs, and the sound/set chime is gone
P8 less grateful ≠ base: ingratitude and baseness are different charges → no man less grateful / no man less **beholden; one word differs
P9 R35's cord … caught and line … light alliterate — chimes R34 does not have, i.e. the manipulation running backwards the three-way was cut to a two-way: let the rope down / let the cable down, one word differing, second and third mentions left as the base's pronoun
P2 R34's device is not a three-member chain; it is two cola with a repeated verb description corrected in r34_device; a factual error about the design's own materials

Refused, on a ground that cells.py check F makes mechanical rather than rhetorical. The critic flagged R35 as audibly patterned at P3 (where it does not repeated), P4 (down / ground), P11 (no … no), P13 (after … after), and as carrying minor echoes at P2 (pit / it) and P8 (mankind / man). Every one of these strings is present, unchanged, in R34 at the same locus — check F asserts the two arms are token-identical outside ≤ 4 slots per locus, and none of these words is in a slot. A chime common to both arms cannot produce a difference between them. It can inflate q1(R35) in absolute terms, which is exactly why §8 makes q1(R35) a measurement and the gate a comparison.

Accepted as an irreducible limit, not repaired. The critic's general prescription — use genuinely equivalent wording, not merely same-length near-synonyms — cannot be met: in English, genuinely equivalent wording is the same wording, which has the same sound. Every de-sounded twin must substitute a near-synonym, and near-synonyms differ in sense. This is a property of the manipulation, not of this execution, and it is carried to §12.11 as a limit rather than argued away. P13's test / proof is refused on the facts: after trial and after proof is a fixed English doublet in which proof means the act of proving, the sense the critic assigns to test.

A screen change the second pass forced. The comparative degree words (less, least, more, most, than, nothing) are now treated as part of the shared frame rather than as content, since they sit in both arms at every locus where they occur; without that, check C rejected P8 for a less / less echo that R34 has identically.

12. What this design cannot do, written before it is run

  1. One work, one chapter, one hand, one language pair.
  2. n = 13 per arm. A locus-level rate moves in steps of 1/13.
  3. One hand wrote R34 and R35 knowing the hypothesis. The token-level twinning, the word count equality, the machine-checked negative rule and F3 are the defences. They are much stronger than S191's and they are not a randomised assignment.
  4. P2 is confounded by construction. Matched members cannot be added without adding the words they are made of, so R35 − base measures members plus their words, never members alone. No design that adds members escapes this; it is named, not solved.
  5. R35 cannot remove the shared syntactic frame, because the frame is the parallelism.
  6. The R35 "no sound" property is screened, not certified. Single-vowel assonance is not machine-screened, no blind phonological panel was run, and Q2 still has no same-sitting positive control of its own — the last because building one would be instrument work.
  7. No independent proposition audit was run on the common additions. They are recorded per locus so a later reader can judge them, and they are common to both arms, which bounds the damage to P2 and M1 and leaves the primary untouched.
  8. Seats are models, not readers. Q1 measures conspicuity only. Nothing here says any device in any arm is good English, and nothing here recommends writing like R34.
  9. The jury is not calibrated (config/models.md, Tier D NOT PASSED). Every claim is provisional and internal-judgment-only.
  10. S191's published Victorian panel is not re-bought and is cited as context, never as this run's evidence.
  11. Exact synonymy does not exist, so a de-sounded twin is never a pure sound manipulation. At every locus R35 substitutes a near-synonym, and near-synonyms differ in sense. The residual semantic difference is one word at eleven of thirteen loci and two at the other two, it was written to no Q2 prediction, and its direction is not predictable — but it is real, the pre-run critic named it at five loci, and no design that removes a sound can avoid it.