Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260808a-set-forms/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260808a-set-forms
statusfrozen
created2026-08-08
updated2026-08-08
sensesaccuracy, voice, style-correspondence, cultural-mediation
internal-judgment-onlytrue
provisionaltrue
linksworkshop/experiments/E-20260808a-set-forms/materials/loci.md, workshop/translations/szent-peter-esernyoje/R04-span2-v1/translation.md, workshop/translations/szent-peter-esernyoje/R06-span2-v1/translation.md, wiki/arms/ARM-two-hands.md, wiki/findings/results/RS-20260807c-two-hands.md, config/models.md, config/budget.md

E-20260808a — do the fixed forms of a language mark where two hands part?

ARM-two-hands step 2, translation limb. Frozen before any API call. The renderings and the locus classification were frozen first, at b583be6 and ea3531e.

Question

RS-20260807c-two-hands §4 found four loci the lead had labelled plain and at which two independent hands nevertheless parted, and three of the four contained a set form the lead had not noticed when it labelled them. That is n = 3, post hoc, on one chapter.

The transportable claim, and the one under test: set-formhood predicts where two independent hands part, even when nothing about the stretch looks difficult. A translator scanning a source for hard places looks for the class this design calls MARK — puns, culture-bound objects, code-switches, things English visibly has no slot for. If SET — archaic morphology, frozen idiom, proverb, fossilised politeness formula, Latin tag, reduplicative, frequentative — behaves like MARK rather than like PLAIN, then the advance list every translator makes is systematically short, and short in a nameable way.

What this teaches about translating literature (subject rule): it names a class of place where translations of the same source will differ, and it is a class defined on the source language's inventory of fixed forms rather than on the translator's sense of difficulty.

Materials

Contamination, measured before the design was written (materials/contamination.json): LEAD × WORS over ¶1–22 gives 12 shared 7-grams, 0 twelve-grams, longest run 11 tokens, against 160 / 27 / 24 for an independent published pair. Clean. (Span 1 of the same work and the same comparator: 2 / 0 / 8.)

Pairs

block pair loci why
A — PRIMARY M1a × M2a 49 (whole chapter) lead-free, same era on both sides, full coverage. Set by critic amendment A1; the frozen design had a 1900 × 2026 pair here and the critic's BLOCKING 5 showed the era gap would produce the predicted effect on its own
B — replication M1b × M2b 49 an independent second draw from the same two models (amendment A2)
C — controls built on M2b 12 4 REPEAT, 4 WRONG, 4 unmanipulated filler, in a block of their own so the primary block's texts are byte-identical to the filed renderings (amendment A5)
D — descriptive WORS × M1a 34 (those inside ¶1–22) the human hand, carrying no registered prediction, because it is the pair the era confound bites

The frozen design's blocks A and C were WORS × M1 and LEAD × WORS. Both are gone from the inferential structure; the human comparator survives as block D and as the D1 census.

Procedure

  1. Generate M1 and M2, unbriefed, one rendering each, temperature 1.0.
  2. Three seats — P1, P2, P3 — none of which produced any rendering in this run. Each seat receives each block as one call: both full English versions, labelled only Version A and Version B, side order randomised per seat per block, authorship never stated.
  3. For each locus the seat is given the Hungarian stretch and its paragraph number, and returns: verdict ∈ {same, different}; quote_a and quote_b — the stretch it aligned in each version; and where different, direction ∈ {both_carry, a_drops, b_drops}.
  4. Judgment is never parallelised across loci within a seat; each block is one call.

Predictions, registered

Amendment A3 binds all of these: the inferential unit is the LOCUS, and a locus's verdict is the majority of the three seats. Seats reduce measurement noise; they are not replicates.

Failure criteria, registered

Limits, declared before running

  1. The lead read WORS before these thresholds were written. The loci and their classes were frozen blind (ea3531e, before the comparator was opened), and P1's threshold is transported verbatim from the previous experiment rather than chosen here — but the choice to restrict the primary to ¶1–22 was made after reading Worswick, because his chapter stops there. This is the run's principal bias exposure and it is stated here, not in a footnote.
  2. P4 carries no threshold for the same reason.
  3. One chapter, one work, one language pair, one human hand. Worswick is not a sample of human translators, and M1/M2 are not a sample of translators at all.
  4. Class sizes are unequal and small — 8 MARK, 19 SET, 10 PLAIN, of which 8 MARK, 11 SET, 7 PLAIN fall inside Worswick's coverage.
  5. Tier D is NOT PASSED. Nothing here judges quality, ranks a translation, or scores a sense. Different is not worse.
  6. The direction field inherits RS-20260807c §6's confound whole: all seats are 2026 models, two of the four hands are 2026 models, and a shared prior favouring machine prose is a live alternative to any reading of who is said to drop things. Reported descriptively only.

Descriptive census, no prediction attached

D1 — what Worswick has no text for. Of the 37 loci, which have no counterpart in his chapter at all. This is computed from his text and needs no seat. It is reported because his chapter's omission of ¶23–¶30 was discovered in the course of this run and is a fact about the translation, not about the instrument.

Cost, pre-flight

Worst case built from max_tokens, per note (abc).

stage calls cap worst case
pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) 1 16,000 $0.06 → $0.025595 actual
M1a M1b (P5), M2a M2b (P4) renderings 4 4,000 $0.13
seat blocks A/B/D, 3 seats × 3 blocks 9 12,000 $0.78
seat block C (controls), 3 seats 3 4,000 $0.08
re-dispatch headroom — — $0.08
declared ceiling, raised by amendment A8 before dispatch $1.10

Headroom at declaration: $5.00 (UTC day 2026-08-08 opens with no rows); $4.974405 at the moment A8 raised the ceiling, the critic call having already been paid.