Repository path: workshop/experiments/E-20260808a-set-forms/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260808a-set-forms |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| senses | accuracy, voice, style-correspondence, cultural-mediation |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260808a-set-forms/materials/loci.md, workshop/translations/szent-peter-esernyoje/R04-span2-v1/translation.md, workshop/translations/szent-peter-esernyoje/R06-span2-v1/translation.md, wiki/arms/ARM-two-hands.md, wiki/findings/results/RS-20260807c-two-hands.md, config/models.md, config/budget.md |
E-20260808a — do the fixed forms of a language mark where two hands part?
ARM-two-hands step 2, translation limb. Frozen before any API call. The renderings and the
locus classification were frozen first, at b583be6 and ea3531e.
Question
RS-20260807c-two-hands §4 found four loci the lead had labelled plain and at which two
independent hands nevertheless parted, and three of the four contained a set form the lead had not
noticed when it labelled them. That is n = 3, post hoc, on one chapter.
The transportable claim, and the one under test: set-formhood predicts where two independent hands part, even when nothing about the stretch looks difficult. A translator scanning a source for hard places looks for the class this design calls MARK — puns, culture-bound objects, code-switches, things English visibly has no slot for. If SET — archaic morphology, frozen idiom, proverb, fossilised politeness formula, Latin tag, reduplicative, frequentative — behaves like MARK rather than like PLAIN, then the advance list every translator makes is systematically short, and short in a nameable way.
What this teaches about translating literature (subject rule): it names a class of place where translations of the same source will differ, and it is a class defined on the source language's inventory of fixed forms rather than on the translator's sense of difficulty.
Materials
- Source. Mikszáth Kálmán, «Szent Péter esernyője» (1895), Part I ch. 2 «Glogova régen», whole,
30 paragraphs, 726 Hungarian words (
span-2-source.txt, MEK-00954). LEAD—T-szent-peter-esernyoje-R04-span2-v1, whole chapter, 995 words. Frozenea3531e.WORS— B. W. Worswick, 1900, ch. II (Project Gutenberg #31945, public domain), 757 words. Worswick's chapter renders Hungarian ¶1–¶22 only; ¶23–¶30 have no counterpart in his text.M1,M2— two unbriefed panel renderings of the whole chapter, generated for this run,P5andP4(config/models.md). Each is told only that it is translating Hungarian literary prose of 1895 into English; no regime, no locus list, no sight of any other rendering.- Loci. 37, classified by the rule frozen in
materials/loci.md: 8 MARK, 19 SET, 10 PLAIN, PLAIN drawn with no selection at all (first finite clause of every paragraph containing neither other class).
Contamination, measured before the design was written (materials/contamination.json):
LEAD × WORS over ¶1–22 gives 12 shared 7-grams, 0 twelve-grams, longest run 11 tokens,
against 160 / 27 / 24 for an independent published pair. Clean. (Span 1 of the same work and the
same comparator: 2 / 0 / 8.)
Pairs
| block | pair | loci | why |
|---|---|---|---|
| A — PRIMARY | M1a × M2a |
49 (whole chapter) | lead-free, same era on both sides, full coverage. Set by critic amendment A1; the frozen design had a 1900 × 2026 pair here and the critic's BLOCKING 5 showed the era gap would produce the predicted effect on its own |
| B — replication | M1b × M2b |
49 | an independent second draw from the same two models (amendment A2) |
| C — controls | built on M2b |
12 | 4 REPEAT, 4 WRONG, 4 unmanipulated filler, in a block of their own so the primary block's texts are byte-identical to the filed renderings (amendment A5) |
| D — descriptive | WORS × M1a |
34 (those inside ¶1–22) | the human hand, carrying no registered prediction, because it is the pair the era confound bites |
The frozen design's blocks A and C were WORS × M1 and LEAD × WORS. Both are gone from the
inferential structure; the human comparator survives as block D and as the D1 census.
Procedure
- Generate
M1andM2, unbriefed, one rendering each, temperature 1.0. - Three seats —
P1,P2,P3— none of which produced any rendering in this run. Each seat receives each block as one call: both full English versions, labelled only Version A and Version B, side order randomised per seat per block, authorship never stated. - For each locus the seat is given the Hungarian stretch and its paragraph number, and returns:
verdict∈ {same,different};quote_aandquote_b— the stretch it aligned in each version; and wheredifferent,direction∈ {both_carry,a_drops,b_drops}. - Judgment is never parallelised across loci within a seat; each block is one call.
Predictions, registered
Amendment A3 binds all of these: the inferential unit is the LOCUS, and a locus's verdict is the
majority of the three seats. Seats reduce measurement noise; they are not replicates.
P1— PRIMARY. On block A, over the 49 loci, on seat-majority verdicts: SET − PLAIN ≥ 0.25 in the proportion called different. The threshold is transported verbatim fromE-20260807cP3and is not chosen here (see §Limits). Reported separately againstPLAIN-AandPLAIN-B; if the two PLAIN subsets disagree, the primary is read againstPLAIN-Balone, which is the within-paragraph control.P2— replication. The same criterion on block B, the independent second draw.P3— permutation. Exact one-sided permutation over the 49 frozen class labels of block A on the statistic SET − PLAIN, enumerated or Monte-Carlo'd to a stated precision. P < 0.05.P4— secondary, descriptive. SET against MARK. No threshold is registered, because any number chosen here would have been chosen after the comparator was read. Reported as a difference with its seat-level spread and nothing inferred from its size.P5— control on the instrument. MARK − PLAIN ≥ 0.25 at ≥ 2 of 3 seats on block A. IfP5fails,P1is uninterpretable rather than confirmed or refuted: a design that cannot see hands part at a pun cannot be trusted when it says they part at an idiom.
Failure criteria, registered
F1REPEAT. Four loci in block C show the same rendering on both sides. A seat fires if it calls ≥ 2 of its 4 different; if ≥ 2 of 3 seats fire, the primary is withheld. (AmendmentA6: the frozen bar was any of the four, which the critic computed would withhold the primary ~9% of the time against a perfect instrument.)F2WRONG. Four loci in block C carry a deliberately damaged side (intents frozen inloci.md). Detection means different and the damaged side named as the one that drops something. Below 9 of 12 across seats, the primary is withheld (amendmentA6).F3quote verification. Every judgment must quote the stretch it aligned in each version; quotes are checked mechanically against that version. Failing judgments are dropped from every rate. A seat below 0.60 verification on a block has that block withheld.- A withheld primary is not re-derived from a weakened bar after the fact.
Limits, declared before running
- The lead read
WORSbefore these thresholds were written. The loci and their classes were frozen blind (ea3531e, before the comparator was opened), andP1's threshold is transported verbatim from the previous experiment rather than chosen here — but the choice to restrict the primary to ¶1–22 was made after reading Worswick, because his chapter stops there. This is the run's principal bias exposure and it is stated here, not in a footnote. P4carries no threshold for the same reason.- One chapter, one work, one language pair, one human hand. Worswick is not a sample of human
translators, and
M1/M2are not a sample of translators at all. - Class sizes are unequal and small — 8 MARK, 19 SET, 10 PLAIN, of which 8 MARK, 11 SET, 7 PLAIN fall inside Worswick's coverage.
- Tier D is NOT PASSED. Nothing here judges quality, ranks a translation, or scores a sense. Different is not worse.
- The
directionfield inheritsRS-20260807c§6's confound whole: all seats are 2026 models, two of the four hands are 2026 models, and a shared prior favouring machine prose is a live alternative to any reading of who is said to drop things. Reported descriptively only.
Descriptive census, no prediction attached
D1 — what Worswick has no text for. Of the 37 loci, which have no counterpart in his chapter
at all. This is computed from his text and needs no seat. It is reported because his chapter's
omission of ¶23–¶30 was discovered in the course of this run and is a fact about the translation,
not about the instrument.
Cost, pre-flight
Worst case built from max_tokens, per note (abc).
| stage | calls | cap | worst case |
|---|---|---|---|
pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
1 | 16,000 | $0.06 → $0.025595 actual |
M1a M1b (P5), M2a M2b (P4) renderings |
4 | 4,000 | $0.13 |
| seat blocks A/B/D, 3 seats × 3 blocks | 9 | 12,000 | $0.78 |
| seat block C (controls), 3 seats | 3 | 4,000 | $0.08 |
| re-dispatch headroom | — | — | $0.08 |
declared ceiling, raised by amendment A8 before dispatch |
$1.10 |
Headroom at declaration: $5.00 (UTC day 2026-08-08 opens with no rows); $4.974405 at the
moment A8 raised the ceiling, the critic call having already been paid.