Repository path: workshop/experiments/E-20260807c-two-hands/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260807c-two-hands |
| status | frozen |
| created | 2026-08-07 |
| updated | 2026-08-07 |
| senses | accuracy, voice, style-correspondence, cultural-mediation |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-two-hands.md, workshop/translations/szent-peter-esernyoje/R04-v1/translation.md, workshop/translations/szent-peter-esernyoje/R06-v1/translation.md, workshop/translations/szent-peter-esernyoje/span-1-source.txt, wiki/findings/essays/ES-20260806-craft-report-koyhaa-kansaa.md, wiki/findings/results/RS-20260806d-first-span-again.md, config/models.md, config/budget.md |
E-20260807c — two hands on one span
ARM-two-hands step 1. S127, 2026-08-07. Frozen before any API call. Tier D is NOT PASSED;
every judgment reported from this run is provisional and internal-judgment-only.
1. Question
(Re-scoped by amendment A5 after the pre-run critic's BLOCKING 5; the original wording is directly below and is struck rather than deleted.)
~~On one span of literature, how much of an English rendering is fixed by the source and how much is the translator?~~ No scalar answer to that is claimed by this run. What is measured is two things, and the question is the pair of them:
- Surface agreement between independent hands over the whole span, against the only measured within-translator blind self-agreement this project owns; and
- Sense-level agreement at sixteen marked places — do two hands do the same thing where the Hungarian marks something, and where they differ, does either drop what it marks?
Two independent competent hands on the same 650 Hungarian words: where do they agree, where do they diverge, and at the divergences does one of them lose something the source marks?
ES-20260806-craft-report-koyhaa-kansaa §7 named this as the successor question and named its
obstacle: everything RS-20260806d measured is one translator's variance with itself.
2. Materials, and the order they were made in
Source. Mikszáth Kálmán, «Szent Péter esernyője» (1895), Part I ch. 1 «Viszik a kis Veronkát»,
whole: 16 paragraphs, 650 Hungarian words. Copy-text
workshop/translations/szent-peter-esernyoje/span-1-source.txt, from Magyar Elektronikus Könyvtár
MEK-00954, decoded from ISO-8859-2. Public domain (Mikszáth d. 1910).
Hands.
| id | hand | words | provenance |
|---|---|---|---|
LEAD |
lead, R04 close translation |
900 | T-szent-peter-esernyoje-R04-v1, frozen e420b81 |
LEADdraft |
lead, R06 single pass |
908 | T-szent-peter-esernyoje-R06-v1, frozen 4617b00 |
WORS |
B. W. Worswick, 1900 | 793 | Project Gutenberg #31945, public domain, read whole |
M1a,M1b |
panel seat P1, two independent unbriefed calls | — | this run |
M2a,M2b |
panel seat P5, two independent unbriefed calls | — | this run |
Order of work, and every step is a git commit. Chapter read in Hungarian → R06 drafted and
frozen (4617b00), its log carrying the registered sixteen-locus prediction → R04 revised and
frozen (e420b81) → only then Worswick's chapter 1 opened and extracted → contamination measured
→ this design written. No English rendering of this chapter was read by the lead before e420b81.
Contamination, measured before this design was written (materials/contamination.json,
tools/dependence_check.py):
| cell | shared 7-grams | 12-grams | longest run | verdict |
|---|---|---|---|---|
subject — LEAD × WORS |
2 | 0 | 8 | clean |
| ref — an independent published pair (K&M 1915 × Garnett 1920, «Пари») | 160 | 27 | 24 | DEPENDENT? |
ref — the lead against itself, draft and revision (R06 × R04) |
583 | 460 | 95 | DEPENDENT? |
The subject cell is below the independent-published-pair reference on every column by a wide margin. Limit: the reference pair's texts are longer (2,840/2,707 tokens against 900/793), which gives them more opportunities to match; the subject's zero twelve-grams is not a length artifact, since a single one would have shown.
Exposure declared rather than left to the number. The lead knew the novel's English title and
that a Worswick translation exists. It had read R. Nisbet Bain's name in the Gutenberg front matter
during extraction. No line of Worswick's chapter 1 was read before e420b81.
3. The measurement, and why it is defined this way
Tokenisation is tools/ngram_overlap.tokenise plus standalone-digit deletion — the same function
the contamination gate used. Proper names are kept.
For an ordered pair of renderings (a, b), difflib.SequenceMatcher(None, ta, tb, autojunk=False):
agree(a,b) = 2·M / (|ta| + |tb|), where M is the total size of the matching blocks.sites(a,b)= the number of non-equalopcodes — maximal contiguous stretches where the two renderings do not coincide — andsites/100w=sites÷ 6.50.
Whole-span alignment, not paragraph alignment, and this is forced by the materials. Worswick renders
the source's 16 paragraphs as 11: he merges. A paragraph-indexed diff would therefore have to be
aligned by the lead, and lead alignment of another hand's prose is the circularity
RS-20260806d's critic finding 1 convicted. SequenceMatcher over the whole token sequence needs no
alignment decision from anyone.
Temperature 1.0 on every translation call, and the design fails without it. The within-hand
quantity this run needs is the same hand rendering the same source twice, blind to the first
attempt. At temperature 0 a second call is a re-run, not a second translation, and agree would go
to ≈1.0 and make P1 vacuous. Temperature 1.0 is the only setting under which within-hand variance
is a real quantity. Seats judge at temperature 0.0.
4. Registered predictions
Written before any call. Scored exactly as written.
~~P1 — a second hand reaches places one hand does not. min{ agree(M1a,M1b), agree(M2a,M2b) }
> max{ agree over all between-hand pairs }.~~ STRUCK by amendment A1 — the critic's
BLOCKING 1 is right that same-model-twice against cross-model is trivially ordered and measures a
sampler, not a translator.
P1′ (A1) — a second hand diverges from a first by more than a translator diverges from
himself. The within-hand term is the lead's own re-rendering of span 1 of «Köyhää kansaa», blind
to its first rendering (T-koyhaa-kansaa-R05-v1 ¶2–75 × T-koyhaa-kansaa-R05-v2,
RS-20260806d), recomputed under this run's metric. Registered:
sites/100wfor every between-hand pair on this span ≥ 1.5 ×sites/100wfor the lead's blind self-pair, and the same direction holds onagree.
It can fail in both directions and the lead has computed neither number. Cross-work confound
declared in advance: the reference is Finnish→English, a different work, and a hand that had the
closed register in between. It is the only measured blind within-translator number in existence here,
which is the fact ES-20260806 §7 was complaining about. Model self-agreement is reported as an
observation and scores nothing.
~~P2 — the human hand is the outlier.~~ STRUCK by amendment A2 (critic BLOCKING 2 and
ADVISORY 15): Worswick is the only human, the only 1900, the only compressing and the only free hand,
so the prediction cannot fail for a nameable reason. The cells are reported under §Observations with
their length ratios, and nothing is inferred from them.
P3 (A3) — the frozen log predicted where hands would diverge, tested where the lead is not a
party. The R06 log fixed sixteen loci from the Hungarian alone, eight D (predicted divergence)
and eight C (predicted convergence), frozen at 4617b00 before Worswick was opened.
- Primary — the lead-free pair. Seats judge
WORSagainstM1a, two hands that had no part in choosing the loci. Registered: the D "differently" rate exceeds the C rate by ≥ 0.25, the direction holds at ≥ 2 of 3 seats, and the exact permutation test over the sixteen frozen labels (A7) is reported with it. - Secondary, reported beside it and not instead of it — the same judgment for
LEADagainstWORS.
This is the answer to critic BLOCKING 3: if a split the lead wrote predicts divergence between two hands the lead did not write, the circularity is not doing the work.
Q2 — the craft payload, reported, not predicted. At each locus a seat calls differently, it
also says whether both English versions carry what the Hungarian marks by different means, or
one of them loses it (and which). Reported as a distribution and as a direction. No prediction is
registered because the lead has no blind expectation here it could honestly claim to have had.
5. Failure criteria
- ~~
FC1~~ — struck withP1byA1; model self-agreement now scores nothing, so its stability gates nothing. The two values are still reported. FC2— any translation call withfinish_reason ≠ stop, or returning fewer than 600 or more than 1,300 words, is re-dispatched once; if it fails again the hand is dropped,P1andP2are evaluated on what remains, and the drop is reported.FC3(REPEAT control, relaxed byA8) — a seat answering differently on ≥3 of its 4 REPEAT items is confabulating difference and is excluded fromP3; 2 of 4 is reported as a caution and excludes nothing. Four items is a weak control and the result says so.FC4(WRONG control — demoted byA4to a floor gate and nothing more) — if the three seats pooled detect fewer than 9 of 12 WRONG items as differently and loses, the seats are not shown to be reading the Hungarian at all, andP3andQ2are withheld. If it passes, that licenses exactly one sentence — the seats can read Hungarian and can name a loss — and not any claim that their detection was calibrated for difficulty.Q2is reported descriptively either way, because the lead set this control's difficulty and a control whose difficulty the lead sets cannot license a harder task (critic BLOCKING 4).FC5— if fewer than 3 seats return usable block-1 or block-3 JSON, the correspondingP3arm is withheld.FC6(A9) — a seat judgment whose quoted stretch does not occur in the version it is attributed to is excluded from every rate, and the per-seat verification rate is reported.
6. The seat limb
Three seats, none of which is a translating hand in this run (P1 and P5 translate; P2, P3, P4 judge). No model judges its own output. The lead judges nothing.
Block 1 — the real comparison. One call per seat. The seat receives the whole Hungarian
chapter (paragraph-numbered), version A whole and version B whole, and the sixteen loci
quoted in Hungarian. It aligns them itself; nobody aligns them for it. Per locus it returns
same | different, and if different, one of BOTH-CARRY | A-LOSES | B-LOSES, plus a short
reason.
Blinding. Whether A is LEAD or WORS is decided per seat by a hash-fixed coin recorded in
runs/*.map.json; locus order is shuffled per seat by the same hash. Neither version is labelled,
dated or attributed. The seats are told both versions are English translations of the chapter and
nothing else.
Block 3 — the lead-free pair (added by A3). One call per seat, format identical to block 1,
same sixteen loci: A/B are WORS and M1a. This block carries P3's primary.
Every seat answer also carries quote_a and quote_b (A9) — the stretch the seat aligned in
each version, verbatim. Each quote is checked mechanically against the version it is attributed to;
FC6 excludes judgments whose quotes do not occur there.
Block 2 — the controls, same format. One call per seat. A = LEAD whole; B = LEAD whole with
four loci deliberately damaged. Eight loci are asked: the four damaged (WRONG: L02, L07,
L10, L14 — two D and two C) and four untouched (REPEAT: L01, L05, L11, L16 — two D
and two C). The damage and its intent are frozen in materials/loci.json.
Block 2 is ceiling-favourable and is declared so. Because A and B are byte-identical everywhere
except four places, the damaged loci are easier to spot than a genuine between-hand difference.
Passing FC4 is therefore necessary and not sufficient: it shows the seats can read the
Hungarian and can name a loss, not that they would catch a subtle one.
7. Pre-flight budget estimate
Worst case built from max_tokens, not from expected length — note (abc).
| stage | calls | max_tokens |
worst-case price basis | worst case |
|---|---|---|---|---|
| 0 — pre-run critic | 1 (+1 reserve) | 16,000 | S126's measured $0.0598 for one such call | $0.15 |
| 1 — translation hands | 4 (+2 re-dispatch) | 3,000 | P1 $6.00/M out; P5 at the measured worst provider ~$3.30/M (config/models.md pricing caution) |
$0.11 |
| 2 — seat blocks | 9 (+3 re-dispatch) | 6,000 | P4 $15.00/M out, P2 $7.50/M, P3 $6.00/M; inputs ~5,000 tok | $0.84 |
Declared ceiling for this session: $1.10, raised from $0.80 by amendment A3, which adds a third
seat block (three calls) to take P3's primary off a hand-pair the lead is a party to. The reason for
the raise is written here rather than absorbed: the critic's BLOCKING 3 was worth three calls. Today's UTC ledger stands at $0.641307535 of $5.00
with $4.358692465 of headroom, so the declared ceiling is 18% of what remains. The max_tokens
of 6,000 on the seat blocks is set in advance from the measured hidden-reasoning appetite of P2
(note (bhq)), not rediscovered by a truncation.
Lead translation is $0 and is not ledgered (charter §3, A4). The two renderings the whole run rests on cost nothing, and one of them is the only English rendering of this chapter made since 1900 that anybody wrote from the Hungarian.
8. What this design cannot do
- Model sampling variance is not a human translator's variance with herself.
P1's within-hand term is what a model does at temperature 1.0 on two independent calls. The nearest human-side number the project owns isRS-20260806d's lead re-rendering blind to its own first attempt — on a different work, a different language, and a hand that had learned the register in between. Both are reported; neither is the other. - One span, one work, one language pair. Nothing here generalises past Hungarian→English on 650 words of one chapter.
- Period is uncontrolled. Worswick's English is 1900; three of the four other hands are 2026.
This is the same confound
RS-20260807bdeclared and it is not solved here. - Worswick is one human hand, not a sample of human hands. He is also, on the evidence of his
793 words against the source's 650, a free translator;
P2cannot separate that from his humanity or his century. - Tier D is NOT PASSED. Every seat judgment is
provisional; nothing here ranks two translations and no sentence in the result may be read as saying one is better. - The sixteen loci were chosen by the lead (critic BLOCKING 3).
A3tests them on a pair the lead is not in, which answers whether the split predicts anything beyond the lead's own process — it does not make the loci a random or mechanically-derived sample of the chapter, and no claim is made that they are representative of it. Q2's difficulty is uncalibrated (critic BLOCKING 4). The lead authored the damage in the only control that establishes loss-detection, so nothing about the seats' sensitivity to subtle loss is shown.
9. Amendments after the pre-run critic
Ten accepted, two overruled with written reasons; dispositions and the critic's own words in
critic.md. A1 strikes the registered primary and replaces it; A2 strikes P2; A3 adds block
3 and moves P3's primary onto a lead-free pair; A4 demotes FC4; A5 re-scopes §1; A6 adds a
length-matched contamination reference; A7 scores the D/C contrast by exact permutation over the
sixteen frozen labels, with exchangeability declared as its assumption; A8 relaxes FC3; A9
requires and verifies seat quotes and adds FC6; A10 records the two overrulings (ADVISORY 13 on
slug placement, against CLAUDE.md rule 6; ADVISORY 16 on what provisional qualifies).