Repository path: workshop/experiments/E-20260808g-two-pasts/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260808g-two-pasts |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| senses | style-correspondence, voice, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-legend.md, workshop/translations/szent-peter-esernyoje/R05-v1/translation.md, workshop/translations/szent-peter-esernyoje/register.md, workshop/translations/szent-peter-esernyoje/collation.md, config/models.md, config/budget.md |
E-20260808g-two-pasts — does the tag alternation a translator must throw away carry anything?
Frozen before dispatch. The translation limb it hangs on (T-szent-peter-esernyoje-R05-v1,
span C) was frozen and committed at aba3b1a, before this design was written.
1. The question, and where it came from
Translating «Szent Péter esernyője» III forced decision D6 and register rule V6. Mikszáth's
speech tags run in two competing past tenses — the archaic narrative past (mondá, kérdé) and the
ordinary past (mondta, kérdezte) — 26 times in this one chapter, 16 archaic against 10
ordinary. English has one past tense and one bleached tag verb. The translator rendered every tag
plainly and wrote V6 as a declared loss, with the honest note that whether anything is actually
lost by this is not a question the translator can answer.
This experiment asks it. Not does the English keep the form — it obviously cannot — but does the source form carry information that independent translating hands recover and put into their English? If they do not, the loss is nominal and V6 is right. If they do, V6 is wrong and must change by erratum.
The project has five retired attempts at the English loses the marking (framework/v0.2 S1).
Every one of them assumed the source marking meant something and went looking for the loss. This
one asks whether the marking means anything first. It is not a sixth attempt at prediction 1: the
device here is not a relation between persons but a narrator's tag morphology.
2. Contamination, measured after the freeze and before any locus was chosen
tools/dependence_check.py, cells in run/:
| pair | 7-grams | 12-grams | 15-grams | longest run | verdict |
|---|---|---|---|---|---|
| span C (3,479 w) × Worswick 1900 ch. III | 31 | 1 | 0 | 12 | usable as an independent comparator |
span C ¶48–66 (689 w) × the lead's own S077 R10c u1 |
5 | 0 | 0 | 9 | clean |
span C ¶48–66 × the lead's own S077 R10c u2 |
8 | 0 | 0 | 9 | clean |
span C ¶48–66 × the lead's own S077 R10p u1 |
68 | 29 | 18 | 27 | DEPENDENT |
span C ¶48–66 × the lead's own S077 R10p u2 |
51 | 18 | 9 | 19 | DEPENDENT |
Note (bhb), fourth firing, and at a density the note has not recorded before: 29 shared twelve-grams in 689 words against a different session's rendering of the same passage, written 61 sessions earlier and not read. The longest run is 27 contiguous tokens of ordinary narrative prose — "priest's relations had carried off every stick and left only a dog, the late incumbent's favourite, a dog like any other to look at, in shape and…" — not a formula. The lead is therefore not an independent hand on this material and writes none of the arms below.
Against Worswick the rendering is clean: the single 12-token run is "he was tired and went across to the schoolmaster's, for he was", entirely function words and a noun.
3. What has already been measured, before this design was frozen — declared, not confirmatory
Two things were computed from the source while choosing the unit, and neither is registered as a prediction; both are reported as description in the result page's census section, with this disclosure attached.
(a) The alternation is almost entirely lexical. Over the whole novel, on all speech-verb
lemmas: mond 47 archaic / 30 ordinary, kérd 28 / 9, kiált 6 / 14, sóhajt 2 / 9, kezd 2 /
67, and szól 0 / 56 and felel 0 / 41. Two verbs carry 82% of the archaic tokens; the two
commonest tag verbs after them never take the form at all.
(b) Within mond and kérd, the form is bound to the attribution slot. A single mechanical
feature — is the verb preceded, within 40 characters, by the attribution dash — was fitted on
Parts I–II (n = 46) and checked on Parts III–V (n = 66): train 0.848, held out 0.894. The
archaic past is the form of the tag; the ordinary past is what the verb takes in narration and in
reported speech inside dialogue.
This was run before the design was frozen and is therefore not a pre-registered test. The rule is a single binary mechanical feature and the split is by the book's own part boundaries, which bounds the fishing; it does not remove it. P3 below is the pre-registered version, on a work this project has never opened.
4. Materials
Source. «Szent Péter esernyője», Révai 1910 (PG #68911), whole novel — 2,145 paragraphs.
Site pool, selected mechanically and frozen in run/pool.json before any arm was written:
every token of mond/kérd in the archaic past that (i) sits in the attribution slot by the §3(b)
feature, (ii) falls in a paragraph of 120–900 characters, and (iii) is the only occurrence of that
token in its paragraph, so substitution is unambiguous. The pool is 17 sites, spread I:4,
II:6, III:3, IV:2, V:2. No site was chosen or rejected by reading it.
Allocation, by pool index, fixed before dispatch:
- MAIN — 13 sites: pool indices 0–11 and 14. Arm A = as printed. Arm O = byte-identical
except the tag verb is replaced by its ordinary-past counterpart (
mondá→mondta,mondák→mondták,kérdé→kérdezte). - POSITIVE — 4 sites: pool indices 12, 13, 15, 16, all
mondá. Arm a = as printed. Arm b =mondá→kiáltá, an archaic tag verb of the same shape whose meaning differs. A translation that cannot show this difference cannot show any tag difference.
Hands — four, none of them the lead, none Anthropic (config/models.md): H1 deepseek/deepseek-v4-pro,
H2 mistralai/mistral-medium-3-5, H3 openai/gpt-5.6-terra, H4 x-ai/grok-4.5.
Assignment, and why there is no separate REPEAT block. For each site, two hands receive one arm and two the other; the pairing alternates with the site's parity, so no hand is systematically on one arm. Every site therefore yields, at no extra cost:
- 2 SAME-ARM comparisons (H1×H2, H3×H4) — two independent hands given byte-identical source. This is the REPEAT baseline.
- 4 CROSS-ARM comparisons — the effect.
Over 13 MAIN sites: 26 same-arm and 52 cross-arm comparisons.
No hand sees both arms of any site, and each hand receives its 17 passages in one call, in pool order, with no indication that anything has been manipulated.
5. Procedure
snapshot open—GET /api/v1/key.- Pre-run critic, one adversarial pass over this frozen design by a non-panel model
(
nvidia/nemotron-3-ultra-550b-a55b), which has returned BLOCKING findings that changed the result in each of the last ten sessions. Findings are accepted or overruled in writing before dispatch. - Translate: four calls, one per hand, each 17 numbered Hungarian passages → 17 numbered English renderings. Instruction is to translate, nothing else; the tag is never mentioned.
- Extract the attribution clause mechanically from each English rendering, by a script written
before the bodies are read (
extract.py), against a fixed verb list. Items where extraction is ambiguous go to a declaredunextractablebucket and are reported, never guessed. - Classify each extracted tag on three mechanical fields:
VERB (lemma), ORDER (
inverted= verb before subject /canonical/no-subject), ARCHAISM (any of quoth, saith, made answer, thou, thee, thy, -eth in the tag). - Recognition probe, non-gating, one cheap call: is the work identifiable from one passage?
snapshot close; reconcile per-responseusage.costagainst the key delta.
Judgment is not parallelised. No model judges its own output; the classification in step 5 is a script, not a seat.
6. Predictions, registered
P1 — the primary. On the MAIN sites, the cross-arm disagreement rate on the composite (VERB, ORDER, ARCHAISM) exceeds the same-arm disagreement rate by no more than 0.10.
P1 holds → independent hands do not put the source's tense choice into their English, and V6's declared loss is not a measurable loss. P1 fails → the alternation is carried, and V6 must change by erratum.
Registered as a null-favouring primary and reported as such: holding P1 is failing to reject. The same-arm rate is the yardstick precisely because it prices what "no difference" looks like when two hands render identical Hungarian.
P2 — the gate. On the POSITIVE sites, cross-arm minus same-arm disagreement on VERB is ≥ 0.50.
P3 — the pre-registered replication, mechanical, $0. The §3(b) tag-slot rule, unchanged,
predicts the form at ≥ 0.80 of mond/kérd past-tense tokens in Mikszáth, «A vén gazember»
(Révai, 1910; Project Gutenberg #69738) — a work this project has never opened and which is not
downloaded at the time of this freeze.
P4 — descriptive, no prediction. Worswick 1900's chapter III renderings of the 26 tag sites, censused against the source form. Reported as a census; nothing is predicted.
7. Failure criteria, frozen
- F1 — P2 < 0.50 → the instrument cannot see a tag difference at all and P1 is withheld.
- F2 — more than 3 of the 13 MAIN sites are
unextractablein 2 or more hands → P1 is withheld for want of data. - F3 — any hand's rendering shares ≥ 8 contiguous tokens with Worswick 1900 at a site → that item is flagged; at ≥ 3 such items the hand is dropped and P1 recomputed without it.
- F4 — if a hand returns fewer than 17 numbered items, or renumbers, its call is re-dispatched once; a second failure drops the hand and P1 is recomputed on three.
- No criterion is weakened after it fires. If a gate fires the number is still reported.
8. Pre-flight cost estimate
Built from max_tokens, not from an expected length (note (abc)).
| stage | calls | cap | worst case |
|---|---|---|---|
| critic | 1 | 16,000 | $0.11 |
| translation | 4 | 8,000 | $0.42 |
| recognition | 1 | 1,000 | $0.02 |
| re-dispatch reserve | 2 | 8,000 | $0.21 |
| declared worst case | $0.76 |
Today's UTC ledger stands at $3.580596 of $5.00, headroom $1.419404. The worst case fits with $0.66 to spare. The mechanical limbs (P3, P4, §3) cost $0.
9. What this cannot show
- Four 2026 models are not four human translators, and their shared training distribution is the standing limitation on every generation stage this project runs. A null here is these hands did not carry it, not it is uncarriable.
- The novel is a Hungarian classic and the hands may know it. The recognition probe measures this and does not gate on it; F3 catches the case where a hand reproduces the published English.
- P1 is a null-favouring primary and a small study cannot prove a null. The same-arm baseline makes the comparison meaningful, but the honest statement of a holding P1 is no effect distinguishable from the noise floor of two hands rendering identical text, with the floor's own size reported beside it.
- §3(a) and §3(b) were computed before this freeze and are description, not test.
10. Amendments, after the pre-run critic and before dispatch
Critic: google/gemini-3.6-flash (panel P2, not a hand in this run), $0.035775.
Verdict NEEDS-REDESIGN, 3 BLOCKING and 1 ADVISORY. All four findings accepted, none overruled.
The registered critic was nvidia/nemotron-3-ultra-550b-a55b and it returned finish_reason:
length with zero content at a 16,000 cap, $0.0613476. Note (bhf): zero content means change
the seat, not raise the cap. The dead body is in runs/discarded/ and is ledgered as waste.
A1 — the hand-pair confound (critic BLOCKING 1). Accepted; §4's assignment is replaced.
The design as frozen made the same-arm pairs always H1×H2 and H3×H4, and the cross-arm pairs always across those two groups. Four different model families have different baseline habits in tag placement and tag-verb choice, so every baseline difference between the groups would have been counted as an arm effect, and every baseline similarity inside a group would have depressed the noise floor the primary is measured against. The primary would have been biased toward failing — that is, toward finding an effect that is a model-family difference.
Replaced. The 2–2 split now rotates with the site index over the three possible pairings —
{H1H2 | H3H4}, {H1H3 | H2H4}, {H1H4 | H2H3} — so every one of the six model pairs appears in
both configurations across the site set. The primary is computed per model pair and averaged
over pairs, so a pair's baseline agreement enters both sides of the comparison. Arm orientation
still alternates with site parity.
A2 — the positive control was the wrong control (critic BLOCKING 2). Accepted; a second one is added and it becomes the gate.
mondá → kiáltá is a lexical change. Models carry lexical changes trivially, so passing it
would have licensed nothing about whether these hands can carry a morphological one — which is
exactly what P1's null needs.
Added: POS-TENSE. Same lemma, same slot, tense morphology the only difference:
mondá → mondja, kérdé → kérdezi (archaic past → present). English must move said → says.
This is now the gate. POS-LEX is retained as a secondary check with no gate attached.
Re-allocation, by pool index, fixed before dispatch: MAIN = 0–9 and 14 (11 sites) · POS-TENSE = 10, 11, 13, 15 (4 sites) · POS-LEX = 12, 16 (2 sites).
P2 is restated: on the POS-TENSE sites, cross-arm minus same-arm disagreement on VERB or TENSE of the English tag is ≥ 0.50. F1 now fires on POS-TENSE, not on POS-LEX.
A3 — the composite was going to collapse (critic BLOCKING 3). Accepted.
ARCHAISM (quoth, saith, thou…) will be zero in essentially every cell for both arms, which would
have shrunk the composite to (VERB, ORDER) while presenting it as three fields and letting P1 hold
partly by dilution.
The registered composite is now (VERB, ORDER). ARCHAISM is reported as a bare count with no
prediction attached, and a fourth descriptive field PERIOD-MARK — any of quoth, saith, methinks,
'tis, nay, whereat, thereupon, betimes, made answer anywhere in the rendering — is reported the
same way. Neither enters the primary.
A4 — the extraction list was an open degree of freedom (critic ADVISORY 4). Accepted.
The verb list and the extraction rule are frozen here, before dispatch, and extract.py is written
and committed before any body is read.
Tag verbs (lemma, matched in past or present): say · ask · answer · reply · cry · shout · exclaim · remark · observe · add · retort · declare · murmur · mutter · whisper · continue · respond · return · tell · speak · call · insist · note · comment · interject · venture · resume · protest · admit · agree · explain · repeat · sigh · laugh · grumble · boast · snap · jest · joke · banter · utter · pronounce · state · mention · inquire · query · demand · question · wonder · put in · chime in · go on · break in · make answer.
Extraction rule. Split the English rendering on dash and quotation boundaries; the attribution
clause is the first resulting segment, at or after the first speech segment, that contains one of
the listed verbs. Zero candidates, or a candidate containing two listed verbs with no way to order
them, → unextractable, reported and never guessed.
ORDER. inverted if the verb is immediately followed by a determiner, possessive, title,
capitalised name or personal pronoun; canonical if such a token immediately precedes the verb;
no-subject otherwise.
What the amendments cost
Nothing in cash — the call count is unchanged. MAIN loses two sites, 13 → 11, which is the price of a gate that tests the right thing. Cross-arm comparisons on MAIN: 44. Same-arm: 22.