Repository path: workshop/experiments/E-20260808h-trajectory/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260808h-trajectory |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| senses | style-correspondence, voice, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-trajectory.md, framework/v0.2/README.md, workshop/translations/duel/R06-v1/translation.md, config/models.md, config/budget.md, wiki/findings/results/RS-20260808b-discordance-fails.md |
E-20260808h-trajectory — when a source changes how one person addresses another partway through a passage, does English move?
Frozen before dispatch. The translation limb it hangs on (T-duel-R06-v1, «Дуэль» XV
¶660–700) was written from the Russian alone and committed at 0bad3cc, before this design
existed and before Garnett was opened.
1. The question, and where it came from
framework/v0.2 §6 opened §8 Q-d on 2026-08-08:
Where a source marks a relation by a change in form across a text — a switch of pronoun, an onset of honorifics, a drop of a title — does English carry the change? And what does a published translator do with it?
It was opened because RS-20260808b retired prediction 1 with the observation that every
instrument this project has built is per utterance, and the thing that mattered in Chekhov's
«Толстый и тонкий» lived in the text's trajectory: three readers shown each utterance alone
called both halves concordant, and they were right, because each half is.
Q-d also wrote down the design problem in advance: it is a question about a text's arc, so it may not be answered by putting isolated utterances to readers; whatever instrument it uses has to show the reader the trajectory, which is the opposite of the blinding every design since S101 has needed.
This design's answer to that. Blinding has been protecting three different things, and only one of them has to go:
| what was blinded | keep? |
|---|---|
| which hand wrote the English | keep — raters never learn |
| what the hypothesis is, that a manipulation exists at all, that these are translations | keep — the rater prompt names no device |
| context — the rest of the passage | drop — the rater sees the whole span, which is the point |
So the trajectory is shown and the hypothesis is not. Nothing about showing a reader a whole passage requires telling them what to look for in it.
And the order is inverted, per note (bku). The project has retired one prediction after five attempts that each began by assuming a source form meant something and going to look for its loss. This one manipulates the source first and asks whether independent hands' English moves at all.
2. What the translation limb contributed, and the wire
The lead translated «Дуэль» XV ¶660–700 — the quarrel in which Laevsky goes from ты to вы with his oldest friend at ¶681 and Samoylenko follows him at ¶686 in the middle of a sentence: «Что ты… что вы сказали?» English has one second-person pronoun and cannot make that move.
The translator's log (frozen, D1–D5) records what it did instead: it put the relation on
vocatives (Samoylenko's голубчик / братец are kept and then allowed to vanish) and on
politeness formulae ("I must ask you…", "Be so good as to…"), and at ¶686 it added a word
("what did you— what did you say, sir?"), recording that no nil-addition option kept both the
self-interruption and the correction.
The wire, one sentence. The translation limb generates the problem — a translator facing a change of address that the target language cannot make as a change of address — and the study limb asks whether any of the things a translator might do there is recoverable by a reader who is not told to look for it.
3. Contamination, measured on the frozen artifact
tools/dependence_check.py, cells in run/contamination.json, run after the freeze commit
0bad3cc and before any site was selected. The project has no prior rendering of «Дуэль» —
workshop/translations/ carries no duel directory; note (bhb)'s repository check was run
before the translation, not after it.
The comparator is Constance Garnett, The Duel (Project Gutenberg #223, 1916), chapter XV.
The lead does not write any arm of this experiment regardless of the figure — the four hands are
independent models. The measurement governs how T-duel-R06-v1 may be cited in §7's census, and
that is all it governs.
4. Materials
Source. «Дуэль» (1891), ru.wikisource copy of ПСС т. 7 (FEB), PD-old-70;
materials/duel.txt, 993 paragraphs, 31,931 words, whitespace-normalised by materials/prep.py.
Candidate pool, generated mechanically by pool.py, committed before any span was read for
content. A candidate is a maximal contiguous paragraph run within one chapter such that:
- 110–260 words;
- all second-person address tokens in the run are of one register —
ты-side orвы-side — counting the pronoun paradigms and, for verbs, unambiguous 2sg endings (-ешь/-ёшь/-ишь/-шься) and 2pl imperatives (-йте/-ите); - ≥ 6 address tokens in the run, ≥ 2 in each word-count half;
- disjoint from every other candidate and from ¶660–700, the span the lead translated, so no hand ever renders the lead's passage.
The pool is 35 candidates — 18 ты, 17 вы. No candidate was chosen or rejected by reading it.
5. Procedure
snapshot open—GET /api/v1/key. Done: usage 80.804849541.- Stage 1 — independent scene screen. All 35 candidates go to one non-hand, non-rater seat
(
nvidia/nemotron-3-ultra-550b-a55b) with a neutral prompt that names no device and no hypothesis: for each passage, how many distinct people are being spoken to (ONE/MORE THAN ONE/UNCLEAR), the name the addressee is called in the text or NONE, and the name of the person doing the speaking. This removes the lead's discretion over the one judgement the manipulation depends on — that the passage has a single addressee, so a change of address reads as this person's manner changed rather than he turned to somebody else. - Stage 2 — site selection, by a rule fixed here and applied to stage 1's output without
further judgement. From the candidates marked
ONE, walk the pool in index order taking alternately aтыcandidate and aвыcandidate until 8 of each are held. In that list of 16, ordered by pool index, positions 0, 4, 8, 12 arePOSand the other twelve areMAIN. If fewer than 8 of either register survive the screen, the deficit is made up from the other register and the imbalance is reported. - Stage 3 — the manipulation. For each site the split point is the paragraph boundary
nearest the word-count midpoint. Only text at or after the split is touched.
- MAIN — the pronoun trajectory. Arm F = as printed. Arm S = every address token at
or after the split converted to the other register, with agreement corrected (pronoun
paradigm, 2sg↔2pl verb forms, imperatives, reflexives, past-tense number). Nothing else
changes: no word is added or removed.
- POS — the vocative trajectory, the positive control. At or after the split, the first
utterance receives a vocative in both arms — familiar in one, formal in the other
(first name / голубчик against name+patronymic / господин + surname), chosen to fit the
dyad the screen named. Pronouns are left alone. Both arms of a POS site are therefore
modified, and they differ from each other in exactly one constituent.
- Direction:
тыsites move toward formality in arm S,выsites move away from it, in both blocks. Every rating is oriented so that positive = in the direction the manipulation went, which makes a rater's standing tendency to read later text as cooler cancel across the two directions rather than accumulate. - Manipulated texts are committed torun/sites.jsonbefore dispatch. - Stage 4 — the hands. Four, none the lead, none Anthropic: H1
deepseek/deepseek-v4-pro, H2openai/gpt-5.6-terra, H3x-ai/grok-4.5, H4mistralai/mistral-medium-3-5. Each receives all 16 spans in one call, numbered, in site order, and is asked only to translate. No hand sees both arms of any site: at site i, hands(i mod 4)and((i+1) mod 4)get arm F and the other two get arm S, so each hand carries 8 of each and every one of the six model pairs appears in both the same-arm and the cross-arm configuration across the site set (the repairE-20260808gA1 made, imported rather than rediscovered). - Stage 5 — the reading probe, the measure. Three seats, none a hand: R1
google/gemini-3.6-flash, R2qwen/qwen3.7-max, R3z-ai/glm-5.2. Two calls each, 16 items per call, one rendering per site per call, rendering index(r + 2k + i) mod 4at site i for rater r, call k — so no call ever contains two renderings of the same span, every rendering is rated at least once, and arms mix within every call. Calls are stateless; a seat carries nothing between them. The prompt, verbatim, names no device and does not say these are translations:Below are N short passages of English prose fiction. They are unrelated to one another. For each passage, consider how the speaker treats the person he or she is speaking to. Comparing the later part of the passage with the earlier part, does the way the speaker treats that person become more distant and formal, become warmer and more familiar, or stay about the same? Answer with exactly one of:
MUCH MORE DISTANT,SOMEWHAT MORE DISTANT,ABOUT THE SAME,SOMEWHAT MORE FAMILIAR,MUCH MORE FAMILIAR. Then quote up to twelve words from the passage that most influenced your answer, or writeNONE. Output exactly N lines and nothing else, each<number>|<answer>|<quote>.
Coding: MUCH MORE DISTANT +2 · SOMEWHAT MORE DISTANT +1 · ABOUT THE SAME 0 ·
SOMEWHAT MORE FAMILIAR −1 · MUCH MORE FAMILIAR −2, then oriented per §5.4.
7. Stage 6 — the source-side gate. One seat that is neither a hand nor a rater
(moonshotai/kimi-k3) receives the 16 Russian spans — one arm per site, 8 F and 8 S by site
parity — and answers the same question in the same format, plus, per item, OK or
FLAG:<what> for anything ungrammatical or unnatural in the Russian. The flag field is what
catches a botched agreement in the lead's manipulation; the rating field is what establishes
that the manipulation is legible as a change of manner in the source.
8. Stage 7 — mechanical census, $0. Over the post-split English of every rendering: vocative
count, honorific tokens (sir, madam, ma'am, Mr, Mrs, Miss, my dear, old man,
brother, my friend), contraction rate, mean sentence length. Arm S against arm F.
9. Stage 8 — the published census, $0. Garnett 1916 at the two authorial switches of ¶681 and
¶686, beside T-duel-R06-v1. Descriptive; nothing is predicted.
10. snapshot close; reconcile per-response usage.cost against the key delta.
Judgment is not parallelised. No model rates its own output; the hands, the raters, the screen, the gate and the critic are five disjoint sets of seats.
6. Predictions, registered
P0 — the source-side gate. On the 12 MAIN sites, the oriented source-side rating of arm S minus arm F is ≥ +1.00.
P1 — the instrument gate. On the 4 POS sites, the oriented English-side mean of arm S minus arm F is ≥ +0.75.
P2 — the primary, null-favouring. On the 12 MAIN sites, the oriented English-side mean of arm S minus arm F is ≤ +0.40.
P2 holds → four independent hands do not put a source's pronoun-borne change of address into their English in any form a reader recovers, while the same reader recovers a vocative-borne change from the same hands on the same material. P2 fails → the change is carried, and Q-d's answer is that English does have devices for it.
Registered as null-favouring and reported as such: holding P2 is failing to reject. P1 is the yardstick that makes P2 readable — it is the same manipulation (how the speaker addresses the other person, changed at the same point in the same passages, judged by the same seats on the same scale) realised through a device English possesses. Note (bkt): a control one component away is a demonstration that the instrument works on something else. English has no T/V pronoun, so a strictly exponent-matched control does not exist; the vocative is the nearest thing there is, inside the same system, and this design says so rather than pretending otherwise.
P3 — descriptive, no prediction. Stage 7's mechanical census.
P4 — descriptive, no prediction. Stage 8's published census.
7. Failure criteria, frozen
- F1 — P0 < +1.00 → the manipulation is not legible as a change of manner even to a reader of the Russian, and P2 is withheld.
- F2 — P1 < +0.75 → the instrument cannot see a carried change of address at all, and P2 is withheld.
- F3 — the arm-F oriented mean on MAIN is itself ≥ +0.75 → the unmanipulated baseline is already being read as a change and the contrast is uninterpretable; P2 is withheld and the baseline is reported.
- F4 — a hand returns other than 16 numbered items, or renumbers → one re-dispatch; a second failure drops the hand and every quantity is recomputed on three.
- F5 — any hand's rendering shares ≥ 8 contiguous tokens with Garnett 1916 at a site → that item is flagged; at ≥ 3 such items the hand is dropped and P2 recomputed without it.
- F6 — more than 2 unparseable lines in a rater call → one re-dispatch; a second failure drops that call and the coverage loss is reported.
- F7 — the source-side gate returns
FLAGon more than 2 arm-S MAIN items → the manipulation is defective; those sites are dropped and P2 is recomputed, with the drop reported. - No criterion is weakened after it fires. If a gate fires the number is still reported.
8. Pre-flight cost estimate
Built from max_tokens, not from expected length (note (abc)).
| stage | calls | model | cap | worst case |
|---|---|---|---|---|
| 1 screen | 1 | nemotron-3-ultra | 4,000 | $0.03 |
| 4 hands | 4 | deepseek / gpt-5.6-terra / grok-4.5 / mistral-medium | 9,000 | $0.19 |
| 5 raters | 6 | gemini-3.6-flash ×2, qwen3.7-max ×2, glm-5.2 ×2 | 2,500 | $0.14 |
| 6 source gate | 1 | kimi-k3 | 3,000 | $0.05 |
| 0 critic | 1 | nemotron-3-ultra | 12,000 | $0.09 |
| re-dispatch reserve | 3 | — | — | $0.17 |
| declared worst case | $0.67 |
Today's UTC ledger stands at $3.749282 of $5.00, headroom $1.250718. The worst case fits with $0.58 to spare. Stages 7–8, the whole translation limb and all analysis cost $0.
9. What this cannot show
- Four 2026 models are not four human translators. A holding P2 is these hands did not carry it, not it is uncarriable. The published census in stage 8 is the only human evidence here and it is n = 1 translator at n = 2 sites.
- A small study cannot prove a null. P2's meaning comes entirely from P1 standing beside it: a difference of the same kind, in the same passages, judged by the same seats, that these hands do carry. Without P1 clearing its bar the run says nothing, which is why F2 exists.
- Arm S of a MAIN site is not Chekhov. A manufactured switch may be less motivated than an authorial one, and an unmotivated switch is one a hand might normalise away. The source-side gate measures whether it is legible, not whether it is natural; this is a real limit on generalising to authorial switches, and the stage 8 census on the one authorial switch in hand is the only check on it.
- POS's arms are both manufactured. Neither is the printed text.
- The raters are told nothing, but they are still models. A reader who has seen a great deal of translated Russian may have a prior about what "later in the passage" does.
- The whole design is one language pair and one author. Q-d is not answered by it; step 2 of
ARM-trajectoryis the second pair.
10. Amendments, after the pre-run critic and before dispatch
Three critic seats were needed. The registered critic nvidia/nemotron-3-ultra-550b-a55b
returned finish_reason: length with zero content at a 12,000 cap ($0.046932); the declared
fallback moonshotai/kimi-k3 did the same at 10,000 ($0.168210). Note (bhf) — zero content
means change the seat, not raise the cap — fired twice in one stage, and both dead bodies are
ledgered as waste. The pass that returned is mistralai/mistral-medium-3-5 (non-panel, the
screen seat, not a hand and not a rater), $0.031467, 11,198 characters: verdict
NEEDS-REDESIGN, 5 BLOCKING and 5 ADVISORY.
A0 — the roster, forced by the two dead seats. Recorded before dispatch.
Nine seats were available and eleven roles wanted filling, so two things moved:
| role | seat | why |
|---|---|---|
| hands | 3, not 4 — deepseek/deepseek-v4-pro, openai/gpt-5.6-terra, x-ai/grok-4.5 |
mistral-medium-3-5 became the critic |
| raters | unchanged — gemini-3.6-flash, qwen3.7-max, glm-5.2 |
none is a hand |
| source-side gate | deepseek-v4-pro and gpt-5.6-terra, two seats, complementary arms |
kimi-k3 was dead. Both are hands, but they rate the Russian source, never any English rendering, so no seat judges its own output; calls are stateless |
| screen, critic | mistral-medium-3-5 |
two calls, no rendering ever seen |
With 3 hands each site is rendered 3 times, split 2/1 between arms, the singleton arm alternating
with the site's position and the singleton hand rotating — arm counts per hand came out 8/8, 9/7,
7/9 and every site got both arms (assign.py).
A1 — the POS block was 3 ты and 1 вы (critic BLOCKING 1). Accepted; the selection rule is replaced.
The frozen rule (positions 0, 4, 8, 12 of the pooled ordered list) mechanically returned three
ты sites and one вы, confounding block with register in the one comparison the primary is read
against. Replaced before any arm was rendered: POS = the 1st and 5th site of each register
in the ordered list → sites 0, 5 (ты) and 7, 15 (вы), 2 and 2. MAIN is now 6 and 6.
A2 — the raters were not told where "later" begins (critic BLOCKING 4). Accepted, and it changed the whole pipeline.
The split point is a fact of the design and was going to be invisible to the raters, who would each
choose their own boundary and dilute the contrast with noise the design could not see. Fixed by
making the boundary explicit and identical in both arms: every item is printed as PART ONE /
PART TWO at the true split. To make that possible in the English, every source paragraph is
numbered [site.paragraph] and the hands are required to return one line per input line — which
also gives a mechanical completeness check. All three hands returned 16 of 16 sites complete, 151
of 151 paragraphs, 0 junk lines.
A3 — P2 is also reported split by register (critic BLOCKING 2). Accepted in part; the orientation itself is overruled.
Orienting each rating so that positive means in the direction the manipulation went is what makes the two directions comparable and is the design's stated guard against a rater's standing tendency to read later text as cooler; the critic's claim that it puts the two registers on different scales is not accepted. What is accepted is that the pooling could hide a register-specific effect, so P2 is reported pooled and split, with the raw (unoriented) arm means beside it.
A4 — F3 restated (critic BLOCKING 5). Accepted as a clarification, not a change.
F3 reads on the oriented mean of the arm-F renderings on MAIN, which is not zero by construction — arm F is the printed text and gets whatever rating it gets.
A5 — P1 stays where it is (critic BLOCKING 3). Overruled, with the reason.
The critic is right that a vocative is not a pronoun and asks for a control "closer to the primary's mechanism"; the control it proposes — a title against a bare pronoun — is the vocative control. §6 already states that English has no T/V pronoun and therefore no exponent-matched control exists. Lowering P1's bar before dispatch on the speculation that vocatives are weak would be weakening a gate to make a null easier to read, which is the one thing a gate may not do. P1 stays at +0.75, and if it fails the run says so.
A6 — advisories
ADVISORY 2 (grammaticality of arm S) accepted and answered inside the design: the residual
checker returned 0 residual forms across all 16 sites, every hand edit is committed in
overrides.json as an ordered phrase rule whose find-string is a hard error if absent, and the
source-side gate carries a per-item OK/FLAG field. ADVISORY 4 (thresholds are round numbers)
accepted as a limit and reported: no power analysis was run, and the observed rater spread is
reported beside every figure. ADVISORY 3 is A0 above. ADVISORY 1 was moot — both registers
had 8 eligible sites. ADVISORY 5 (cost): the declared worst case stands at $0.67; the actual was
$0.644078.
A7 — deviations from the frozen text, recorded rather than smoothed
- §5.2 wrote the screen's question as how many distinct people are being spoken to (
ONE/MORE THAN ONE); the prompt as dispatched asked how many people take part in the conversation (TWO/MORE-THAN-TWO). Same criterion — a two-party exchange — differently encoded. 29 of 35 candidates came backTWO. - §5.4 said the POS vocative goes into "the first utterance at or after the split". As applied:
where the tail already carries a vocative from one participant to the other, that vocative is
replaced; where it does not, one is inserted, identically placed in both arms. Where a tail
carried two such slots they were moved together, so that neither arm carries a change of its own.
Fixed in writing in
overrides.jsonbefore any arm was rendered. - Three dispatch-level failures were re-sent beyond F4's single re-dispatch, all of them
parameter-level rather than content-level:
deepseekandgrokeach spent an entire cap on hidden reasoning and returned empty content, and were re-sent with reasoning disabled (deepseek, worked) and at low reasoning effort (grok, worked after a 400 on the first parameter form). Every body, dead or alive, is inrun/and in the ledger.