Repository path: workshop/experiments/E-20260810t-drift-source-present/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260810t-drift-source-present |
| status | frozen |
| created | 2026-08-10 |
| updated | 2026-08-10 |
| senses | consistency |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-terminology-drift.md, wiki/findings/results/RS-20260809i-terminology-drift.md, workshop/experiments/E-20260809i-terminology-drift/design.md, workshop/translations/tini-budz/R04-v1/translation.md, workshop/translations/tini-polonyna/R04-v1/translation.md, wiki/goodness-senses.md |
E-20260810t-drift-source-present — is the zero a fact about the sense, or about a source-blind reader?
ARM-terminology-drift step 2 of 2, T2. Frozen before dispatch; the pre-run critic pass is
§9 and every amendment it forces is recorded there with the wording it replaced.
1. The question
RS-20260809i (S147) found that a jury scoring consistency on the English alone rates a
passage that renders the same realia term two different ways exactly as coherent as a
composition-matched passage that renders every term one way: mean difference 0.000, nine of
twelve units tied, P = 0.875, against a comparator carrying the identical number of foreign tokens.
The same seats, the same scale and the same passages dropped 2.000 points when one proper name
was spelled two ways.
That result has a rival reading of equal standing, and the result page registered it as the run's
sharpest limit (§5.1). voice was scored with the Ukrainian on the page and consistency was not,
because consistency is defined as target-internal. So the finding may be about the sense —
terminology drift is not a coherence event — or about the reader — a reader of the English alone
has no way to know that polonyna and the high pasture answer to one Ukrainian word, and drift is
visible as drift only against a source.
This design separates them. The same operator, the same scale, the same rubric string, the same four seats, in two conditions that differ in exactly one thing: whether the Ukrainian is on the page.
Subject-rule statement (continue-prompt §4.5). What this unit teaches about evaluating
translations: whether internal coherence is a property a reader can assess without the source at
all, or whether the sense the typology calls target-internal is in fact a source-relative sense
wearing a target-internal definition. That is a claim about a sense of good, not about the
project's apparatus.
2. Materials
Two units of the same novella, six blocks, 123 marked class sites.
| unit | blocks | words | class sites | status |
|---|---|---|---|---|
tini-polonyna (S147) |
BA BB BC |
717 Ukr | 54 | arms imported unchanged from E-20260809i/materials/arms.json |
tini-budz (this session) |
BD BE BF |
769 Ukr | 69 | built here; T-tini-budz-R04-v1, frozen at 2a484ab |
tini-budz is Kotsiubynsky's milking and cheese-making sequence, four pages further on in the same
novella — the same institution, the same class of culture-bound vocabulary, translated in this
session by the lead under R04 with its R06 draft frozen first. Provenance, the one elided verse
passage, the four editorial glosses that were read, and the within-block repetition table are on
workshop/translations/tini-budz/README.md. It doubles the units of analysis from 12 to 24 and
supplies a second, independently translated drift realization — RS-20260809i §5.3 named the
single realization as a limit and could not afford a second.
Arms. Generated mechanically by materials/arms.py (57 assertions, all passing before
dispatch) from materials/template.txt, which is itself built by substitution on the frozen R04
text (materials/build_template.py) rather than retyped:
| arm | what it is | uniformity |
|---|---|---|
DOM |
every class item in its English form | class-uniform, 0.000 foreign |
MATCH |
item-uniform; COPY-token count equal to DRIFT's in every block |
item-uniform |
DRIFT |
within a block, every item occurring twice or more takes both forms | item-NON-uniform |
FALSE |
MATCH plus one planted propositional error per block; no drift |
item-uniform, and factually wrong |
SPELL |
DOM with one proper name spelled two ways, nothing else changed |
the positive control |
MATCH is the whole design, and it is E-20260809i's pre-run critic's BLOCKING 1: it carries the
same number of foreign tokens as DRIFT in every block, from as nearly as possible the same items,
and differs from it only in whether the variation falls between items or within them. Residual
per-item L1 on the new unit: 10 / 6 / 8 tokens.
Composition on tini-budz, fraction of class tokens in COPY form:
| block | sites | COPY |
MIXED |
MATCH |
DRIFT |
DOM |
|---|---|---|---|---|---|---|
BD |
26 | 1.000 | 0.692 | 0.615 | 0.615 | 0.000 |
BE |
20 | 1.000 | 1.000 | 0.700 | 0.700 | 0.000 |
BF |
23 | 1.000 | 1.000 | 0.739 | 0.739 | 0.000 |
Fifteen items drift within a block (against S147's smaller set): BD diynytsia ×6, vivchar ×5,
strunka ×4, kozar ×3, zahoroda ×3, honinnyk ×2; BE vatah ×5, staya ×2, berbenytsia ×2, budz ×2,
zhentytsia ×2; BF vatah ×7, spuzar ×3, budz ×3, staya ×2.
Two declared asymmetries between the units.
MIXEDis class-uniform inBEandBF. Ontini-budzthe base translation domesticated only two class items (вівчар→ shepherd,загорода→ pen) and both occur only inBD. So the class-non-uniform exhibit this project has been reading since S012 lives, on this unit, in one block of three.tini-polonyna'sMIXED(0.875 / 0.789 / 0.842) remains the exhibit;MIXEDhere is reported and carries no registered test either way.SPELLexists forBA,BBandBDonly.BChas no proper noun occurring twice (S147's finding);BE's single name occurs once andBFcontains none. The control therefore runs on 3 blocks × 4 seats = 12 units per condition, against 24 for the primary.
3. The two conditions, and what separates them
SB — source-blind |
SP — source-present |
|
|---|---|---|
| header | "you are not being shown the original" | "You will be shown the Ukrainian passage, and then one complete English rendering of it" |
| body | the English arm alone | the Ukrainian block, then the English arm |
| rubric string | CONS_DEF, imported verbatim from E-20260809i/run.py and asserted equal to it |
identical |
| the one changed clause | "…and you have no source text to check it against: you are being asked only whether the passage hangs together with itself." | "…and you are NOT being asked whether it is a faithful rendering of the Ukrainian, which is given for reference only: you are being asked only whether the passage hangs together with itself." |
The manipulation is information, not question. Both conditions ask the same target-internal question and both explicitly refuse the accuracy reading. What changes is whether the reader can know that two English strings answer to one Ukrainian word. If the sense is genuinely target-internal, that knowledge is irrelevant and the two conditions agree.
4. Procedure
- One arm per call. No seat is told that any other rendering exists, or that a condition exists (E-20260809i critic, BLOCKING 4).
- Both conditions dispatched in ONE interleaved pass, order shuffled once from frozen seed
20260810and stored atmaterials/dispatch-order.jsonbefore any call. RunningSBto completion and thenSPwould confound the condition with anything that changes at a provider over the hour the run takes, and the condition contrast is the primary. - Four seats, the same slugs in the same roles as S147:
openai/gpt-5.6-terra,google/gemini-3.6-flash,deepseek/deepseek-v4-pro,x-ai/grok-4.5. Not five or six: theSBarm is a replication of S147's C stage, and comparability is worth more here than the marginal power two more seats would buy. - Unit of analysis = block × seat. 6 × 4 = 24 for every primary and secondary; 3 × 4 = 12 for the control.
- Cells. 2 conditions × 4 seats × (4 arms × 6 blocks +
SPELL× 3 blocks) = 216 calls (DOM,MATCH,DRIFT,FALSE, andSPELLon three blocks). Total re-dispatches capped at 20 (critic MINOR 12); beyond that the run stops and reports incomplete. temperature: 0.0,max_tokens: 4000. A body that returnsfinish_reason: lengthwith no content is moved toruns/discarded/and re-dispatched once at 8,000 with reasoning disabled; the dead body is kept and ledgered (notes (bkw), (blj)). A body that returnslengthwith readable content is NOT silently accepted — note (bls), S151: it is counted, named in the result, and its score used only if the JSON object is complete.
5. Registered predictions and bars
All tests are exact sign tests by enumeration, one-sided where a direction is registered, ties excluded from the test and reported. Nothing below may be altered after any body is read.
Let, for unit u = (block, seat):
d_SB(u) = consistency(MATCH,SB,u) − consistency(DRIFT,SB,u),
d_SP(u) = consistency(MATCH,SP,u) − consistency(DRIFT,SP,u).
| # | test | statistic | bar | what it decides |
|---|---|---|---|---|
| P1a | THE PRIMARY, leg 1. Δ(u) = d_SP(u) − d_SB(u) > 0 |
sign test, 24 units, one-sided | P ≤ 0.05 | drift is penalised more when the source is visible |
| P1b | THE PRIMARY, leg 2 (critic BLOCKING 1). McNemar on drift-responsiveness: units with d_SP > 0 ≥ d_SB against units with d_SB > 0 ≥ d_SP |
exact binomial on the discordant units, one-sided | P ≤ 0.05 | the interaction is not just P2 wearing a difference |
P1 is ESTABLISHED only if BOTH legs clear 0.05. |
||||
| P2 | the simple effect with the source present: d_SP > 0 |
sign test, 24 units, one-sided | P ≤ 0.05 | consistency sees drift at all, given the source |
| P3 | the S147 null, re-measured: d_SB |
sign test, 24 units, two-sided, plus the mean | no verdict word. Reported as: mean, signs, exact P, and whether \|mean\| falls inside 25% of the SB SPELL effect (critic BLOCKING 2) |
whether S147's exact zero is a property of the instrument or of that run |
| C1 | THE GATE — scale room, in EACH condition separately: DOM > SPELL |
sign test, 12 units, one-sided | P ≤ 0.05 in both | whether a null in that condition means anything |
| C2 | descriptive. d_SB and d_SP split by unit (polonyna / budz) |
means and signs | — | whether the new passage behaves like the old |
| C3 | descriptive. SB means on BA BB BC against S147's stored figures |
difference of means | — | cross-session test–retest on identical prompts |
| C4 | THE MECHANISM CONTROL (critic BLOCKING 3). MATCH > FALSE, in each condition, and the interaction |
sign test, 24 units, one-sided | P ≤ 0.05 | whether the seats use the source for fidelity while told not to |
| D1 | registered descriptive. per-seat Δ; finish_reason rate by condition; fractional per-item L1 |
tables | — | critic SERIOUS 6, MINOR 9, SERIOUS 7 |
| D2 | descriptive. the arm-mean table, per condition, per block | means | — | the composition row S147 §2 made readable |
Every P below indexes stochasticity within these four fixed model seats only (critic SERIOUS 6). None of them is an inference to a population of readers, and none may be reported as one.
Registered expectation, stated so it can be wrong. I expect P1 to fire and P3 to replicate — that is, the zero is about the reader. The alternative outcome that would most change the project's sense list is P1 and P2 both failing while C1 passes in both conditions: that would say terminology drift is not a coherence event for these readers even when they can see it is drift, and the parenthesis's first half would have to be struck rather than qualified.
6. Failure criteria and withholding rules
- F1 — the gate. If
C1fails in a condition, every null in that condition is uninterpretable and is reported as such. IfC1fails in either condition,P1is withheld entirely: a difference of differences needs both differences to mean something.P3may still be reported ifC1passes underSB. - F2 — completeness. If fewer than 95% of the 216 cells return a parseable integer in range, the run is reported as incomplete and every figure carries the completeness rate. Below 90%, no registered test is reported at all.
- F3 — G1 parity.
mistralai/mistral-medium-3-5(not a jury seat) checks, with the Ukrainian shown, six real arm pairs on the new unit plus four planted propositional errors. Bar: at least 4 of 6 real pairs equivalent AND at least 3 of 4 planted errors caught. Below either, the twelve DOM strings written fortini-budzare declared unverified and everybudzcontrast involvingDOMorMATCHis withheld — a DOM string that says something different would sit inMATCHandDRIFTa different number of times and is therefore a live confound, not a cosmetic one. The old unit's DOM forms passed this gate at S147 (8 of 8 pairs, 4 of 4 planted) and are imported unchanged. - F4 — no post-hoc analysis choice. Test, units, direction, bar and tie handling are fixed above.
analyse.pyis written before dispatch andverify.pyrecomputes every reported number by independent routes.
6a. G1 parity — run before dispatch, and it took two seats
Outcome: 4 of 4 planted errors caught, 5 of 6 real pairs equivalent. Nothing withheld.
The first seat failed the task, not the materials. mistralai/mistral-medium-3-5 returned
non-equivalent for all ten pairs — the six real ones and the four planted ones alike — citing in
every case exactly the lexical differences the prompt excludes ("A uses 'milking gate' vs B uses
'strunka'"). A gate that never returns equivalent has no discriminative power and cannot bar
anything; its three bodies are kept in runs/discarded/ and ledgered at $0.030216. The seat was
a poor choice on my part: config/models.md already records it as the panel's weakest and
not panel material.
The second seat is moonshotai/kimi-k3, which passed this identical task at S147 (8 of 8, 4 of
4) and is not a jury seat. $0.047766.
| pair | verdict | note |
|---|---|---|
BD MATCH/DRIFT |
equivalent, empty difference list | |
BE MATCH/DRIFT |
equivalent, empty difference list | |
BF MATCH/DRIFT |
equivalent, empty difference list | the primary contrast, verified in all three blocks |
BD DOM/MIXED |
equivalent | |
BF DOM/MIXED |
equivalent | |
BE DOM/MATCH |
non-equivalent | "A says fourteen head of sheep, B says fourteen drobyeta" |
| four planted errors | all four caught, each named correctly | black shaggy → fair clean-shaven; third day → third hour; fourteen → forty; sun goes down → sun comes up |
Item-level inspection, as the amended F3 requires at 5 of 6. The single flag is on drobyeta,
whose DOM form is head of sheep at the BE site («Мосійчук має штирнадцять дробєт»). That
rendering is denotationally correct, and the "difference" the seat named is the
foreign-versus-English difference the prompt explicitly excludes — the same error the first seat made
ten times out of ten, made once here. The DOM string is upheld and nothing is withheld, and the
flag is recorded rather than argued away. It touches no MATCH/DRIFT contrast in any case.
7. Pre-flight cost estimate
Worst case is built from the caps the requests actually permit, not from expected output (note (abc)).
| stage | calls | cap | worst-case line |
|---|---|---|---|
| pre-run critic | 1 | 16,000 | $0.10 |
| G1 parity | 3 | 8,000 | $0.06 |
| scoring | 216 | 4,000 (8,000 on one retry each) | $2.40 |
| re-dispatch allowance | ≤ 20 | 8,000 | $0.30 |
| declared ceiling | $2.86 |
Expected, from S147's realised $0.0053/call over 128 calls with half of them source-bearing: ≈ $1.50. Today's UTC ledger stands at $0.867877 of $5.00 before this run, so the ceiling fits the headroom with $1.27 to spare.
8. Declared limits, written before the run
- The translator knew the design.
tini-polonynawas translated beforeE-20260809iexisted;tini-budzwas translated in this session withRS-20260809i§8's successor design already written down. The consequence is bounded and stated:MIXEDon the new unit is the lead's own class handling and carries no weight in any registered test (§5, D1).MATCHandDRIFTare generated from the marked sites and are invariant to which form the base chose at any site, so the primary is untouched by what the translator knew. - Four model seats, not readers. Tier D is NOT PASSED; every figure is
provisionalandinternal-judgment-only. - One work, one class, one translator, two passages. The class is Hutsul pastoral realia and nothing here generalises past it without being shown to.
- Two drift realizations, not many. One seeded assignment per unit. Better than S147's one; not a distribution.
SPmay buy its effect with an accuracy halo. The prompt refuses the accuracy reading explicitly, and refusing it in words is not the same as preventing it. IfP1fires, the honest statement is the source's presence changes the score, and the mechanism — drift becoming visible, or accuracy leaking in — is not settled by this design. That is registered here so it cannot be quietly dropped from the result.C1's 12 units are half the primary's 24, so the gate is the weakest-powered test in the run and aC1failure could itself be a power failure. IfC1fails, that reading is reported beside the withholding.
9. Pre-run critic — NEEDS-AMENDMENT, 12 findings, 3 BLOCKING
One independent adversarial pass over the frozen design and its predecessor result,
moonshotai/kimi-k3 — not a jury seat and not S147's critic. 11 accepted, 1 accepted in part,
none overruled. Two calls: the first returned finish_reason: length with zero content, the
whole 16,000 cap spent on hidden reasoning ($0.266061, dead, ledgered); the re-dispatch with
reasoning disabled cost $0.053280. Note (bkw)'s remedy, fifth occurrence.
The amendment that matters, and it changed what was dispatched: BLOCKING 3. §8.5 conceded that
the source-present condition might buy its effect from an accuracy halo, and then §5 registered an
interpretation — "the zero is about the reader" — that assumed it had not. Nothing in the frozen
design separated the two. The remedy is the FALSE arm: MATCH, with no drift at all, plus
one planted propositional error per block that contradicts the Ukrainian and is invisible from
the English alone (materials/false_arm.py, 30 assertions; the six edits and the reason each is
source-blind-invisible are tabled there). It is registered as C4. MIXED was dropped to pay
for it — 48 cells out, 48 cells in, the total unchanged — and that is the right trade twice over,
because the critic's MINOR 11 showed MIXED was weaker than §8.1 had admitted.
| # | severity | what it attacked | disposition |
|---|---|---|---|
| 1 | BLOCKING | P1's Δ sign test degenerates to P2 when d_SB is near zero on most units, so the interaction is not separately tested |
accepted. P1 is now two legs, both required: P1a, the Δ sign test as frozen, and P1b, McNemar on drift-responsiveness — units responsive under SP only against units responsive under SB only. The interaction is established only if both clear 0.05 |
| 2 | BLOCKING | P3's replication rule (\|mean\| ≤ 0.25) is arbitrary and uncalibrated; a real 0.2-point cost would be declared a replication of zero |
accepted. The word replicated is struck. The band is now 25% of the SPELL effect measured in the same condition — data-anchored, by a rule fixed before dispatch — and the mean, the signs and the exact two-sided P are reported whatever it says |
| 3 | BLOCKING | nothing separates drift became visible from accuracy leaked in; P1 firing would be read the first way with no evidence against the second |
accepted — the FALSE arm, above. If C4 fires under SP and not SB, the seats are demonstrably using the source for fidelity while told not to, and a P1 firing is open to the same reading; if C4 does not fire, the halo account is much weakened |
| 4 | SERIOUS | SPELL covers BA BB BD; the gate does not gate BE BF, where a third of the primary's data sits |
accepted in part. The claim is restricted: C1 licenses interpretation of nulls on the blocks where SPELL exists, and the gap on BE/BF is declared. Not accepted: building a different control on BE/BF would make C1 a heterogeneous gate whose failure could not be attributed to anything — worse than a declared gap. Per-block DOM means are reported as ceiling evidence instead |
| 5 | SERIOUS | C1 on 12 units needs ~10 of 12 non-tied one way; a tie-driven failure would withhold P1 even if the interaction is real |
accepted. Registered tie-contingent rule: if C1 fails and ≥ 8 of its 12 units are tied with DOM at ceiling (7), that is a saturation finding and P1 is reported with the caveat rather than withheld |
| 6 | SERIOUS | four fixed model seats are not exchangeable replicates of a reader; the framing invites that misreading | accepted. §5 now says every P indexes stochasticity within these four seats only, and a per-seat Δ table is registered so a one-seat-driven P1 is visible |
| 7 | SERIOUS | residual per-item L1 of 10/6/8 means MATCH and DRIFT differ in which items carry foreign tokens, not only in within/between |
accepted, with one correction to the finding. The critic compared raw L1 against S147's 4/6/6 without dividing by site count; as fractions the two units are comparable (this unit 0.385 / 0.300 / 0.348; S147 0.211 / 0.316 / 0.375). The fractional table is now registered as a descriptive, and the honest statement of the contrast — matched on token count, not on which items carry them — is in §8 |
| 8 | SERIOUS | the design leans on S147's voice movement (0.583, P = 0.109), which missed its own bar |
accepted. P2 is registered as the first adequately-powered test of whether source-present consistency sees drift at all, standing on its own and not on S147's direction |
| 9 | MINOR | SP prompts are ~2× longer, so any length-driven effect loads onto the condition |
accepted. finish_reason rates by condition are a registered reported check |
| 10 | MINOR | F3's 4-of-6 bar can pass with a third of pairs non-equivalent |
accepted. Tightened: 6 of 6 for full release; 4–5 of 6 triggers item-level inspection and withholding limited to contrasts involving the failing DOM strings; < 4 of 6 blanket withholding as before |
| 11 | MINOR | §8.1's invariance claim covers form choice but not site selection — the lead chose which items repeat, knowing repetition is what is tested | accepted, and it is the sharpest of the twelve. §8.1 now says so, and it is part of why MIXED was the arm dropped |
| 12 | MINOR | the retry allowance is uncapped in the procedure text | accepted. Total re-dispatches capped at 20; beyond that the run stops and reports incomplete |