Repository path: workshop/experiments/E-20260809i-terminology-drift/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260809i-terminology-drift |
| status | frozen |
| created | 2026-08-09 |
| updated | 2026-08-09 |
| senses | consistency, voice, naturalness, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-terminology-drift.md, wiki/goodness-senses.md, workshop/translations/tini-polonyna/R04-v1/translation.md, workshop/experiments/E-20260809i-terminology-drift/materials/arms.py, config/models.md, config/budget.md |
E-20260809i-terminology-drift — the consistency / voice parenthesis, both halves
Frozen before dispatch. Every number this design names is registered here. The base translation
and both translator's logs were frozen earlier and independently, at commits b007b18 (draft) and
8145531 (revision), before this file existed.
⚑ READ §9 FIRST — this design was amended before dispatch
The independent pre-run critic returned
NEEDS-REDESIGNwith nine BLOCKING findings, and it was right about the load-bearing one: the additive null in §5 would have let the primary fire with no drift effect present at all. §§3–6 below are the design as first frozen and are kept verbatim as the record. Where §9 conflicts with them, §9 governs and is what ran. Nothing was dispatched under the superseded version, and the critic saw the version in §§1–8.
1. Question
wiki/goodness-senses.md, consistency entry, final clause: "Distinguish from voice (voice can
be consistent while terminology drifts, and vice versa)."
Does a blind jury's consistency score respond to terminology drift while its voice score
does not, and does voice respond to a genericising operator while consistency does not?
A second, older question rides on the same materials. The consistency entry has said since S013
that "a calibrated jury scoring class-uniform against class-inconsistent handling of the same realia
set would discriminate" (TH-20260724-translation-distance-axes C4). RS-20260803d (S097)
measured the census — what translators do — and said explicitly that the jury test was not it.
This design runs it.
2. Materials
T-tini-polonyna-R04-v1 — Kotsiubynsky, «Тіні забутих предків», the polonyna movement, 717
Ukrainian words in three blocks of 250 / 244 / 223, rendered by the lead under R04. The project's
first translation from Ukrainian; eighteenth source language. Contamination none as a reachability
statement (workshop/translations/tini-polonyna/README.md §Contamination), with the reason the
design does not turn on it stated there: every arm is a variant of one lead rendering, and every
comparison is within-run and within-translator.
The class. Twenty culture-bound items naming the polonyna as an institution, occurring at 54
sites across the three blocks. materials/template.txt is the frozen base with every one of those
54 sites replaced by a marker carrying the item id, a COPY form and a DOM form, and nothing else
touched.
Twelve items occur two or more times inside a single block — skalka ×4, polonyna ×5, ×4,
vatah ×3, ×2, ×2, staya ×3, ×2, vatra ×3, khudibka ×3, stoyishche ×2, vivchar ×2 — which
is the property the passage was selected for, checked by counting before it was read closely
(workshop/translations/tini-polonyna/README.md). Within-item drift is therefore visible inside a
single block and does not depend on a rater holding three blocks in mind.
3. The arms
Generated mechanically by materials/arms.py, which runs 63 assertions and refuses to emit
anything that fails one.
| arm | construction | uniformity |
|---|---|---|
MIXED |
each item takes the form the base used, everywhere | item-uniform, class-non-uniform |
COPY |
every item copy-opaque (transliterated), everywhere | class-uniform |
DOM |
every item domesticated to an English equivalent, everywhere | class-uniform |
DRIFT |
within a block, occurrence k takes COPY when k is odd, DOM when k is even | item-non-uniform |
GEN |
DOM with the narration genericised; class strings and all quoted speech byte-identical to DOM |
item-uniform |
SPELL |
DOM with proper-name spellings doubled and nothing else |
item-uniform |
Assertions that make this checkable rather than assertable, all passing:
MIXEDis byte-identical to the frozenR04-v1text (whitespace folded).MIXED,COPY,DOM,DRIFTare byte-identical outside the 54 class sites, proven by re-rendering each arm from a sentinel skeleton.- Every item occurring ≥2 times in a block appears in both forms in that block in
DRIFT— 12 of 12. GENcarries eachDOMclass-form string exactly as often asDOMdoes, and carries no COPY form anywhere.SPELLreduces toDOMexactly by undoing the name spellings; no other difference exists.
Composition — the fraction of the 54 class tokens standing in COPY form:
| block | sites | COPY |
MIXED |
DRIFT |
DOM |
|---|---|---|---|---|---|
BA |
16 | 1.000 | 0.875 | 0.750 | 0.000 |
BB |
19 | 1.000 | 0.789 | 0.684 | 0.000 |
BC |
19 | 1.000 | 0.842 | 0.737 | 0.000 |
| all | 54 | 1.000 | 0.833 | 0.722 | 0.000 |
This is the design's central fact. MIXED and DRIFT sit close on the only axis the four
lexical arms vary along — how much foreign residue the English carries — and differ almost entirely
in whether the variation is between items or within them. The two class-uniform arms bracket
them and supply the additive null.
SPELL exists for BA and BB only. BC contains no proper noun occurring twice, so no
name-spelling doubling is constructible there without adding content. This is a declared scope limit
on the positive control, not a result.
GEN's operator (narration only; every string inside quotation marks is byte-identical to
DOM): exclamation → statement; ellipsis → full stop; inversion → canonical order; paratactic
and…and chains → subordination or segmentation; idiosyncratic phrasing → ordinary phrasing
(gathered into themselves → concentrated, He had to hurry → It was necessary for him to
hurry). Held: register band, every proposition, every image (pride embraced Ivan's soul, the
fire wound its adder's body, the earth sighs gladly, the Beskyd knitted his brows, vestments of
rose and gold), the present→past tense shift in BC, and all terminology.
4. Procedure
Seats (config/models.md): J1 openai/gpt-5.6-terra, J2 google/gemini-3.6-flash,
J3 deepseek/deepseek-v4-pro. The same three seats as S134, S141 and S146.
Four stages, one sense each, each in its own call. This is not economy — it is wiki/goodness-senses.md
usage rule 6, generalised. S135 measured a within-call halo of about +0.38 of seven when senses
are scored together, and the two senses under test here are the two whose separation is the
question, so scoring them in one call would put the halo exactly where the finding is.
| stage | sense | source shown | arms | orderings | calls |
|---|---|---|---|---|---|
V |
voice |
Ukrainian present | MIXED COPY DOM DRIFT GEN |
2 | 18 |
C |
consistency |
English alone | the same five, plus SPELL in BA/BB |
2 | 18 |
N |
naturalness |
English alone | the five | 1 | 9 |
A |
accuracy |
Ukrainian present | the five | 1 | 9 |
consistency is scored on the English alone because it is a target-internal-coherence sense; voice
and accuracy are source–target relations and get the Ukrainian. Definitions are quoted verbatim
from wiki/goodness-senses.md; for consistency, the entry's definitional sentences only ("Internal
coherence across the whole text… the long work is its test bed"), the rest of the paragraph being
evidence rather than definition. That truncation is declared here and is the one place a rubric
string was cut.
Unit of analysis: block × seat × ordering — 18 units for stages V and C, 9 for N and A,
12 for the SPELL control. All tests are exact one-sided sign tests over units.
Order. Arms are presented in a per-call permutation from a seed frozen in materials/orders.json
before dispatch; the second ordering is the reverse of the first.
5. Predictions, registered
Write p(arm) for that arm's fraction of class tokens in COPY form, per block. The additive null
for an arm of composition p is the linear interpolation of the two class-uniform arms:
pred(arm) = DOM + (COPY − DOM) · p(arm)— computed per unit.dip(arm) = pred(arm) − observed(arm), so a positive dip is a cost not explained by composition.
| # | prediction | statistic | bar |
|---|---|---|---|
| P1 | primary. Item-level drift costs consistency beyond composition |
dip(DRIFT) on consistency > 0 |
sign test, 18 units, one-sided P ≤ 0.05 |
| P2 | primary — the parenthesis. The drift cost is larger on consistency than on voice |
dip(DRIFT)(consistency) − dip(DRIFT)(voice) > 0 |
sign test, 18 units, P ≤ 0.05 |
| P3 | Genericising the narration costs voice |
DOM − GEN on voice > 0 |
sign test, 18 units, P ≤ 0.05 |
| P4 | the reverse half. That cost is larger on voice than on consistency |
(DOM−GEN)voice − (DOM−GEN)consistency > 0 |
sign test, 18 units, P ≤ 0.05 |
| P5 | the C4 jury test. Class-level non-uniformity, item-uniform, costs consistency |
dip(MIXED) on consistency > 0 |
sign test, 18 units, P ≤ 0.05 |
| P6 | Item-level drift costs more than class-level | dip(DRIFT) > dip(MIXED) on consistency |
sign test, 18 units, P ≤ 0.05 |
| P7 | specificity, reported not gated. The drift dip is a consistency effect |
dip(DRIFT) on naturalness and on accuracy |
reported with exact P; a dip ≥ 0.50 on either weakens P1 and is said to |
P2 and P4 are the load-bearing pair, and they are the two that survive the obvious objection. A convex response to proportion copied would manufacture a dip with no drift effect at all — but it would manufacture it on every sense, and P2 and P4 are differences of dips within a unit. The double dissociation is established only if both P2 and P4 hold.
Standardised secondary. Because the two senses need not use the 1–7 scale alike, P2 and P4 are recomputed with each sense's per-seat SD across arms divided out, and both figures are reported. If raw and standardised disagree in sign, neither is reported as establishing the claim.
6. Gates and failure criteria
| id | gate | bar | if it fails |
|---|---|---|---|
| G1 | Parity. An independent adjudicator (P4 moonshotai/kimi-k3, not a jury seat), shown the Ukrainian, judges eight registered arm pairs propositionally equivalent, and is shown a WRONG arm carrying four planted content errors as its positive control |
≥ 7 of 8 pairs equivalent and ≥ 3 of 4 planted errors caught by name | F1: P1, P5, P6 withheld |
| G2 | Mechanical operator integrity, local and free — the 63 assertions in arms.py |
all pass | run does not dispatch |
| G3 | Positive control on the consistency scale. SPELL — name spellings doubled, nothing else — must cost consistency |
DOM − SPELL > 0, sign test over 12 units, P ≤ 0.05 |
F1: P1, P5, P6 withheld. A scale that does not notice Ivan/Iwan is not measuring internal coherence |
| G4 | Cell completeness | ≥ 90% of cells in a stage parsed | F2: that stage reported with n and no P value |
| G5 | Not one scale. Pearson r between consistency and voice over all arm × unit cells |
if r ≥ +0.90 the double dissociation is not established whatever P2 and P4 return | F3 |
F1, F2, F3 are registered failure criteria and are not movable after the fact. G2 passed before this file was frozen; the other four are decided by the run.
7. What this design cannot establish
- Tier D is NOT PASSED. No jury verdict here carries evidential weight; every figure is
provisionalandinternal-judgment-only(charter §2.4, §5). - The population is three model seats, not readers and not translators.
- One work, one language pair, one class, one translator. A
consistencyeffect found here is an effect on this class in this passage. GENand the lexical arms are not the same kind of operator — one rewrites narration, the others substitute at 54 sites — so P2 and P4 are two separate one-directional results that together make a double dissociation, not one symmetrical measurement.DRIFTis an artificial arm. No translator produces it; the observed pattern isMIXED, which is why P5 and P6 are registered alongside P1 and why the result page must not report P1 as though it priced a real translational habit.
8. Budget
Declared ceiling $1.20. UTC day 2026-08-09 stood at $3.379187 of $5.00 before this session,
so the ceiling fits the $1.620813 headroom with $0.42 to spare. Pre-flight worst case is built
from max_tokens, not from expected output (note (abc)).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| critic | 1 | 16,000 | $0.10 |
V |
18 | 4,000 | $0.34 |
C |
18 | 4,000 | $0.34 |
N |
9 | 3,000 | $0.13 |
A |
9 | 3,000 | $0.13 |
| G1 parity | 4 | 6,000 | $0.14 |
| 59 | $1.18 |
Key usage is snapshotted before and after; per-response usage.include costs are primary and the
snapshot delta is the cross-check. A stage that will not fit the remaining headroom is dropped whole
and the drop is reported, in the order A, N — never V, C or the gates.
9. Amendment record — the independent pre-run critic, 2026-08-09, before any jury call
One adversarial pass, openai/gpt-5.6-terra, $0.05209325, over the frozen §§1–8 and the full text
of every arm in two blocks. Verdict NEEDS-REDESIGN, 12 findings, 9 of them BLOCKING.
Raw at runs/critic.txt.
Nothing was overruled outright. Nine findings are accepted in full, three in part — in each of the three the critic's own remedy was a larger instrument than the day's budget buys, and what was adopted instead is stated below with the residue left standing as a limit.
9.1 What changed, finding by finding
| # | critic's finding | disposition |
|---|---|---|
| 1 | The linear-interpolation null is asserted, not justified; any convex response to the proportion of foreign tokens manufactures the registered dip with no drift effect at all, and it can do so on one sense and not another, so P2 is not immune either | ACCEPTED IN FULL. The interpolation is gone from the inference. A new arm MATCH is built: item-uniform, and carrying exactly DRIFT's number of foreign tokens in every block (12/16, 13/19, 14/19), drawn from as nearly as possible the same items (residual per-item L1 = 4 / 6 / 6). The primary is now the direct contrast DRIFT vs MATCH |
| 2 | DRIFT's odd/even rule makes form a deterministic function of mention number — every first mention foreign, every later one domestic — so a penalty could be a first-mention effect |
ACCEPTED IN SUBSTANCE. The rule is replaced by a seeded random assignment under three constraints: both forms per repeated item; per-item COPY count unchanged, so composition is fixed by construction; and first mentions deliberately split (in every block at least one repeated item opens foreign and at least one opens domestic). Partial: the critic also wanted several independent drift realizations. One is run. The residue is a named limit |
| 3 | block × seat × ordering is not 18 independent units; the reverse-order replicate is a repeated measurement of one seat/block, so the registered P-values are fictitious |
ACCEPTED IN FULL. The second ordering is abolished. The unit is block × seat — 12 units — and what the ordering cost is spent on a fourth seat, J4 x-ai/grok-4.5, a fourth lab. Zero differences are dropped from the sign test and n reduced, and both n and the number dropped are reported with every P |
| 4 | Showing a seat all five variants of one passage side by side makes the manipulation obvious; this measures spot-the-odd-one-out, not scoring a translation | ACCEPTED IN FULL. One arm per call. No seat is told that any other rendering exists, and no seat sees the same block twice on the same scale |
| 5 | Subtracting raw 1–7 scores across two separately elicited ordinal scales assumes a commensurability the design itself denies; the SD-standardised secondary does not repair it | ACCEPTED IN SUBSTANCE. All cross-sense tests are re-specified as scale-free: per unit, each sense's MATCH → DRIFT movement is reduced to its sign, and P2 is an exact McNemar test on the discordant units. Raw magnitudes are reported as descriptive only. Partial: the critic wanted a calibrated latent-variable model; that is not affordable here and the limit is named |
| 6 | GEN is a bundled rewrite that also moves propositions, focalization and register — it seemed → it seemed to him, he felt → he realised, taken up with → the production of |
ACCEPTED. The three sites named are repaired (it seemed, he felt at once, the making of living fire) and DOM-vs-GEN is a gated parity pair in two blocks. The bundling itself stands and is not defended as a single-property operator: §7 already said GEN and the lexical arms are not the same kind of operator, and P3/P4 are reported as a claim about a bundled genericising rewrite |
| 7 | G1/G3 gate the wrong claims. SPELL failing would not invalidate the drift finding and SPELL passing would not validate it; and F1 withheld P1 but left P2 reportable after a failure that supposedly undermines it |
ACCEPTED IN FULL. G3 is abolished as a gate and SPELL is demoted to a reported manipulation check (P6). F1 now withholds P1 and P2 and P5 together — every claim resting on the lexical arms falls with the parity gate or none does |
| 8 | G5's pooled r ≥ 0.90 cutoff is arbitrary, pools incomparable things, and can veto or permit the conclusion for irrelevant reasons | ACCEPTED IN FULL. G5 is abolished as a gate; the correlation is reported descriptively with no threshold attached |
| 9 | The "one lexical axis" claim in §3 is false — COPY and DOM forms differ in familiarity, specificity, morphology and length, and frequent items are not interchangeable doses | ACCEPTED. The claim is withdrawn. It was load-bearing only for the interpolation null, which finding 1 already removed; MATCH needs no such axis, only equal token counts and near-equal item profiles, both of which are measured rather than assumed |
| 10 | P7's 0.50 threshold is an escape hatch, not a decision rule | ACCEPTED — P7 is dropped, with its stages. See §9.3 |
| 11 | Two "primary" claims, six tests, no family definition or error-control plan | ACCEPTED. The confirmatory family is {P1, P2} for the first half and {P3, P4} for the second, each conjunctive — both members must hold or the half is not established, which needs no multiplicity correction because a conjunction cannot inflate α. Everything else is descriptive and labelled so |
| 12 | SPELL is too narrow and too brittle to bear a gate's weight |
ACCEPTED — subsumed by finding 7's demotion |
9.2 The design that ran
Arms. DOM · MIXED · MATCH · DRIFT · GEN, plus SPELL in BA/BB for stage C only.
COPY is built and asserted but is not scored — it existed to anchor the dead interpolation, and
128 calls is what the budget buys. Its absence is a loss of the all-foreign reference point and is
recorded as such.
| arm | composition (fraction of class tokens foreign) | uniformity |
|---|---|---|
DOM |
0.000 | class-uniform |
MATCH |
0.750 / 0.684 / 0.737 by block | item-uniform |
DRIFT |
0.750 / 0.684 / 0.737 by block — identical | item-NON-uniform |
MIXED |
0.875 / 0.789 / 0.842 | item-uniform, class-non-uniform |
GEN |
0.000, class strings byte-identical to DOM |
item-uniform |
Stages. Two, one sense each, one arm per call: V (voice, Ukrainian shown) 60 calls;
C (consistency, English alone) 68 calls. 128 jury calls, dispatched in one frozen shuffled
order across stages, blocks, seats and arms.
Seats. J1 openai/gpt-5.6-terra · J2 google/gemini-3.6-flash · J3
deepseek/deepseek-v4-pro · J4 x-ai/grok-4.5. Unit = block × seat = 12.
9.3 Predictions as they now stand
| # | prediction | statistic | bar |
|---|---|---|---|
| P1 | confirmatory. Item-level drift costs consistency against a composition-matched item-uniform arm |
consistency: MATCH > DRIFT, sign test over 12 units |
one-sided P ≤ 0.05 (≥ 10 of 12 with no ties) |
| P2 | confirmatory — the parenthesis. The same manipulation does not do this to voice |
exact McNemar on units discordant in the sign of MATCH − DRIFT between the two senses |
one-sided P ≤ 0.05 |
| P3 | confirmatory. A genericising rewrite costs voice |
voice: DOM > GEN, sign test over 12 units |
one-sided P ≤ 0.05 |
| P4 | confirmatory — the reverse half. It does not do this to consistency |
exact McNemar on units discordant in the sign of DOM − GEN between the two senses |
one-sided P ≤ 0.05 |
| P5 | descriptive. Local flatness of the composition response over the item-uniform arms | consistency and voice: MIXED vs MATCH (0.833 vs 0.722 foreign) |
reported, no bar |
| P6 | descriptive. Manipulation check on the consistency scale |
consistency: DOM vs SPELL, 8 units (BA,BB × 4 seats) |
reported, no bar |
The first half of the parenthesis is established only if P1 and P2 both hold; the second only if P3 and P4 both hold. Either half may hold without the other, and the double dissociation needs all four.
Registered handling of ties: a unit whose two scores are equal contributes no sign and is dropped from that test, with n reduced and the count of dropped units reported beside every P value.
9.4 Gates as they now stand
| id | gate | bar | if it fails |
|---|---|---|---|
| G1 | Parity, adjudicator moonshotai/kimi-k3 (P4, not a seat), Ukrainian shown: eight registered pairs — MATCH/DRIFT in all three blocks, DOM/GEN in two, plus MIXED/MATCH, DOM/MIXED, DOM/MATCH — with a WRONG arm carrying four planted content errors as its positive control |
≥ 7 of 8 equivalent and ≥ 3 of 4 planted errors named | F1: P1, P2 and P5 all withheld together |
| G2 | Mechanical integrity — 72 assertions in arms.py, including MIXED byte-identical to the frozen base, the five generated arms identical outside the 54 sites, MATCH composition equal to DRIFT's, and DRIFT's first mentions split |
all pass | does not dispatch — passed before dispatch |
| G4 | Cell completeness per stage | ≥ 90% parsed | F2: that stage reported with n and no P value |
9.5 What the amendment does not fix, and is a limit on the result
- One drift realization, not several (finding 2, partial). A
DRIFTeffect could still be an artefact of this particular assignment of forms to positions. - No calibrated cross-sense metric (finding 5, partial). The sign-based McNemar is scale-free, but two senses can differ in how readily any manipulation moves them for reasons that are not about what they measure, and nothing here separates that from a genuine dissociation.
GENremains a bundle (finding 6). P3 and P4 price a genericising rewrite, not a property.COPYis unscored and the composition axis is anchored only at 0.000 and ~0.72–0.83.- Stages
NandAare gone, so this run has nonaturalnessoraccuracyreading at all and cannot say whether the manipulation moved them. The specificity claim rests entirely onvoiceas the comparison sense. - Four model seats. Not readers, not translators.
- Tier D is NOT PASSED. No figure here carries evidential weight.
9.6 Budget after amendment
128 jury calls (was 54) at a smaller per-call size, plus 4 parity calls, plus the critic already
spent. Worst case rebuilt from max_tokens: $1.15, against the ceiling of $1.20 declared in
§8 and headroom of $1.568813 after the critic. Stage C is dispatched before stage V; if
headroom runs out, V is truncated seat-last and the truncation is reported, since C carries P1
and V is needed only for P2.