Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260810t-drift-source-present/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260810t-drift-source-present
statusfrozen
created2026-08-10
updated2026-08-10
sensesconsistency
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-terminology-drift.md, wiki/findings/results/RS-20260809i-terminology-drift.md, workshop/experiments/E-20260809i-terminology-drift/design.md, workshop/translations/tini-budz/R04-v1/translation.md, workshop/translations/tini-polonyna/R04-v1/translation.md, wiki/goodness-senses.md

E-20260810t-drift-source-present — is the zero a fact about the sense, or about a source-blind reader?

ARM-terminology-drift step 2 of 2, T2. Frozen before dispatch; the pre-run critic pass is §9 and every amendment it forces is recorded there with the wording it replaced.

1. The question

RS-20260809i (S147) found that a jury scoring consistency on the English alone rates a passage that renders the same realia term two different ways exactly as coherent as a composition-matched passage that renders every term one way: mean difference 0.000, nine of twelve units tied, P = 0.875, against a comparator carrying the identical number of foreign tokens. The same seats, the same scale and the same passages dropped 2.000 points when one proper name was spelled two ways.

That result has a rival reading of equal standing, and the result page registered it as the run's sharpest limit (§5.1). voice was scored with the Ukrainian on the page and consistency was not, because consistency is defined as target-internal. So the finding may be about the sense — terminology drift is not a coherence event — or about the reader — a reader of the English alone has no way to know that polonyna and the high pasture answer to one Ukrainian word, and drift is visible as drift only against a source.

This design separates them. The same operator, the same scale, the same rubric string, the same four seats, in two conditions that differ in exactly one thing: whether the Ukrainian is on the page.

Subject-rule statement (continue-prompt §4.5). What this unit teaches about evaluating translations: whether internal coherence is a property a reader can assess without the source at all, or whether the sense the typology calls target-internal is in fact a source-relative sense wearing a target-internal definition. That is a claim about a sense of good, not about the project's apparatus.

2. Materials

Two units of the same novella, six blocks, 123 marked class sites.

unit blocks words class sites status
tini-polonyna (S147) BA BB BC 717 Ukr 54 arms imported unchanged from E-20260809i/materials/arms.json
tini-budz (this session) BD BE BF 769 Ukr 69 built here; T-tini-budz-R04-v1, frozen at 2a484ab

tini-budz is Kotsiubynsky's milking and cheese-making sequence, four pages further on in the same novella — the same institution, the same class of culture-bound vocabulary, translated in this session by the lead under R04 with its R06 draft frozen first. Provenance, the one elided verse passage, the four editorial glosses that were read, and the within-block repetition table are on workshop/translations/tini-budz/README.md. It doubles the units of analysis from 12 to 24 and supplies a second, independently translated drift realization — RS-20260809i §5.3 named the single realization as a limit and could not afford a second.

Arms. Generated mechanically by materials/arms.py (57 assertions, all passing before dispatch) from materials/template.txt, which is itself built by substitution on the frozen R04 text (materials/build_template.py) rather than retyped:

arm what it is uniformity
DOM every class item in its English form class-uniform, 0.000 foreign
MATCH item-uniform; COPY-token count equal to DRIFT's in every block item-uniform
DRIFT within a block, every item occurring twice or more takes both forms item-NON-uniform
FALSE MATCH plus one planted propositional error per block; no drift item-uniform, and factually wrong
SPELL DOM with one proper name spelled two ways, nothing else changed the positive control

MATCH is the whole design, and it is E-20260809i's pre-run critic's BLOCKING 1: it carries the same number of foreign tokens as DRIFT in every block, from as nearly as possible the same items, and differs from it only in whether the variation falls between items or within them. Residual per-item L1 on the new unit: 10 / 6 / 8 tokens.

Composition on tini-budz, fraction of class tokens in COPY form:

block sites COPY MIXED MATCH DRIFT DOM
BD 26 1.000 0.692 0.615 0.615 0.000
BE 20 1.000 1.000 0.700 0.700 0.000
BF 23 1.000 1.000 0.739 0.739 0.000

Fifteen items drift within a block (against S147's smaller set): BD diynytsia ×6, vivchar ×5, strunka ×4, kozar ×3, zahoroda ×3, honinnyk ×2; BE vatah ×5, staya ×2, berbenytsia ×2, budz ×2, zhentytsia ×2; BF vatah ×7, spuzar ×3, budz ×3, staya ×2.

Two declared asymmetries between the units.

  1. MIXED is class-uniform in BE and BF. On tini-budz the base translation domesticated only two class items (вівчар → shepherd, загорода → pen) and both occur only in BD. So the class-non-uniform exhibit this project has been reading since S012 lives, on this unit, in one block of three. tini-polonyna's MIXED (0.875 / 0.789 / 0.842) remains the exhibit; MIXED here is reported and carries no registered test either way.
  2. SPELL exists for BA, BB and BD only. BC has no proper noun occurring twice (S147's finding); BE's single name occurs once and BF contains none. The control therefore runs on 3 blocks × 4 seats = 12 units per condition, against 24 for the primary.

3. The two conditions, and what separates them

SB — source-blind SP — source-present
header "you are not being shown the original" "You will be shown the Ukrainian passage, and then one complete English rendering of it"
body the English arm alone the Ukrainian block, then the English arm
rubric string CONS_DEF, imported verbatim from E-20260809i/run.py and asserted equal to it identical
the one changed clause "…and you have no source text to check it against: you are being asked only whether the passage hangs together with itself." "…and you are NOT being asked whether it is a faithful rendering of the Ukrainian, which is given for reference only: you are being asked only whether the passage hangs together with itself."

The manipulation is information, not question. Both conditions ask the same target-internal question and both explicitly refuse the accuracy reading. What changes is whether the reader can know that two English strings answer to one Ukrainian word. If the sense is genuinely target-internal, that knowledge is irrelevant and the two conditions agree.

4. Procedure

5. Registered predictions and bars

All tests are exact sign tests by enumeration, one-sided where a direction is registered, ties excluded from the test and reported. Nothing below may be altered after any body is read.

Let, for unit u = (block, seat): d_SB(u) = consistency(MATCH,SB,u) − consistency(DRIFT,SB,u), d_SP(u) = consistency(MATCH,SP,u) − consistency(DRIFT,SP,u).

# test statistic bar what it decides
P1a THE PRIMARY, leg 1. Δ(u) = d_SP(u) − d_SB(u) > 0 sign test, 24 units, one-sided P ≤ 0.05 drift is penalised more when the source is visible
P1b THE PRIMARY, leg 2 (critic BLOCKING 1). McNemar on drift-responsiveness: units with d_SP > 0 ≥ d_SB against units with d_SB > 0 ≥ d_SP exact binomial on the discordant units, one-sided P ≤ 0.05 the interaction is not just P2 wearing a difference
P1 is ESTABLISHED only if BOTH legs clear 0.05.
P2 the simple effect with the source present: d_SP > 0 sign test, 24 units, one-sided P ≤ 0.05 consistency sees drift at all, given the source
P3 the S147 null, re-measured: d_SB sign test, 24 units, two-sided, plus the mean no verdict word. Reported as: mean, signs, exact P, and whether \|mean\| falls inside 25% of the SB SPELL effect (critic BLOCKING 2) whether S147's exact zero is a property of the instrument or of that run
C1 THE GATE — scale room, in EACH condition separately: DOM > SPELL sign test, 12 units, one-sided P ≤ 0.05 in both whether a null in that condition means anything
C2 descriptive. d_SB and d_SP split by unit (polonyna / budz) means and signs — whether the new passage behaves like the old
C3 descriptive. SB means on BA BB BC against S147's stored figures difference of means — cross-session test–retest on identical prompts
C4 THE MECHANISM CONTROL (critic BLOCKING 3). MATCH > FALSE, in each condition, and the interaction sign test, 24 units, one-sided P ≤ 0.05 whether the seats use the source for fidelity while told not to
D1 registered descriptive. per-seat Δ; finish_reason rate by condition; fractional per-item L1 tables — critic SERIOUS 6, MINOR 9, SERIOUS 7
D2 descriptive. the arm-mean table, per condition, per block means — the composition row S147 §2 made readable

Every P below indexes stochasticity within these four fixed model seats only (critic SERIOUS 6). None of them is an inference to a population of readers, and none may be reported as one.

Registered expectation, stated so it can be wrong. I expect P1 to fire and P3 to replicate — that is, the zero is about the reader. The alternative outcome that would most change the project's sense list is P1 and P2 both failing while C1 passes in both conditions: that would say terminology drift is not a coherence event for these readers even when they can see it is drift, and the parenthesis's first half would have to be struck rather than qualified.

6. Failure criteria and withholding rules

6a. G1 parity — run before dispatch, and it took two seats

Outcome: 4 of 4 planted errors caught, 5 of 6 real pairs equivalent. Nothing withheld.

The first seat failed the task, not the materials. mistralai/mistral-medium-3-5 returned non-equivalent for all ten pairs — the six real ones and the four planted ones alike — citing in every case exactly the lexical differences the prompt excludes ("A uses 'milking gate' vs B uses 'strunka'"). A gate that never returns equivalent has no discriminative power and cannot bar anything; its three bodies are kept in runs/discarded/ and ledgered at $0.030216. The seat was a poor choice on my part: config/models.md already records it as the panel's weakest and not panel material.

The second seat is moonshotai/kimi-k3, which passed this identical task at S147 (8 of 8, 4 of 4) and is not a jury seat. $0.047766.

pair verdict note
BD MATCH/DRIFT equivalent, empty difference list
BE MATCH/DRIFT equivalent, empty difference list
BF MATCH/DRIFT equivalent, empty difference list the primary contrast, verified in all three blocks
BD DOM/MIXED equivalent
BF DOM/MIXED equivalent
BE DOM/MATCH non-equivalent "A says fourteen head of sheep, B says fourteen drobyeta"
four planted errors all four caught, each named correctly black shaggy → fair clean-shaven; third day → third hour; fourteen → forty; sun goes down → sun comes up

Item-level inspection, as the amended F3 requires at 5 of 6. The single flag is on drobyeta, whose DOM form is head of sheep at the BE site («Мосійчук має штирнадцять дробєт»). That rendering is denotationally correct, and the "difference" the seat named is the foreign-versus-English difference the prompt explicitly excludes — the same error the first seat made ten times out of ten, made once here. The DOM string is upheld and nothing is withheld, and the flag is recorded rather than argued away. It touches no MATCH/DRIFT contrast in any case.

7. Pre-flight cost estimate

Worst case is built from the caps the requests actually permit, not from expected output (note (abc)).

stage calls cap worst-case line
pre-run critic 1 16,000 $0.10
G1 parity 3 8,000 $0.06
scoring 216 4,000 (8,000 on one retry each) $2.40
re-dispatch allowance ≤ 20 8,000 $0.30
declared ceiling $2.86

Expected, from S147's realised $0.0053/call over 128 calls with half of them source-bearing: ≈ $1.50. Today's UTC ledger stands at $0.867877 of $5.00 before this run, so the ceiling fits the headroom with $1.27 to spare.

8. Declared limits, written before the run

  1. The translator knew the design. tini-polonyna was translated before E-20260809i existed; tini-budz was translated in this session with RS-20260809i §8's successor design already written down. The consequence is bounded and stated: MIXED on the new unit is the lead's own class handling and carries no weight in any registered test (§5, D1). MATCH and DRIFT are generated from the marked sites and are invariant to which form the base chose at any site, so the primary is untouched by what the translator knew.
  2. Four model seats, not readers. Tier D is NOT PASSED; every figure is provisional and internal-judgment-only.
  3. One work, one class, one translator, two passages. The class is Hutsul pastoral realia and nothing here generalises past it without being shown to.
  4. Two drift realizations, not many. One seeded assignment per unit. Better than S147's one; not a distribution.
  5. SP may buy its effect with an accuracy halo. The prompt refuses the accuracy reading explicitly, and refusing it in words is not the same as preventing it. If P1 fires, the honest statement is the source's presence changes the score, and the mechanism — drift becoming visible, or accuracy leaking in — is not settled by this design. That is registered here so it cannot be quietly dropped from the result.
  6. C1's 12 units are half the primary's 24, so the gate is the weakest-powered test in the run and a C1 failure could itself be a power failure. If C1 fails, that reading is reported beside the withholding.

9. Pre-run critic — NEEDS-AMENDMENT, 12 findings, 3 BLOCKING

One independent adversarial pass over the frozen design and its predecessor result, moonshotai/kimi-k3 — not a jury seat and not S147's critic. 11 accepted, 1 accepted in part, none overruled. Two calls: the first returned finish_reason: length with zero content, the whole 16,000 cap spent on hidden reasoning ($0.266061, dead, ledgered); the re-dispatch with reasoning disabled cost $0.053280. Note (bkw)'s remedy, fifth occurrence.

The amendment that matters, and it changed what was dispatched: BLOCKING 3. §8.5 conceded that the source-present condition might buy its effect from an accuracy halo, and then §5 registered an interpretation — "the zero is about the reader" — that assumed it had not. Nothing in the frozen design separated the two. The remedy is the FALSE arm: MATCH, with no drift at all, plus one planted propositional error per block that contradicts the Ukrainian and is invisible from the English alone (materials/false_arm.py, 30 assertions; the six edits and the reason each is source-blind-invisible are tabled there). It is registered as C4. MIXED was dropped to pay for it — 48 cells out, 48 cells in, the total unchanged — and that is the right trade twice over, because the critic's MINOR 11 showed MIXED was weaker than §8.1 had admitted.

# severity what it attacked disposition
1 BLOCKING P1's Δ sign test degenerates to P2 when d_SB is near zero on most units, so the interaction is not separately tested accepted. P1 is now two legs, both required: P1a, the Δ sign test as frozen, and P1b, McNemar on drift-responsiveness — units responsive under SP only against units responsive under SB only. The interaction is established only if both clear 0.05
2 BLOCKING P3's replication rule (\|mean\| ≤ 0.25) is arbitrary and uncalibrated; a real 0.2-point cost would be declared a replication of zero accepted. The word replicated is struck. The band is now 25% of the SPELL effect measured in the same condition — data-anchored, by a rule fixed before dispatch — and the mean, the signs and the exact two-sided P are reported whatever it says
3 BLOCKING nothing separates drift became visible from accuracy leaked in; P1 firing would be read the first way with no evidence against the second accepted — the FALSE arm, above. If C4 fires under SP and not SB, the seats are demonstrably using the source for fidelity while told not to, and a P1 firing is open to the same reading; if C4 does not fire, the halo account is much weakened
4 SERIOUS SPELL covers BA BB BD; the gate does not gate BE BF, where a third of the primary's data sits accepted in part. The claim is restricted: C1 licenses interpretation of nulls on the blocks where SPELL exists, and the gap on BE/BF is declared. Not accepted: building a different control on BE/BF would make C1 a heterogeneous gate whose failure could not be attributed to anything — worse than a declared gap. Per-block DOM means are reported as ceiling evidence instead
5 SERIOUS C1 on 12 units needs ~10 of 12 non-tied one way; a tie-driven failure would withhold P1 even if the interaction is real accepted. Registered tie-contingent rule: if C1 fails and ≥ 8 of its 12 units are tied with DOM at ceiling (7), that is a saturation finding and P1 is reported with the caveat rather than withheld
6 SERIOUS four fixed model seats are not exchangeable replicates of a reader; the framing invites that misreading accepted. §5 now says every P indexes stochasticity within these four seats only, and a per-seat Δ table is registered so a one-seat-driven P1 is visible
7 SERIOUS residual per-item L1 of 10/6/8 means MATCH and DRIFT differ in which items carry foreign tokens, not only in within/between accepted, with one correction to the finding. The critic compared raw L1 against S147's 4/6/6 without dividing by site count; as fractions the two units are comparable (this unit 0.385 / 0.300 / 0.348; S147 0.211 / 0.316 / 0.375). The fractional table is now registered as a descriptive, and the honest statement of the contrast — matched on token count, not on which items carry them — is in §8
8 SERIOUS the design leans on S147's voice movement (0.583, P = 0.109), which missed its own bar accepted. P2 is registered as the first adequately-powered test of whether source-present consistency sees drift at all, standing on its own and not on S147's direction
9 MINOR SP prompts are ~2× longer, so any length-driven effect loads onto the condition accepted. finish_reason rates by condition are a registered reported check
10 MINOR F3's 4-of-6 bar can pass with a third of pairs non-equivalent accepted. Tightened: 6 of 6 for full release; 4–5 of 6 triggers item-level inspection and withholding limited to contrasts involving the failing DOM strings; < 4 of 6 blanket withholding as before
11 MINOR §8.1's invariance claim covers form choice but not site selection — the lead chose which items repeat, knowing repetition is what is tested accepted, and it is the sharpest of the twelve. §8.1 now says so, and it is part of why MIXED was the arm dropped
12 MINOR the retry allowance is uncapped in the procedure text accepted. Total re-dispatches capped at 20; beyond that the run stops and reports incomplete