Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260816g-device-function/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260816g-device-function
statusfrozen
created2026-08-16
updated2026-08-16
sensesstyle-correspondence, perceived-source-carriage, accuracy
provisionaltrue
linkswiki/arms/ARM-device-function.md, workshop/translations/zhongli/R04-v1/translation.md, framework/v0.2/README.md, wiki/findings/results/RS-20260816f-night-seam.md, wiki/method-notes.md, config/models.md, config/budget.md

E-20260816g — does a translator's account of what a device is doing predict what a reader loses when it goes?

Frozen before dispatch. ARM-device-function step 1 (T5).

1. Question

framework/v0.2 §7's one actionable instruction is a visibility diagnostic: count the marks your English contains, count the ones the source licensed, and the difference is what you invented. Two results say visibility is not function. RS-20260816f: a device that is wholly redundant for detection (56 of 56 boundaries found without it) is load-bearing for placement (7/7 against 1/7, P = 0.00233). Note (bpu): at a depictive device the two markings do not assert the same thing, so function cannot be reached by asking which rendering is better marked.

Both of those results reached function the same way — subtractively: write the passage without the device and put both to a reader. Whether that is worth telling a practitioner to do depends on one prior fact, which the project has never measured:

Does the translator already know? Does a translator's own declared account of what a device is for predict which reader-side quantity moves when the device is removed?

On the two occasions the project has checked in passing, the answer was no (RS-20260816f §6, and §7.14's withdrawal of §7.12 item 3). This design asks it on purpose, with the account registered in advance.

2. Materials

T-zhongli-R04-v1 — 蒲松齡〈種梨〉 ("Planting Pears"), Liaozhai zhiyi juan 1, 575 Chinese characters, translated whole by the lead under R04 and frozen at 7ace478, with the R06 draft frozen first at aaf5697. 686 English words. Contamination against Giles 1880 measured after both freezes: 0 / 0 / 0 shared 7-, 12-, 15-grams, longest common run 6 tokens, clean.

The declaration is §Device declaration of that file, committed before this design existed. It assigns each of fourteen stretches exactly one job from a closed vocabulary — MANNER, STANCE, ORNAMENT, NONE — or marks it SHAM. Three loci per job, three NONE, two SHAM.

The two arms are built by build_arms.py from the frozen translation string:

Markers are numbered M01–M14 in text order; the declaration ids L01–L14 are in declaration order and appear nowhere a seat can see. The mapping is materials/map.json.

The plain alternatives were written under a rule fixed in the declaration before the labels could influence it: replace the marked stretch with the plainest English rendering of the same events I can write; change nothing outside the stretch; add and remove no event. Where the device is depictive the plain rendering necessarily commits to less — note (bpu) — and that is a property of the material. This design does not assume a propositional constant and does not need one: it measures which reader-side quantity moves, not whether two renderings are equally good.

3. Procedure

Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5, QR qwen/qwen3.7-max (P4 and P5 are out, notes (bps), (bne); GL is out on long prompts).

Each body receives one arm only, blind: no mention of a second version, of the Chinese, of the translator, or of any hypothesis. For each of the fourteen markers it answers three yes/no probes:

probe wording put to the seat
manner Does the marked stretch tell you anything about how something was done, or what it looked or sounded like, beyond the bare fact that it happened?
stance Does the marked stretch convey an attitude — the narrator's, or a character's — toward what it describes?
ornament Does the marked stretch make you suppose the original had a figure of speech, a set phrase, or sound-play at this point?

and one whole-body item: list any marker that is ungrammatical or does not read as English.

Stage 0 — an independent check on the plain alternatives, bought before anything else (the critic's BLOCKING 2 and MAJOR 5 and 8). Two seats that score nothing in this run (P2, P3) are shown, for each of the twelve altered loci, the Chinese stretch and the two English renderings labelled A and B in an order fixed by the locus index, and are asked (i) whether both report the same events, and (ii) whether either is not English. (ii) is a gate: any locus called not-English by both seats is rewritten and the stage re-run before dispatch. (i) is a measurement and gates nothing, per notes (bpu) and (bkr) — a content-parity gate at device-marked sites fails at those sites for every hand, including published ones, and this project has lost two runs to exactly that. Its numbers go to the result page's limits.

Stage 1 (probe, note (bps)) — replicate 1, both arms, all four seats: 8 bodies. finish_reason and parse are read before anything further is bought; a seat that fails is dropped and recorded. Stage 2 — replicates 2 and 3, both arms, live seats: 16 bodies. Total 24.

max_tokens 2000, reasoning cap 900 — strictly less, per note (bpv). temperature 0.8, not 0 (the critic's BLOCKING 3): at temperature 0 a second call to the same seat is a copy of the first, and three such calls would be pseudoreplicates counted as three readers. Replicates are genuine samples and the seat, not the body, is the unit of the estimator (§4). Strict JSON; an unparseable body is dead, reported dead, counted in nothing, never re-rolled.

No fourth replicate is bought under any circumstance. If a seat loses more than one body in an arm, the seat is reported with the bodies it has and the estimator's per-seat mean is taken over them; a seat with zero usable bodies in either arm is dropped from the estimator entirely and named.

3a. Frozen parse rule

Strip a leading ``` fence, take the outermost { … }, json.loads. Expected shape {"markers": {"M01": {"manner": true, "stance": false, "ornament": false, "why": "…"}, …}, "unenglish": ["M07"]}. A body missing any of the fourteen markers, or with a non-boolean in any of the three probe fields, is dead.

4. The statistic, and what it is and is not an estimate of

Seat-clustered, per the critic's BLOCKING 3. For seat s, locus i, probe q, let y_s(i,q) be the seat's mean yes-rate over its usable bodies in an arm. Then

Δ(i,q) = mean over seats of [ y_s^FULL(i,q) − y_s^FLAT(i,q) ]

so a seat contributes once however many bodies it supplied. For the nine job loci, d(i) is the declared job's probe, and the primary statistic is

T = mean over the nine job loci of [ Δ(i, d(i)) − mean of Δ(i,q) over the other two probes ]

The estimand, narrowed on the record (the critic's BLOCKING 1 and 2, MAJOR 5 and 9). The translator wrote the translation, chose the loci, assigned the labels and wrote the plain alternatives. That dependency cannot be removed inside this session's budget and it is not pretended away. What it means is that T does not estimate "the translator knows what the device does". The translator controls that something is removed at each locus. He does not control which of three properties blind readers report losing, and that is what T reads. So the estimand is:

When this translator removes what he says a device is doing, does blind readers' reported loss land on the property he named, rather than on one of the other two?

The permutation null is correspondingly narrow: the declared label at a locus is unrelated to which probe moves there. It is a valid randomisation of the label-to-locus assignment and nothing more.

The two directions license very different things, and this asymmetry is registered. A hold licenses almost nothing — a translator's labels and the probes share a vocabulary, and the author wrote both sides of the contrast. A failure is the informative outcome: it says that even with the author writing the subtraction in his own favour, readers do not lose the property he named. That is the outcome §7 would be able to act on, and it is the one this design is powered to see.

What the rewrites remove, stated rather than assumed (MAJOR 5). Several plain alternatives take more than a decoration with them: L04 drops a form of address, L07 drops a number as well as a hyperbole, L09 drops a sound and a repetition, L03 drops an explicit pace. The estimand above is therefore the total effect of these twelve particular rewrites, not of a purified device. Stage 0 measures how far event-preservation actually holds and the result page reports it.

5. Predictions, registered

6. Failure criteria — what withholds what

7. What each outcome licenses

Narrowed after the critic's MAJOR 9. Every line below is about this translator, this tale, these twelve rewrites and these four model seats, and will be written that way.

Nothing here licenses a sentence about human readers. The seats are language models, Tier D is NOT PASSED, and every evaluative sentence carries internal-judgment-only.

8. Budget

Worst case built from the caps the requests permit (note abc), not from expected output.

stage calls caps in/out worst
pre-run critic (P1) 1 ~6,000 / 4,000 $0.021874 actual
stage 0 — plain-alternative check (P2, P3) 2 ~1,500 / 3,000 $0.034
stages 1–2 — 3 replicates × 4 seats × 2 arms 24 ~2,000 / 2,000 $0.305
total 27 $0.361

Worst case $0.361 < stop-loss $0.39 < declared ceiling $0.42. The stop-loss is enforced inside run.py, not by intention (note (bpq)). The UTC day opened for this session at $4.135637 of $5.00, leaving $0.864363.

9. Pre-run critic

An independent adversarial critic (P1, a seat that judges nothing in this run) received design v1, both arms and the frozen declaration, and returned NEEDS REDESIGN, 9 findings, 4 BLOCKING (critic-findings.json, $0.021874). v1 was never dispatched. Seven findings taken in full or in substance, two refused in part on the record.

# sev what it said what was done
1 BLOCKING the permutation has no randomisation basis: the labels and the probes come from the same author and the same semantic distinctions, so a positive T may show only that they resemble each other taken in substance. The estimand is narrowed in §4 and the confirmatory claim is withdrawn: a hold now licenses nothing, and the design is declared to be powered for the failure. The P value is kept, with its null restated as the narrow one it is. Refused in part: the remedy (independent blinded labels, held-out loci) is a different and larger experiment and is named as step 2's option, not bought here.
2 BLOCKING one hand wrote the translation, the loci, the labels and the plain alternatives; "plainest rendering" leaves latitude to remove exactly the labelled property taken in substance, same §4 narrowing: the author controls that something is removed, not which of three properties readers report losing. Plus a purchase: stage 0 buys an independent check that the alternatives are English and a measurement of event-preservation. Refused in part: an independently authored alternative set needs a second translator seat and a source-language adjudicator; out of budget, recorded as the first thing step 2 should buy.
3 BLOCKING temperature-0 replicates are pseudoreplicates; the effective n is the four models taken in full. temperature 0.8, three replicates, and the seat is the unit of the estimator (§4).
4 BLOCKING the three probes may all fire together on any vivid stretch; and the ornament probe asks a blind reader about the original first half taken in full: F4 is widened to a separation criterion that catches co-movement, ceilings and floors, not just an all-zero table. Second half refused on the record: what a reader is made to suppose about the original is not a defect of the probe, it is the quantity framework/v0.2 §7.14 is about — a supplied device tells your reader the original was doing something, whether or not it was. The wording is changed to name the inference honestly (make you suppose) rather than removed.
5 MAJOR several rewrites remove address, number, sound or pace as well as the device taken in full: §4 restates the estimand as the total effect of these twelve rewrites and names what each of the four takes with it.
6 MAJOR the sham floor is not a per-locus floor and "0.25 = two bodies of eight" is wrong for unpaired arms taken in full: P3 is now a measured floor with no absolute threshold, and F1 is relative to it. The arithmetic error is corrected in the open.
7 MAJOR a null cannot be distinguished from low power or saturation taken in substance: F4 widened, per-locus tables published in every case, and a registered sentence that a null will not be written as equivalence.
8 MAJOR F2 is not a sufficient check and the design does not say how the tests are recomputed when a locus is withheld taken in full: the recomputation rule is fixed in F2, with a cap of two withheld job loci, plus the pre-dispatch English gate in stage 0.
9 MAJOR the outcome licenses exceed what one translator, one story and four evaluator models can support taken in full: §7 rewritten, every line scoped to these materials.