Repository path: workshop/experiments/E-20260816g-device-function/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260816g-device-function |
| status | frozen |
| created | 2026-08-16 |
| updated | 2026-08-16 |
| senses | style-correspondence, perceived-source-carriage, accuracy |
| provisional | true |
| links | wiki/arms/ARM-device-function.md, workshop/translations/zhongli/R04-v1/translation.md, framework/v0.2/README.md, wiki/findings/results/RS-20260816f-night-seam.md, wiki/method-notes.md, config/models.md, config/budget.md |
E-20260816g — does a translator's account of what a device is doing predict what a reader loses when it goes?
Frozen before dispatch. ARM-device-function step 1 (T5).
1. Question
framework/v0.2 §7's one actionable instruction is a visibility diagnostic: count the marks your
English contains, count the ones the source licensed, and the difference is what you invented. Two
results say visibility is not function. RS-20260816f: a device that is wholly redundant for
detection (56 of 56 boundaries found without it) is load-bearing for placement (7/7 against 1/7,
P = 0.00233). Note (bpu): at a depictive device the two markings do not assert the same thing, so
function cannot be reached by asking which rendering is better marked.
Both of those results reached function the same way — subtractively: write the passage without the device and put both to a reader. Whether that is worth telling a practitioner to do depends on one prior fact, which the project has never measured:
Does the translator already know? Does a translator's own declared account of what a device is for predict which reader-side quantity moves when the device is removed?
On the two occasions the project has checked in passing, the answer was no (RS-20260816f §6, and
§7.14's withdrawal of §7.12 item 3). This design asks it on purpose, with the account registered in
advance.
2. Materials
T-zhongli-R04-v1 — 蒲松齡〈種梨〉 ("Planting Pears"), Liaozhai zhiyi juan 1, 575 Chinese
characters, translated whole by the lead under R04 and frozen at 7ace478, with the R06 draft frozen
first at aaf5697. 686 English words. Contamination against Giles 1880 measured after both
freezes: 0 / 0 / 0 shared 7-, 12-, 15-grams, longest common run 6 tokens, clean.
The declaration is §Device declaration of that file, committed before this design existed. It
assigns each of fourteen stretches exactly one job from a closed vocabulary — MANNER, STANCE,
ORNAMENT, NONE — or marks it SHAM. Three loci per job, three NONE, two SHAM.
The two arms are built by build_arms.py from the frozen translation string:
FULL— the translation verbatim, with the fourteen stretches wrapped in{{M01: …}}markers.FLAT— byte-identical toFULLexcept inside the twelve non-SHAMmarkers, where the stretch is replaced by the plain alternative inbuild_arms.py. The twoSHAMmarkers are identical in both arms.
Markers are numbered M01–M14 in text order; the declaration ids L01–L14 are in declaration
order and appear nowhere a seat can see. The mapping is materials/map.json.
The plain alternatives were written under a rule fixed in the declaration before the labels could influence it: replace the marked stretch with the plainest English rendering of the same events I can write; change nothing outside the stretch; add and remove no event. Where the device is depictive the plain rendering necessarily commits to less — note (bpu) — and that is a property of the material. This design does not assume a propositional constant and does not need one: it measures which reader-side quantity moves, not whether two renderings are equally good.
3. Procedure
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5,
QR qwen/qwen3.7-max (P4 and P5 are out, notes (bps), (bne); GL is out on long prompts).
Each body receives one arm only, blind: no mention of a second version, of the Chinese, of the translator, or of any hypothesis. For each of the fourteen markers it answers three yes/no probes:
| probe | wording put to the seat |
|---|---|
manner |
Does the marked stretch tell you anything about how something was done, or what it looked or sounded like, beyond the bare fact that it happened? |
stance |
Does the marked stretch convey an attitude — the narrator's, or a character's — toward what it describes? |
ornament |
Does the marked stretch make you suppose the original had a figure of speech, a set phrase, or sound-play at this point? |
and one whole-body item: list any marker that is ungrammatical or does not read as English.
Stage 0 — an independent check on the plain alternatives, bought before anything else (the
critic's BLOCKING 2 and MAJOR 5 and 8). Two seats that score nothing in this run (P2, P3) are
shown, for each of the twelve altered loci, the Chinese stretch and the two English renderings
labelled A and B in an order fixed by the locus index, and are asked (i) whether both report the
same events, and (ii) whether either is not English. (ii) is a gate: any locus called not-English
by both seats is rewritten and the stage re-run before dispatch. (i) is a measurement and gates
nothing, per notes (bpu) and (bkr) — a content-parity gate at device-marked sites fails at
those sites for every hand, including published ones, and this project has lost two runs to exactly
that. Its numbers go to the result page's limits.
Stage 1 (probe, note (bps)) — replicate 1, both arms, all four seats: 8 bodies. finish_reason
and parse are read before anything further is bought; a seat that fails is dropped and recorded.
Stage 2 — replicates 2 and 3, both arms, live seats: 16 bodies. Total 24.
max_tokens 2000, reasoning cap 900 — strictly less, per note (bpv). temperature 0.8, not
0 (the critic's BLOCKING 3): at temperature 0 a second call to the same seat is a copy of the
first, and three such calls would be pseudoreplicates counted as three readers. Replicates are
genuine samples and the seat, not the body, is the unit of the estimator (§4). Strict JSON; an
unparseable body is dead, reported dead, counted in nothing, never re-rolled.
No fourth replicate is bought under any circumstance. If a seat loses more than one body in an arm, the seat is reported with the bodies it has and the estimator's per-seat mean is taken over them; a seat with zero usable bodies in either arm is dropped from the estimator entirely and named.
3a. Frozen parse rule
Strip a leading ``` fence, take the outermost { … }, json.loads. Expected shape
{"markers": {"M01": {"manner": true, "stance": false, "ornament": false, "why": "…"}, …},
"unenglish": ["M07"]}. A body missing any of the fourteen markers, or with a non-boolean in any of
the three probe fields, is dead.
4. The statistic, and what it is and is not an estimate of
Seat-clustered, per the critic's BLOCKING 3. For seat s, locus i, probe q, let
y_s(i,q) be the seat's mean yes-rate over its usable bodies in an arm. Then
Δ(i,q) = mean over seats of [ y_s^FULL(i,q) − y_s^FLAT(i,q) ]
so a seat contributes once however many bodies it supplied. For the nine job loci, d(i) is the
declared job's probe, and the primary statistic is
T = mean over the nine job loci of [ Δ(i, d(i)) − mean of Δ(i,q) over the other two probes ]
The estimand, narrowed on the record (the critic's BLOCKING 1 and 2, MAJOR 5 and 9). The
translator wrote the translation, chose the loci, assigned the labels and wrote the plain
alternatives. That dependency cannot be removed inside this session's budget and it is not pretended
away. What it means is that T does not estimate "the translator knows what the device does".
The translator controls that something is removed at each locus. He does not control which
of three properties blind readers report losing, and that is what T reads. So the estimand is:
When this translator removes what he says a device is doing, does blind readers' reported loss land on the property he named, rather than on one of the other two?
The permutation null is correspondingly narrow: the declared label at a locus is unrelated to which probe moves there. It is a valid randomisation of the label-to-locus assignment and nothing more.
The two directions license very different things, and this asymmetry is registered. A hold licenses almost nothing — a translator's labels and the probes share a vocabulary, and the author wrote both sides of the contrast. A failure is the informative outcome: it says that even with the author writing the subtraction in his own favour, readers do not lose the property he named. That is the outcome §7 would be able to act on, and it is the one this design is powered to see.
What the rewrites remove, stated rather than assumed (MAJOR 5). Several plain alternatives take
more than a decoration with them: L04 drops a form of address, L07 drops a number as well as a
hyperbole, L09 drops a sound and a repetition, L03 drops an explicit pace. The estimand above is
therefore the total effect of these twelve particular rewrites, not of a purified device. Stage 0
measures how far event-preservation actually holds and the result page reports it.
5. Predictions, registered
P1(primary).T > 0, with exact one-sided permutation P ≤ 0.05. The null relabels the nine job loci with the observed multiset{MANNER×3, STANCE×3, ORNAMENT×3}— 1,680 assignments, enumerated exhaustively, no sampling.P2. The threeNONEloci move less than the nine job loci:mean over NONE of max_q |Δ(i,q)| < mean over job loci of max_q |Δ(i,q)|, with an exact permutation over which three of the twelve non-SHAMloci carry theNONElabel (C(12,3) = 220, enumerated), P ≤ 0.05.P3(the floor, restated after the critic's MAJOR 6). At the twoSHAMloci — identical text in both arms — the floor ismean over the two of max_q |Δ(i,q)|, computed with the same seat-clustered estimator. No absolute threshold is claimed; the floor is a measured quantity and it is whatP1andP2are read against. (The frozen v1 of this design said 0.25 "i.e. no more than two bodies of eight differ", which is arithmetically wrong for unpaired arms. Corrected before dispatch.)P4. Theornamentprobe fires more often onFULLthan onFLATatL07,L08,L09specifically (the three loci where the English is answering a named Chinese figure — 萬目攢視, 倏而花倏而實, 丁丁). Per-locus, directional, no P value claimed at n = 8 per cell.
6. Failure criteria — what withholds what
F1. If the sham floor (P3) is at or above the mean over the nine job loci ofmax_q |Δ(i,q)|, the instrument cannot resolve a locus andP1andP2are both withheld. The run then reports the floor and nothing else.F2. AFLATmarker flagged as not-English by both stage-0 seats is rewritten before dispatch. After dispatch, a marker flagged by three or more of the twenty-four bodies has everyΔat that locus withheld. How the tests are recomputed is fixed here (the critic's MAJOR 8):P1is recomputed on the surviving job loci with the permutation taken over the surviving label multiset, and if more than two job loci are withheldP1is withheld entirely;P2's permutation universe becomes the surviving non-SHAMloci. No other exclusion rule exists and none may be invented after the run.F3. Fewer than sixteen usable bodies, or fewer than two live seats in either arm: everything is withheld and the run is reported as a dead run.F4(widened after the critic's BLOCKING 4 and MAJOR 7). The three probes must be shown to move separately on this material, orTis measuring nothing. Letsep = mean over the nine job loci of [ max_q Δ(i,q) − min_q Δ(i,q) ]. Ifsepis at or below the sham floor — whether because nothing moved, because everything moved together, or because a ceiling or floor pinned the probes —P1is not interpretable as a null about the translator and is reported as an instrument outcome. The per-locus, per-probe table and the raw FULL and FLAT rates are published in every case, so that a null can be read for saturation rather than asserted as equivalence. A null here is not an equivalence claim and will not be written as one.
7. What each outcome licenses
Narrowed after the critic's MAJOR 9. Every line below is about this translator, this tale, these twelve rewrites and these four model seats, and will be written that way.
P1holds → on these materials the declared job is recoverable from the subtraction. This licenses no practitioner instruction on its own, for the reason set out in §4: the author wrote both sides. It is reported as consistent-with, and §7 gains a sentence saying the question is open.P1fails,F4not fired → on these materials the translator's account did not predict where the reader's loss landed, even though he wrote the subtraction. That is a bound worth having, and withRS-20260816f§6 and §7.14's withdrawal it becomes the third case in a row. §7 then gains: your reason for a device is not evidence about what it does — the only way to find out is to write the passage without it and look. Stated as a diagnostic, with its three instances named.P2fails whileP1holds → the translator can name a job but cannot tell when a choice is inert. Narrower, and still writable.F4fires → the probes do not separate on this material; the run reports an instrument outcome and §7 gains nothing. This is a live possibility and is not a failure of the arm.
Nothing here licenses a sentence about human readers. The seats are language models, Tier D is
NOT PASSED, and every evaluative sentence carries internal-judgment-only.
8. Budget
Worst case built from the caps the requests permit (note abc), not from expected output.
| stage | calls | caps in/out | worst |
|---|---|---|---|
pre-run critic (P1) |
1 | ~6,000 / 4,000 | $0.021874 actual |
stage 0 — plain-alternative check (P2, P3) |
2 | ~1,500 / 3,000 | $0.034 |
| stages 1–2 — 3 replicates × 4 seats × 2 arms | 24 | ~2,000 / 2,000 | $0.305 |
| total | 27 | $0.361 |
Worst case $0.361 < stop-loss $0.39 < declared ceiling $0.42. The stop-loss is enforced inside
run.py, not by intention (note (bpq)). The UTC day opened for this session at $4.135637 of
$5.00, leaving $0.864363.
9. Pre-run critic
An independent adversarial critic (P1, a seat that judges nothing in this run) received design v1,
both arms and the frozen declaration, and returned NEEDS REDESIGN, 9 findings, 4 BLOCKING
(critic-findings.json, $0.021874). v1 was never dispatched. Seven findings taken in full or in
substance, two refused in part on the record.
| # | sev | what it said | what was done |
|---|---|---|---|
| 1 | BLOCKING | the permutation has no randomisation basis: the labels and the probes come from the same author and the same semantic distinctions, so a positive T may show only that they resemble each other |
taken in substance. The estimand is narrowed in §4 and the confirmatory claim is withdrawn: a hold now licenses nothing, and the design is declared to be powered for the failure. The P value is kept, with its null restated as the narrow one it is. Refused in part: the remedy (independent blinded labels, held-out loci) is a different and larger experiment and is named as step 2's option, not bought here. |
| 2 | BLOCKING | one hand wrote the translation, the loci, the labels and the plain alternatives; "plainest rendering" leaves latitude to remove exactly the labelled property | taken in substance, same §4 narrowing: the author controls that something is removed, not which of three properties readers report losing. Plus a purchase: stage 0 buys an independent check that the alternatives are English and a measurement of event-preservation. Refused in part: an independently authored alternative set needs a second translator seat and a source-language adjudicator; out of budget, recorded as the first thing step 2 should buy. |
| 3 | BLOCKING | temperature-0 replicates are pseudoreplicates; the effective n is the four models | taken in full. temperature 0.8, three replicates, and the seat is the unit of the estimator (§4). |
| 4 | BLOCKING | the three probes may all fire together on any vivid stretch; and the ornament probe asks a blind reader about the original |
first half taken in full: F4 is widened to a separation criterion that catches co-movement, ceilings and floors, not just an all-zero table. Second half refused on the record: what a reader is made to suppose about the original is not a defect of the probe, it is the quantity framework/v0.2 §7.14 is about — a supplied device tells your reader the original was doing something, whether or not it was. The wording is changed to name the inference honestly (make you suppose) rather than removed. |
| 5 | MAJOR | several rewrites remove address, number, sound or pace as well as the device | taken in full: §4 restates the estimand as the total effect of these twelve rewrites and names what each of the four takes with it. |
| 6 | MAJOR | the sham floor is not a per-locus floor and "0.25 = two bodies of eight" is wrong for unpaired arms | taken in full: P3 is now a measured floor with no absolute threshold, and F1 is relative to it. The arithmetic error is corrected in the open. |
| 7 | MAJOR | a null cannot be distinguished from low power or saturation | taken in substance: F4 widened, per-locus tables published in every case, and a registered sentence that a null will not be written as equivalence. |
| 8 | MAJOR | F2 is not a sufficient check and the design does not say how the tests are recomputed when a locus is withheld |
taken in full: the recomputation rule is fixed in F2, with a cap of two withheld job loci, plus the pre-dispatch English gate in stage 0. |
| 9 | MAJOR | the outcome licenses exceed what one translator, one story and four evaluator models can support | taken in full: §7 rewritten, every line scoped to these materials. |