Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260810z-idiom-reach/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260810z-idiom-reach
statusfrozen
created2026-08-10
updated2026-08-10
sensesstyle-correspondence
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-idiom-reach.md, workshop/regimes/R24-source-first-located-low.md, workshop/regimes/R23-device-crossed-low.md, workshop/regimes/R22-placeless-low.md, workshop/translations/botchan/R24-v1/translation.md, workshop/translations/botchan/R22-v1/translation.md, workshop/experiments/E-20260810c-register-reach/design.md, wiki/findings/results/RS-20260810c-register-reach.md, wiki/findings/results/RS-20260809g-device-cross.md, framework/v0.2/README.md, config/models.md, config/budget.md, wiki/method-notes.md

E-20260810z — an unexercised permission and an absent device: telling them apart

Frozen 2026-08-10 (S155) before dispatch. ARM-idiom-reach step 1. The translation limb (T-botchan-R24-v1, its W5+ table and R24) was committed at 96ed0aa before this file existed; code.py was written and self-tested before it too; this file was committed at 95db9a2 before the pre-run critic saw it.

AMENDED after the pre-run critic, before any generation call. NEEDS-REDESIGN, 16 findings, 7 BLOCKING; 12 accepted, 4 accepted-in-part, 5 individual remedies overruled in writing. Amendments A1–A11 are in force and are marked in place below. The adjudication is critic.md; the pre-amendment text is git show 95db9a2:.../design.md.

What the critic changed, in one paragraph. The design compared a newly generated source-first +I arm against a −I arm generated at S150 in a different call, which conflates the permission with independent-pass variability (BLOCKING 1, 2, 11): a contemporaneous source-first −I arm Nsrc is added and the primary moves inside the resulting 2 × 2 (A1). The positive control was an arm instructed to maximise the measured outcome, which validates nothing (BLOCKING 5, 6, 14): it is replaced by the lead's own frozen R22 and R24 English at each site, a documented negative/positive pair (A4, A5). A loc answer was counted without checking the named marker was in the text (13): it now must be (A6).

1. Question

framework/v0.2 §7.2 refuses a register-carriage recommendation for the third time, and RS-20260810c §6 names the one thing that would move it:

is the located idiom unavailable in Japanese narration, or is it merely unreachable by minimal revision? The instrument that would separate them writes the +I arm from the source, not from the ∅, and pays the tie-rate cost that the minimal-pair procedure was adopted to avoid.

This is that instrument.

When two independent hands are given the located-idiom permission and write from the Japanese rather than revising their own placeless English, does the English they produce become placeable — and does it become more placeable than the same permission buys inside a revision?

(The second half is the critic's A1: the question is only answerable as an interaction, against a source-first −I arm dispatched in the same stage.)

2. What is carried from E-20260810c unchanged, and why that is the whole design

Everything except the one clause under test. The sites, the hands, the judges, the loc prompt, the item-batching rule and the respelling protocol are the previous run's, and three of the six model arms rated here are the previous run's frozen output, re-rated in this batch.

carried what
sites the same 30, runs/sites.json of E-20260810c, verbatim, ids JA-01–JA-30. No new census. The 0.874 admission rate and the narrowed estimand of RS-20260810c §3 carry with them
hands H1 = x-ai/grok-4.5, H2 = qwen/qwen3.7-max
judges L1 = nvidia/nemotron-3-ultra-550b-a55b, L2 = z-ai/glm-5.2 — the same two as S145 and S150, which is what makes loc(B) a reproduction rather than a new measurement
loc prompt LOC_HEAD verbatim, all three questions (err, add, loc)
batching rating is split by site, with all arms of a site kept together in one call (RS-20260806g: loc and REG are batch-sensitive, and a difference is only read within one call)
−S protocol code.py imports E-20260810c/code.py rather than copying it, so that "respelling" means here exactly what it meant there
note (bkr) no content-parity check gates anything. err and add are asked inside the reproduced prompt and reported descriptively
A21 where an arm's English at a site is byte-identical to that hand's ∅, that is recorded as a non-application and is not treated as a device

The one change is the procedure by which the +I arm is written. In E-20260810c the A cell was a minimal revision of the hand's own ∅. Here Asrc is the same clause, the same model, the same temperature and the same item format, written from the Japanese with no English in front of it.

3. What this run does NOT attempt, registered so a reader does not look for it

There is no REG stage and no register comparison. P1a — how much of English's downward movement each device buys — is not asked here and is not withheld here; it is simply not this run's question. §7.2's obstacle was the manipulation, and until a +I arm exists that actually differs from its baseline there is nothing to compare. This run is a reach run. If it passes, the register comparison is the successor and it needs a REG panel; if it fails, the register comparison has no arm to run on.

No new census, no new sites, no new criterion. Reusing a frozen list removes a degree of freedom the lead would otherwise have after having translated this span twice.

4. Arms

Six model arms per hand, plus two lead instrument controls — fourteen English renderings per site.

arm switches procedure provenance
∅ −I −S one pass, from the source, per site frozen, E-20260810c/runs/gen-{H}-null.parsed.json
A +I −S minimal revision of that hand's ∅ frozen, E-20260810c/runs/gen-{H}-A.parsed.json
B −I +S minimal revision of that hand's ∅ frozen, E-20260810c/runs/gen-{H}-B.parsed.json
Nsrc −I −S one pass, from the source, per site — A1, added by the critic new
Asrc +I −S one pass, from the source, per site — R24's procedure new
Lsrc +I −S, dose 2 one pass, from the source, preferring the placeable option wherever English offers one new
NLEAD −I −S the lead's T-botchan-R22-v1 English at this site — instrument control (A4) frozen, S150
PLEAD +I −S the lead's T-botchan-R24-v1 English at this site — instrument control (A4) frozen, 96ed0aa

A1 — the design is a 2 × 2 and the primary lives inside it.

−I +I
minimal revision ∅ (frozen) A (frozen)
source-first Nsrc Asrc

Nsrc and Asrc are dispatched in the same stage, to the same models, at the same temperature, over the same items, in the same format, and their prompt strings differ in exactly one clause — machine-checked before dispatch: the unified diff of the two is two lines and both are the clause. Nothing else about them differs, which is what loc(Asrc) − loc(Nsrc) therefore isolates.

A4 — NLEAD and PLEAD are instrument controls and enter G3b and nothing else. They are one rendering per site with no hand, drawn from two frozen lead translations of this exact span: NLEAD written under a rule that forbade located means, PLEAD under one that permitted them, with a per-site log naming the located item taken at 45 unprimed sites. The paragraph-to-site alignment is the lead's, made once, and code.assert_lead_sites() proves every span is a contiguous substring of the frozen translation — the boundaries are the lead's choice and not one word is. Three of the thirty spans are byte-identical across the pair (JA-11, JA-14, JA-27), which bounds G3b at 0.90 rather than 1.00 and is stated here rather than discovered.

Nsrc's prompt is E-20260810c's BASE + CLAUSE["null"] + TAIL, and Asrc's is the same with CLAUSE["A"], both with the REVISE block absent. The REVISE block is what makes A a revision of ∅; its absence is what makes Asrc a translation.

Lsrc's prompt is Asrc's with one sentence appended to the clause, printed here in full so the dose is on the record:

At every site where such an option exists, PREFER it: where English offers both a placeless wording and one a reader would place, choose the one a reader would place.

Lsrc is an instrument dose, not a translation regime, and no artifact is filed for it. Its role is narrowed by A6: it measures prompt-following reach — how far a hand told to prefer placeable English moves the placement rate — and it is not offered as a measurement of what the material affords. The critic's BLOCKING 5 is right that an arm instructed to maximise the measured outcome cannot validate the instrument, and A4's PLEAD/NLEAD pair does that job instead.

5. Roles, and no model does double duty

role slug also used as
pre-run critic openai/gpt-5.6-terra nothing else in this run
hand H1 x-ai/grok-4.5 nothing else
hand H2 qwen/qwen3.7-max nothing else
judge L1 nvidia/nemotron-3-ultra-550b-a55b nothing else
judge L2 z-ai/glm-5.2 nothing else

The lead is not in the run. T-botchan-R24-v1 is rated by nobody, enters no figure, and is reported only as a translator's-chair record beside the panel's (§9). The lead wrote it knowing the arm's question, which disqualifies it from every primary more strongly than any overlap measurement could.

6. The measure

loc, exactly as E-20260810c §7 defines it: for each ⟨site, hand, arm⟩ a judge returns a string naming the marker that would make a reader place the English, or the empty string. A non-empty string is a placement.

loc(arm, judge) = share of the 60 ⟨site, hand⟩ pairs at which that judge returned a non-empty loc string for that arm.

Denominator 60 = 30 sites × 2 hands. Reported per judge and, in the per-hand table, per hand. Absolute levels are never pooled across judges; every difference is read within one judge.

7. Gates

gate bar what it protects
G1 −S purity Nsrc, Asrc and Lsrc contain 0 respellings (code.respellings) at ≥ 28 of 30 sites, each hand an arm that respells is not −I −S vs +I −S; the I/S confound R23 exists to prevent
G2 ~~the procedure changed something~~ demoted to a sanity record, A3 no bar; divergence rates reported the critic's BLOCKING 4 is right that byte divergence between two independently generated arms is guaranteed and shows nothing. F2 is deleted and there is no separate manipulation gate for P1 — see §12.7
G3 batch anchor and reproduction loc(B) ≥ 0.40 on both judges B is frozen and is not one of the arms being compared. S145 measured 0.5714 / 0.5357 and S150 0.5833 / 0.5667 on these two judges. A batch that loses this is not the batch those runs measured in. It no longer pretends to validate the instrument for P1's question (A5)
G3b instrument sensitivity to a standard-spelled located idiom — A5, new loc(PLEAD) − loc(NLEAD) ≥ +0.20 on both judges can this instrument see a located idiom documented to be there, in text with no respelling in it? P1's whole question is standard-spelled located idiom, and G3 says nothing about it
G4 returns 100% of the primary-arm cells — Nsrc, Asrc, ∅, A × 2 hands × 2 judges × the complete-case sites — after at most one re-dispatch (A7) the critic's finding 10: missingness that is arm- or difficulty-correlated moves the primary
~~G5~~ deleted (A5, A6) it was one quantity in two roles and could not corroborate anything (finding 15)

Failure criteria, and what each one does

No bar here is loosened after it fires and none may be moved after any body is read. G3's 0.40 sits below both prior measurements of the same quantity on the same judges; P1's +0.20 is E-20260810c's G2 bar unchanged.

8. Predictions

id statement bar
P1 primary, A1. The located permission, exercised from the source, produces placeable English — measured against a source-first −I arm dispatched in the same stage loc(Asrc) − loc(Nsrc) ≥ +0.20 on both judges
P1b the same effect inside the revision procedure, on the frozen arms, in this batch loc(A) − loc(∅) < +0.20 on both judges
P1c the interaction, and it is what licenses any claim about the procedure (A1, critic 11): does the permission buy more from the source than from a revision? reported as [loc(Asrc) − loc(Nsrc)] − [loc(A) − loc(∅)], per judge. No claim that "S150's null was procedural" may rest on less than this
P1d the pass-to-pass check: two source-first −I passes, one from S150 and one from today, agree reported as loc(Nsrc) − loc(∅), per judge, descriptive
P2 prompt-following reach (A6, narrowed). A hand told to prefer placeable English moves the placement rate loc(Lsrc) − loc(Nsrc) ≥ +0.30 on both judges. Not a measurement of what the material affords
P3 There is no cheap way down reproduces a third time on a frozen arm loc(B) ≥ 0.40 on both judges (the same quantity as G3, declared)
P4 registered mechanical corroborator, no jury, no gate (A3). Frozen located-lexicon hits rank Lsrc > Asrc > Nsrc, and Asrc > A reported on code.lexicon_strict() and code.lexicon(), both
P5 descriptive, no gate: the markers the judges name for each arm, each checked mechanically for presence in the rated text (code.grounded) and for being a respelling reported as a table

Every rate is reported twice (A6): over all non-empty loc calls, and over TEXT-GROUNDED calls only — a marker counts as a placement only if code.grounded() finds it in the English the judge was shown. Where the two readings differ in sign, the grounded one governs and the result says so.

Supporting inference, registered now (A8). A two-sided exact site-level sign-flip permutation over the 30 site-level differences (Asrc minus Nsrc placements, summed over the two hands), per judge, α = 0.05. The site is the analysis unit; ⟨site, hand⟩ is the reporting unit. It is support, not a gate. The "both judges" rule is a conjunctive requirement, which is conservative, and is not offered as a multiplicity correction.

The primary is computed on the complete-case site set (A7) — sites at which every primary arm returned from every hand and both judges rated every one — fixed before the data are read. The add-excluded sensitivity of A2 is reported beside it.

What each outcome means, written before the data

P1 P2 G3b reading
pass — — The device is reachable from the source. With P1c positive, S150's null was procedural and framework/v0.2 §7.2's stated reason must be replaced; with P1c near zero, the permission buys the same little either way and the procedure was never the obstacle
fail pass pass The device is available and merely permitting it does not get it taken. A hand told to prefer placeable English produces it, and the instrument demonstrably reads standard-spelled located idiom; a hand merely permitted it does not, from the source any more than from a revision. The refusal stands and its reason moves off the procedure
fail fail pass The instrument works and neither dose reaches the device on this material. The strongest available statement that the obstacle is the material and the language, not the design
fail — fail F4 fires. P1's null is withheld. The instrument cannot see a located idiom that a frozen log says is there, and P4/P5 are all the run has

All three are informative and the design is indifferent between them, which is the condition R24 was written under and is repeated here.

9. The lead's limb, and what it is for

T-botchan-R24-v1 answers the availability question from inside a translator, on the same span, under the same switch, and its W5+ table is a per-site record of TAKEN / AVAILABLE NOT TAKEN / NONE AVAILABLE. Its counts are computed by code.log_counts() and asserted by verify.py; none is typed by hand (note (bey)).

It is not evidence about the panel's behaviour and not a control for anything. It is quoted beside P2 because a translator saying a located option existed at this site and here it is and a model straining for one are two readings of the same question, and where they disagree that is worth seeing. The lead's two primed rows are excluded from every rate the log yields.

Contamination, measured after the freeze (materials/contamination.json, tools/dependence_check.py):

pair n7 n12 n15 longest run verdict
T-botchan-R24-v1 ~ Morri 1918 9 3 0 14 DEPENDENT?
T-botchan-R22-v1 ~ Morri 1918 (reference, S150) 1 0 0 7 clean
T-botchan-R24-v1 ~ T-botchan-R22-v1 (self) 275 92 50 24 DEPENDENT?

Two things are recorded here before any result is read. The +I rendering sits far closer to the published translation than the −I rendering of the same span by the same translator did — and the 14-token run is "i had studied for three years but to tell the truth i had no", which contains no located idiom at all. And the lead matches itself at 24 contiguous tokens across two sessions with the earlier rendering unopened, which is note (bhb) again and is why T-botchan-R24-v1 is not an independent second opinion of anything. Neither figure gates this run, because the lead's artifact enters no figure in it.

10. Procedure

  1. python3 code.py — self-tests pass. Done before this file was frozen.
  2. Pre-run critic, openai/gpt-5.6-terra, over this design, R24, code.py and the frozen W5+ table. Findings adjudicated in critic.md; every accepted amendment applied before any generation call; every overrule written down with its reason.
  3. Generation — 6 calls: 2 hands × {Nsrc, Asrc, Lsrc}, temperature 0.0, cap 6,000, one call per ⟨hand, arm⟩ over all 30 sites, JSON array out.
  4. Mechanical checks G1 and G2 computed from code.py before any rating call goes out; code.assert_lead_sites() must pass or nothing is dispatched. If F3 fires it fires here.
  5. Rating — 4 blocks × 2 judges = 8 calls (A9); 14 arms per site (6 model arms × 2 hands, plus NLEAD and PLEAD), 420 items, ≤ 112 per call; item order shuffled across the whole block with seed 8155; all arms of a site in the same call (A10, and see §12.7).
  6. analyse.py → analysis/report.json. verify.py recomputes every reported number from the raw bodies and the frozen inputs, and asserts the frozen-input hashes.

Every raw body is written to runs/ before anything is computed from it; a dead body is rotated into runs/discarded/ and never overwritten (note (bhd)).

11. Pre-flight cost estimate — built from max_tokens, not from expected output (note (abc))

Restated with arithmetic after A11 (critic finding 16), and after the critic itself was paid for at $0.02704075 actual.

stage calls cap worst case from max_tokens
critic 1 16,000 $0.02704075 — spent
generation 6 6,000 $0.30
rating 8 16,000 $0.56
base total 15 $0.89
maximum permitted retry path — one full re-dispatch of generation and rating 14 +$0.86
ceiling $1.75

UTC day 2026-08-10 stood at $2.236650490 of $5.00 across seven sessions; headroom $2.763349510, and $2.736308760 after the critic. The ceiling fits with $0.99 to spare. A run that does not fit is split, scaled down or deferred; nothing here needs to be.

12. Declared deviations and known defects of this design

  1. Two hands, both LLMs, one work, one author, one language pair, thirty sentences — S150's limit 2, carried whole, including that the estimand is criterion-positive narration sentences of this span.
  2. P2/G5 and P3/G3 are each one quantity in two roles. Declared here rather than discovered later. G3 is defensible because B is not one of the arms under comparison; G5 is not independent of P2, and F4 is written to make that harmless rather than to hide it.
  3. Lsrc is a dose, not a regime. It is not a serious rendering policy and no claim about good translation rests on it. W8 (not parodic) still binds through BASE.
  4. One post-freeze edit to T-botchan-R24-v1, made before this design existed and recorded here: row A62's item cell named the placeless wording before the located one, which would have put a placeless phrase into the frozen lexicon. The cell was reordered to name the located item first. No translation decision, code, count or coverage code changed; git diff 96ed0aa is the check.
  5. loc is a judgement about English by two models whose competence is asserted by nobody (A10). Tier D is NOT PASSED and every figure here is panel-perceived and provisional.
  6. The frozen ∅, A and B arms were generated in a different call from the one that rates them here. Their re-rating is a fresh measurement of frozen text, which is what makes P1b and P3 reproductions; it is not a re-run of S150 and cannot detect a generation-side defect.
  7. There is no independent manipulation gate on P1, and A3 says so rather than hiding it. Byte divergence between two independently generated arms is guaranteed and gates nothing; a blinded item-level coding of whether a located option was taken is a third panel stage the budget does not hold. The mechanical corroborators are P4 (frozen lexicon, jury-free) and A6's grounding requirement; neither is a gate.
  8. All fourteen arms of a site are rated in one call (A10, critic BLOCKING 7). Arm labels are never shown and item order is shuffled across the whole block, but a judge could in principle read the arms comparatively. The batching is RS-20260806g's rule and is why differences are read within a call at all. The empirical answer is that this exact batching returned loc(A) − loc(∅) = −0.0167 / 0.0000 at S150 — if comparative reading inflates differences, it inflated nothing there. It remains this run's largest instrument limitation.
  9. NLEAD and PLEAD are written by a translator who knew the arm's question, and PLEAD's contamination against Morri 1918 is DEPENDENT? (§9). Neither matters for their one role: G3b asks only whether two judges can see a located idiom that a frozen log documents, in text with no respelling in it. They enter no comparison with any hand and no primary.