Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260810z-idiom-reach.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260810z-idiom-reach
statusfrozen
created2026-08-10
updated2026-08-10
sensesstyle-correspondence
internal-judgment-onlytrue
provisionaltrue
linksworkshop/experiments/E-20260810z-idiom-reach/design.md, workshop/experiments/E-20260810z-idiom-reach/critic.md, workshop/translations/botchan/R24-v1/translation.md, workshop/translations/botchan/R22-v1/translation.md, workshop/regimes/R24-source-first-located-low.md, workshop/regimes/R23-device-crossed-low.md, workshop/regimes/R22-placeless-low.md, wiki/arms/ARM-idiom-reach.md, wiki/findings/results/RS-20260810c-register-reach.md, wiki/findings/results/RS-20260809g-device-cross.md, wiki/findings/results/RS-20260808f-placeless.md, framework/v0.2/README.md, wiki/base/anchors/A-morri-botchan/A-morri-botchan.md, config/models.md, config/budget.md, wiki/method-notes.md

RS-20260810z — there is no placeless English to measure the located device against

E-20260810z-idiom-reach, ARM-idiom-reach step 1, 2026-08-10 (S155). One cell, Japanese → English, 30 frozen sites, 14 arms per site, 2 judges, 840 of 840 rating cells returned. Design frozen at 95db9a2 before the critic saw it; the translation limb frozen at 96ed0aa before the design existed. Pre-run critic NEEDS-REDESIGN, 16 findings, 7 BLOCKING, amendments A1–A11 in force. Verifier 91 checks, 0 failures, 5 of 5 mutation tests caught. $0.172620430 against a declared ceiling of $1.75, key reconciliation exact to 1e-15.

Tier D is NOT PASSED. Every loc figure here is an LLM-panel-perceived judgement about English (A10) and no goodness sense is scored.


1. What the run was for, and the two things it found

framework/v0.2 §7.2 refuses a register-carriage recommendation for the third time because the +I arm at S150 was byte-identical to its own placeless baseline at 93% of hand-sites: an unexercised permission and an absent device look identical. This run wrote the +I arm from the source instead, which is the instrument RS-20260810c §6 named.

The primary FAILS and its null reading is WITHHELD, because the instrument control the critic forced into the design failed on one of the two judges. loc(Asrc) − loc(Nsrc) = +0.0667 and +0.1666 against a bar of +0.20 on both. G3b — can the judges see a located idiom a frozen translator's log documents at 45 of 62 sites? — returned +0.0667 and +0.2000 against +0.20 on both. F4 fires.

What the run does establish is a different thing and it is prior to everything §7 has been arguing about. THE PLACELESS BASELINE IS NOT PLACELESS. Two independent judges place the lead's R22 rendering — written under a rule whose whole content is no lexical item, idiom, or grammatical form that a competent reader would locate — and name, quoting the text back: apologise, trodden, for a song, the state of me, had it brought home to me, give myself airs. Three of those are in the frozen located lexicon mechanically, with no jury at all. Every figure this programme has produced is a difference measured from a floor that is itself located, and one of the markers is a spelling nobody in five sessions has thought about: there is no placeless orthography either. You must choose -ise or -ize.

2. Design, in one table

−I located forbidden +I located permitted +I preferred
minimal revision of the hand's own ∅ ∅ (frozen, S150) A (frozen, S150) —
source-first, one pass Nsrc (new) Asrc (new) Lsrc (new)
the lead, source-first NLEAD = T-botchan-R22-v1 PLEAD = T-botchan-R24-v1 —
(respelling, for the anchor) B = −I +S (frozen, S150)

Hands H1 = x-ai/grok-4.5, H2 = qwen/qwen3.7-max; judges L1 = nemotron-3-ultra, L2 = glm-5.2 — the same two judges as S145 and S150, which is what makes loc(B) a reproduction. loc = share of ⟨site, hand⟩ pairs at which a judge names a marker that would make a reader place the English.

The critic's A1 changed the answer, and by more than the effect. The design as frozen compared Asrc against the S150 ∅; BLOCKING 1 said that conflates the permission with independent-pass variability and demanded a contemporaneous source-first −I arm. It was right: on L2, hand H1, the S150 ∅ is placed at 0.200 and today's Nsrc — same clause, same model, same temperature, same format — at 0.0333. The baseline moves 0.167 between two passes of the identical instruction, which is more than the whole effect under test. P1d = loc(Nsrc) − loc(∅) = −0.0333 / −0.1000, and it is the most consequential number on this page.

3. Gates

gate bar result
G1 −S purity 0 respellings at ≥ 28 of 30 sites, each new hand-arm PASS. 0 respellings in all 180 cells of Nsrc, Asrc, Lsrc
G2 (demoted to a sanity record by A3) none Asrc diverges from ∅ at 0.967 / 0.933; the frozen minimal-revision A at 0.100 / 0.033. The procedure did what it claims
G3 batch anchor loc(B) ≥ 0.40 both judges PASS. 0.5500 / 0.5833 — see §7 on which reading, and on the ambiguity
G3b instrument sensitivity loc(PLEAD) − loc(NLEAD) ≥ +0.20 both judges FAILS. L1 +0.0667, L2 +0.2000
G4 returns 100% of primary-arm cells PASS. 840 / 840 = 1.000

F4 fires: P1's NULL reading is WITHHELD. P1 did not pass either, so nothing is claimed in the other direction. The bar was not moved after it fired, and F4's asymmetry was registered before dispatch, not chosen here.

4. Predictions

id quantity L1 L2 bar verdict
P1 loc(Asrc) − loc(Nsrc) +0.0667 +0.1666 ≥ +0.20 both FAILS, null withheld by F4
P1b loc(A) − loc(∅) 0.0000 0.0000 < +0.20 both PASSES
P1c the interaction +0.0667 +0.1666 reported all of the permission effect is in the source-first procedure
P1d loc(Nsrc) − loc(∅) −0.0333 −0.1000 reported §2
P2 loc(Lsrc) − loc(Nsrc) +0.1167 +0.3000 ≥ +0.30 both FAILS on L1
P3 loc(B) 0.5500 0.5833 ≥ 0.40 both PASSES

Absolute rates, all fourteen arms (denominator 60 ⟨site, hand⟩ pairs for a model arm, 30 sites for a lead arm):

arm L1 L2
Nsrc −I source-first 0.0000 0.0167
∅ −I revision baseline 0.0333 0.1167
A +I minimal revision 0.0333 0.1167
Asrc +I source-first 0.0667 0.1833
Lsrc +I preferred 0.1167 0.3167
B −I +S respelling 0.5500 0.5833
NLEAD the lead, −I 0.0333 0.2000
PLEAD the lead, +I 0.1000 0.4000

Supporting inference (A8, site-level exact sign-flip, support and not a gate): P1 L1 P = 0.125 on 4 non-tied sites, L2 P = 0.0078 on 8. P2 L1 P = 0.0156 on 7, L2 P = 6.1e-05 on 15. P1's bar was a magnitude and it was not met; that a small effect is distinguishable from zero on one judge is not the same claim and is not substituted for it.

A2's add-excluded sensitivity: on the 19 of 30 sites where neither judge flagged an addition in either arm, P1 = +0.0263 / +0.1316 — same direction, smaller. It does not rescue the prediction and was never going to.

5. P1b reproduces exactly, and that is the cleanest thing here

loc(A) − loc(∅) is 0.0000 on both judges — the +I arm produced by minimal revision has identical placement to the placeless draft it was revised from, at both judges, on 60 pairs each. S150 measured −0.0167 and 0.0000 with the same judges in a different batch. The S150 null is not a batch artefact.

And the per-hand table, registered as reporting and not as a test, says the pooled failure is a pooling.

Nsrc → Asrc Nsrc → Lsrc
H1 x-ai/grok-4.5, L1 0.0000 → 0.1333 0.0000 → 0.2333
H1, L2 0.0333 → 0.3000 0.0333 → 0.5333
H2 qwen/qwen3.7-max, L1 0.0000 → 0.0000 0.0000 → 0.0000
H2, L2 0.0000 → 0.0667 0.0000 → 0.1000

One hand takes the permission from the source and the other does not take it at any dose. H2 told it may use every regional slang, local idiom and class-marked grammar English has, and then told to prefer them, produces English that L1 places at zero of sixty either way. This is RS-20260810c's C3 in a new form: a hand that does not apply the manipulation measures nothing, and swapping the hand at S150 did not fix it.

6. P4 — the frozen lexicon, no jury, and the finding that survives everything

The located-item list is extracted mechanically from T-botchan-R24-v1's W5+ table, frozen at 96ed0aa. Strict sublist, 40 items, share of cells containing at least one:

arm rate the items found
Nsrc 0.0000 —
∅ 0.0500
B 0.0500
A 0.0667
Asrc 0.0833 clean off, go round to, the old man ×2, sponge off — all five in H1
Lsrc 0.1500 the five above plus get up to, get it in the neck, telling off, for a song — all nine in H1
NLEAD 0.1000 the state of me, no mind to, for a song
PLEAD 0.3667 eleven — and this figure is CIRCULAR and carries nothing: the lexicon was extracted from this text's own log

PLEAD's 0.3667 is not evidence of anything and is printed only so that nobody reads it as evidence. NLEAD's 0.1000 is not circular in the direction that matters: those three items were found in the R22 text, which was written under a rule forbidding located means and whose 21-row W5 table logs 19 refusals and none of these. At JA-18 the placeless lead rendering says for a song where both models' placeless arms say for next to nothing, and the R24 log lists for a song as row A40, a located choice taken under the +I permission. The same translator used the same idiom under both rules and logged it under one.

7. G3 was ambiguous as registered, and the ambiguity is resolved on a ground that is not the outcome

A6 requires every rate to be reported twice — over all non-empty loc answers, and over text-grounded ones, where the marker the judge names must be present in the English shown. The design said the grounded reading governs a disagreement in sign. G3 is a level, and the design did not say which reading it is read on. That is the RS-20260810x §P-B failure again and it is named rather than smoothed — note (blw).

all calls grounded only
loc(B) L1 / L2 0.5500 / 0.5833 0.1667 / 0.1500
every other arm identical, to four decimals, except ∅/A at L1 (0.0333 → 0.0167)

The mechanism is diagnosed and is arm-specific. B's markers are an', o', 'em — the grounding rule drops tokens under three characters, so a judge quoting B's marker exactly is scored ungrounded. The rule is mis-specified for the respelling arm and for nothing else.

G3 is read on the all-calls figure, and the reason is internal to what G3 is for: it is a reproduction of S145 (0.5714 / 0.5357) and S150 (0.5833 / 0.5667), and those runs had no grounding rule, so a grounded rate is not a reproduction of anything. The choice changes no other number on this page — the primary arms' grounded and all-calls rates are identical — which is why it can be made after the data without the usual objection, and it is stated here so a reader can disagree.

There is no cheap way down now rests on three independent runs.

8. What the judges actually named, and why G3b failed

L1 named a marker 52 times in 420 items; L2 named one 98 times. On PLEAD — the lead's +I English, with a frozen log naming the located item taken at 45 unprimed sites — L1 named three and L2 named twelve. L2's twelve are the log's own items, quoted back: Funk, pinch, had a couple of years on me, went round, got it in the neck, old man, fuss of, state of me, got my back up, clouted, sponging, for a song, favoured, trodden.

G3b did not fail because the located idiom is not in the text. It failed because one of the two judges does not see it. Both judges see the respelling arm at ~0.57. L1's entire vocabulary of placement on standard-spelled text is small: on PLEAD it named Funk, apologise — a spelling — and clouted, telling-off.

This is what A5 was for. Without PLEAD/NLEAD the run would have read P1's small null as the material affords no located idiom, and it would have been wrong: the material affords it, a frozen log documents it, one judge reads it back item by item, and the other is nearly blind to it. The critic's BLOCKING 5, 6 and 14 are the difference between a false conclusion and this one.

9. Three lines that carry the run

Site JA-13, ¶6 — 口惜しかったから、兄の横っ面を張って大変叱られた:

Nsrc I was so mad I slapped my brother's face and got scolded hard. Asrc I was so sore I slapped my brother across the face and got a proper scolding. Lsrc I was so sore I smacked my brother across the face and got a proper telling-off. NLEAD It galled me, so I hit him across the side of the face and was scolded terribly. PLEAD It got my back up, so I clouted him across the side of the head and got a terrible telling-off.

Site JA-18 — 先祖代々の瓦落多を二束三文に売った:

∅ and A, both hands: …sold the family odds and ends for next to nothing. NLEAD, written under the placeless rule: …sold off the family junk, handed down for generations, for a song.

The arm whose rule forbade the located idiom used it; the arms that were offered it declined.

10. Limits

  1. P1's null is withheld and P2 failed. Nothing on this page says the located idiom is unavailable in this material, and §6 and §8 say the opposite is likelier.
  2. Two hands, and one of them contributes nothing (§5). Every model-side placement above the floor is x-ai/grok-4.5's, at both judges.
  3. Two judges who disagree about the object being measured (§8). loc is not calibrated for standard-spelled located idiom, G3b is the first attempt to check whether it is, and it says half the panel is not.
  4. All fourteen arms of a site are rated in one call. A10, overruled remedy: arm labels are never shown and item order is shuffled, but comparative reading is possible. The empirical defence is P1b = 0.0000 twice.
  5. NLEAD/PLEAD are written by a translator who knew the arm's question, and PLEAD's contamination against Morri 1918 is DEPENDENT? — 9 shared 7-grams, 3 twelve-grams, longest run 14 tokens, against 1 / 0 / 7 clean for the −I rendering of the same span by the same hand. §11 says what that is worth. They enter G3b and §6's NLEAD row and nothing else.
  6. The estimand is RS-20260810c's, narrowed: criterion-positive narration sentences of this span, on a criterion that admits 0.874 of it.
  7. P4's PLEAD row is circular and is printed only to be discounted (§6).
  8. One dead body, stage1-loc-c-L2, a provider-side error truncating a JSON array at 424 completion tokens; billed $0 to the project and re-dispatched once under F5. The upstream cost, $0.0144502, was absorbed by the provider and is recorded because it happened, not because it was paid.
  9. No REG stage, no register comparison — registered in §3 of the design as not attempted, so its absence is not a withholding.

11. Two contamination measurements that are a finding, not hygiene

tools/dependence_check.py, after the freeze:

pair n7 n12 n15 longest run verdict
T-botchan-R24-v1 (+I) ~ Morri 1918 9 3 0 14 DEPENDENT?
T-botchan-R22-v1 (−I) ~ Morri 1918 1 0 0 7 clean
T-botchan-R24-v1 ~ T-botchan-R22-v1 275 92 50 24 DEPENDENT?

The same translator, the same source span, the same week: the rendering with the placeless constraint lifted sits far closer to the published translation than the constrained one did. The 14-token run is "i had studied for three years but to tell the truth i had no" — nothing located in it at all. The reading this suggests, and it is a conjecture on one pair and is not a finding: the placeless rule was acting as a de-contamination device, pushing the translator off the readiest English, which is also where a published translator landed. It is recorded so a later run can test it, and RS-20260728b's standing rule is not disturbed: the measurement program on the contamination rule is closed unless a published figure comes to depend on it, and none does.

The self-overlap row is note (bhb) again: 24 contiguous tokens against itself with the earlier rendering unopened, above this project's record against any published human translation. T-botchan-R24-v1 is not an independent second opinion of anything.

12. Cost

$0.172620430 across 15 live bodies, against a declared ceiling of $1.75 (9.9% of it). Key reconciliation EXACT: 89.871470094 → 90.044090524, delta 0.172620430 against a per-response sum of 0.172620430 — agreement to 1e-15.

Waste $0.00 billed. One dead body at $0 (§10.8). Third session running with an exact reconciliation and no paid waste.