Repository path: wiki/findings/results/RS-20260810z-idiom-reach.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260810z-idiom-reach |
| status | frozen |
| created | 2026-08-10 |
| updated | 2026-08-10 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260810z-idiom-reach/design.md, workshop/experiments/E-20260810z-idiom-reach/critic.md, workshop/translations/botchan/R24-v1/translation.md, workshop/translations/botchan/R22-v1/translation.md, workshop/regimes/R24-source-first-located-low.md, workshop/regimes/R23-device-crossed-low.md, workshop/regimes/R22-placeless-low.md, wiki/arms/ARM-idiom-reach.md, wiki/findings/results/RS-20260810c-register-reach.md, wiki/findings/results/RS-20260809g-device-cross.md, wiki/findings/results/RS-20260808f-placeless.md, framework/v0.2/README.md, wiki/base/anchors/A-morri-botchan/A-morri-botchan.md, config/models.md, config/budget.md, wiki/method-notes.md |
RS-20260810z — there is no placeless English to measure the located device against
E-20260810z-idiom-reach, ARM-idiom-reach step 1, 2026-08-10 (S155). One cell, Japanese →
English, 30 frozen sites, 14 arms per site, 2 judges, 840 of 840 rating cells returned.
Design frozen at 95db9a2 before the critic saw it; the translation limb frozen at 96ed0aa
before the design existed. Pre-run critic NEEDS-REDESIGN, 16 findings, 7 BLOCKING, amendments
A1–A11 in force. Verifier 91 checks, 0 failures, 5 of 5 mutation tests caught.
$0.172620430 against a declared ceiling of $1.75, key reconciliation exact to 1e-15.
Tier D is NOT PASSED. Every loc figure here is an LLM-panel-perceived judgement about English
(A10) and no goodness sense is scored.
1. What the run was for, and the two things it found
framework/v0.2 §7.2 refuses a register-carriage recommendation for the third time because the
+I arm at S150 was byte-identical to its own placeless baseline at 93% of hand-sites: an
unexercised permission and an absent device look identical. This run wrote the +I arm from the
source instead, which is the instrument RS-20260810c §6 named.
The primary FAILS and its null reading is WITHHELD, because the instrument control the critic forced into the design failed on one of the two judges.
loc(Asrc) − loc(Nsrc)= +0.0667 and +0.1666 against a bar of +0.20 on both.G3b— can the judges see a located idiom a frozen translator's log documents at 45 of 62 sites? — returned +0.0667 and +0.2000 against +0.20 on both.F4fires.What the run does establish is a different thing and it is prior to everything §7 has been arguing about. THE PLACELESS BASELINE IS NOT PLACELESS. Two independent judges place the lead's
R22rendering — written under a rule whose whole content is no lexical item, idiom, or grammatical form that a competent reader would locate — and name, quoting the text back: apologise, trodden, for a song, the state of me, had it brought home to me, give myself airs. Three of those are in the frozen located lexicon mechanically, with no jury at all. Every figure this programme has produced is a difference measured from a floor that is itself located, and one of the markers is a spelling nobody in five sessions has thought about: there is no placeless orthography either. You must choose -ise or -ize.
2. Design, in one table
−I located forbidden |
+I located permitted |
+I preferred |
|
|---|---|---|---|
minimal revision of the hand's own ∅ |
∅ (frozen, S150) |
A (frozen, S150) |
— |
| source-first, one pass | Nsrc (new) |
Asrc (new) |
Lsrc (new) |
| the lead, source-first | NLEAD = T-botchan-R22-v1 |
PLEAD = T-botchan-R24-v1 |
— |
| (respelling, for the anchor) | B = −I +S (frozen, S150) |
Hands H1 = x-ai/grok-4.5, H2 = qwen/qwen3.7-max; judges L1 = nemotron-3-ultra,
L2 = glm-5.2 — the same two judges as S145 and S150, which is what makes loc(B) a
reproduction. loc = share of ⟨site, hand⟩ pairs at which a judge names a marker that would make
a reader place the English.
The critic's A1 changed the answer, and by more than the effect. The design as frozen
compared Asrc against the S150 ∅; BLOCKING 1 said that conflates the permission with
independent-pass variability and demanded a contemporaneous source-first −I arm. It was right:
on L2, hand H1, the S150 ∅ is placed at 0.200 and today's Nsrc — same clause, same
model, same temperature, same format — at 0.0333. The baseline moves 0.167 between two
passes of the identical instruction, which is more than the whole effect under test.
P1d = loc(Nsrc) − loc(∅) = −0.0333 / −0.1000, and it is the most consequential number on
this page.
3. Gates
| gate | bar | result |
|---|---|---|
G1 −S purity |
0 respellings at ≥ 28 of 30 sites, each new hand-arm | PASS. 0 respellings in all 180 cells of Nsrc, Asrc, Lsrc |
G2 (demoted to a sanity record by A3) |
none | Asrc diverges from ∅ at 0.967 / 0.933; the frozen minimal-revision A at 0.100 / 0.033. The procedure did what it claims |
G3 batch anchor |
loc(B) ≥ 0.40 both judges |
PASS. 0.5500 / 0.5833 — see §7 on which reading, and on the ambiguity |
G3b instrument sensitivity |
loc(PLEAD) − loc(NLEAD) ≥ +0.20 both judges |
FAILS. L1 +0.0667, L2 +0.2000 |
G4 returns |
100% of primary-arm cells | PASS. 840 / 840 = 1.000 |
F4 fires: P1's NULL reading is WITHHELD. P1 did not pass either, so nothing is claimed in
the other direction. The bar was not moved after it fired, and F4's asymmetry was registered
before dispatch, not chosen here.
4. Predictions
| id | quantity | L1 |
L2 |
bar | verdict |
|---|---|---|---|---|---|
P1 |
loc(Asrc) − loc(Nsrc) |
+0.0667 | +0.1666 | ≥ +0.20 both | FAILS, null withheld by F4 |
P1b |
loc(A) − loc(∅) |
0.0000 | 0.0000 | < +0.20 both | PASSES |
P1c |
the interaction | +0.0667 | +0.1666 | reported | all of the permission effect is in the source-first procedure |
P1d |
loc(Nsrc) − loc(∅) |
−0.0333 | −0.1000 | reported | §2 |
P2 |
loc(Lsrc) − loc(Nsrc) |
+0.1167 | +0.3000 | ≥ +0.30 both | FAILS on L1 |
P3 |
loc(B) |
0.5500 | 0.5833 | ≥ 0.40 both | PASSES |
Absolute rates, all fourteen arms (denominator 60 ⟨site, hand⟩ pairs for a model arm, 30 sites for a lead arm):
| arm | L1 |
L2 |
|---|---|---|
Nsrc −I source-first |
0.0000 | 0.0167 |
∅ −I revision baseline |
0.0333 | 0.1167 |
A +I minimal revision |
0.0333 | 0.1167 |
Asrc +I source-first |
0.0667 | 0.1833 |
Lsrc +I preferred |
0.1167 | 0.3167 |
B −I +S respelling |
0.5500 | 0.5833 |
NLEAD the lead, −I |
0.0333 | 0.2000 |
PLEAD the lead, +I |
0.1000 | 0.4000 |
Supporting inference (A8, site-level exact sign-flip, support and not a gate): P1 L1
P = 0.125 on 4 non-tied sites, L2 P = 0.0078 on 8. P2 L1 P = 0.0156 on 7, L2
P = 6.1e-05 on 15. P1's bar was a magnitude and it was not met; that a small effect is
distinguishable from zero on one judge is not the same claim and is not substituted for it.
A2's add-excluded sensitivity: on the 19 of 30 sites where neither judge flagged an
addition in either arm, P1 = +0.0263 / +0.1316 — same direction, smaller. It does not rescue
the prediction and was never going to.
5. P1b reproduces exactly, and that is the cleanest thing here
loc(A) − loc(∅) is 0.0000 on both judges — the +I arm produced by minimal revision has
identical placement to the placeless draft it was revised from, at both judges, on 60 pairs each.
S150 measured −0.0167 and 0.0000 with the same judges in a different batch. The S150 null is not
a batch artefact.
And the per-hand table, registered as reporting and not as a test, says the pooled failure is a pooling.
Nsrc → Asrc |
Nsrc → Lsrc |
|
|---|---|---|
H1 x-ai/grok-4.5, L1 |
0.0000 → 0.1333 | 0.0000 → 0.2333 |
H1, L2 |
0.0333 → 0.3000 | 0.0333 → 0.5333 |
H2 qwen/qwen3.7-max, L1 |
0.0000 → 0.0000 | 0.0000 → 0.0000 |
H2, L2 |
0.0000 → 0.0667 | 0.0000 → 0.1000 |
One hand takes the permission from the source and the other does not take it at any dose. H2
told it may use every regional slang, local idiom and class-marked grammar English has, and then
told to prefer them, produces English that L1 places at zero of sixty either way. This is
RS-20260810c's C3 in a new form: a hand that does not apply the manipulation measures nothing,
and swapping the hand at S150 did not fix it.
6. P4 — the frozen lexicon, no jury, and the finding that survives everything
The located-item list is extracted mechanically from T-botchan-R24-v1's W5+ table, frozen at
96ed0aa. Strict sublist, 40 items, share of cells containing at least one:
| arm | rate | the items found |
|---|---|---|
Nsrc |
0.0000 | — |
∅ |
0.0500 | |
B |
0.0500 | |
A |
0.0667 | |
Asrc |
0.0833 | clean off, go round to, the old man ×2, sponge off — all five in H1 |
Lsrc |
0.1500 | the five above plus get up to, get it in the neck, telling off, for a song — all nine in H1 |
NLEAD |
0.1000 | the state of me, no mind to, for a song |
PLEAD |
0.3667 | eleven — and this figure is CIRCULAR and carries nothing: the lexicon was extracted from this text's own log |
PLEAD's 0.3667 is not evidence of anything and is printed only so that nobody reads it as
evidence. NLEAD's 0.1000 is not circular in the direction that matters: those three items
were found in the R22 text, which was written under a rule forbidding located means and whose
21-row W5 table logs 19 refusals and none of these. At JA-18 the placeless lead rendering says
for a song where both models' placeless arms say for next to nothing, and the R24 log
lists for a song as row A40, a located choice taken under the +I permission. The same
translator used the same idiom under both rules and logged it under one.
7. G3 was ambiguous as registered, and the ambiguity is resolved on a ground that is not the outcome
A6 requires every rate to be reported twice — over all non-empty loc answers, and over
text-grounded ones, where the marker the judge names must be present in the English shown. The
design said the grounded reading governs a disagreement in sign. G3 is a level, and the
design did not say which reading it is read on. That is the RS-20260810x §P-B failure again and
it is named rather than smoothed — note (blw).
| all calls | grounded only | |
|---|---|---|
loc(B) L1 / L2 |
0.5500 / 0.5833 | 0.1667 / 0.1500 |
| every other arm | identical, to four decimals, except ∅/A at L1 (0.0333 → 0.0167) |
The mechanism is diagnosed and is arm-specific. B's markers are an', o', 'em — the
grounding rule drops tokens under three characters, so a judge quoting B's marker exactly is
scored ungrounded. The rule is mis-specified for the respelling arm and for nothing else.
G3 is read on the all-calls figure, and the reason is internal to what G3 is for: it is a
reproduction of S145 (0.5714 / 0.5357) and S150 (0.5833 / 0.5667), and those runs had no
grounding rule, so a grounded rate is not a reproduction of anything. The choice changes no other
number on this page — the primary arms' grounded and all-calls rates are identical — which is
why it can be made after the data without the usual objection, and it is stated here so a reader
can disagree.
There is no cheap way down now rests on three independent runs.
8. What the judges actually named, and why G3b failed
L1 named a marker 52 times in 420 items; L2 named one 98 times. On PLEAD — the lead's +I
English, with a frozen log naming the located item taken at 45 unprimed sites — L1 named three
and L2 named twelve. L2's twelve are the log's own items, quoted back: Funk, pinch, had a
couple of years on me, went round, got it in the neck, old man, fuss of, state of me, got
my back up, clouted, sponging, for a song, favoured, trodden.
G3bdid not fail because the located idiom is not in the text. It failed because one of the two judges does not see it. Both judges see the respelling arm at ~0.57.L1's entire vocabulary of placement on standard-spelled text is small: onPLEADit named Funk, apologise — a spelling — and clouted, telling-off.
This is what A5 was for. Without PLEAD/NLEAD the run would have read P1's small null as
the material affords no located idiom, and it would have been wrong: the material affords it, a
frozen log documents it, one judge reads it back item by item, and the other is nearly blind to it.
The critic's BLOCKING 5, 6 and 14 are the difference between a false conclusion and this one.
9. Three lines that carry the run
Site JA-13, ¶6 — 口惜しかったから、兄の横っ面を張って大変叱られた:
NsrcI was so mad I slapped my brother's face and got scolded hard.AsrcI was so sore I slapped my brother across the face and got a proper scolding.LsrcI was so sore I smacked my brother across the face and got a proper telling-off.NLEADIt galled me, so I hit him across the side of the face and was scolded terribly.PLEADIt got my back up, so I clouted him across the side of the head and got a terrible telling-off.
Site JA-18 — 先祖代々の瓦落多を二束三文に売った:
∅andA, both hands: …sold the family odds and ends for next to nothing.NLEAD, written under the placeless rule: …sold off the family junk, handed down for generations, for a song.
The arm whose rule forbade the located idiom used it; the arms that were offered it declined.
10. Limits
P1's null is withheld andP2failed. Nothing on this page says the located idiom is unavailable in this material, and §6 and §8 say the opposite is likelier.- Two hands, and one of them contributes nothing (§5). Every model-side placement above the
floor is
x-ai/grok-4.5's, at both judges. - Two judges who disagree about the object being measured (§8).
locis not calibrated for standard-spelled located idiom,G3bis the first attempt to check whether it is, and it says half the panel is not. - All fourteen arms of a site are rated in one call.
A10, overruled remedy: arm labels are never shown and item order is shuffled, but comparative reading is possible. The empirical defence isP1b= 0.0000 twice. NLEAD/PLEADare written by a translator who knew the arm's question, andPLEAD's contamination against Morri 1918 isDEPENDENT?— 9 shared 7-grams, 3 twelve-grams, longest run 14 tokens, against 1 / 0 / 7cleanfor the−Irendering of the same span by the same hand. §11 says what that is worth. They enterG3band §6'sNLEADrow and nothing else.- The estimand is
RS-20260810c's, narrowed: criterion-positive narration sentences of this span, on a criterion that admits 0.874 of it. P4'sPLEADrow is circular and is printed only to be discounted (§6).- One dead body,
stage1-loc-c-L2, a provider-sideerrortruncating a JSON array at 424 completion tokens; billed $0 to the project and re-dispatched once underF5. The upstream cost, $0.0144502, was absorbed by the provider and is recorded because it happened, not because it was paid. - No
REGstage, no register comparison — registered in §3 of the design as not attempted, so its absence is not a withholding.
11. Two contamination measurements that are a finding, not hygiene
tools/dependence_check.py, after the freeze:
| pair | n7 | n12 | n15 | longest run | verdict |
|---|---|---|---|---|---|
T-botchan-R24-v1 (+I) ~ Morri 1918 |
9 | 3 | 0 | 14 | DEPENDENT? |
T-botchan-R22-v1 (−I) ~ Morri 1918 |
1 | 0 | 0 | 7 | clean |
T-botchan-R24-v1 ~ T-botchan-R22-v1 |
275 | 92 | 50 | 24 | DEPENDENT? |
The same translator, the same source span, the same week: the rendering with the placeless
constraint lifted sits far closer to the published translation than the constrained one did. The
14-token run is "i had studied for three years but to tell the truth i had no" — nothing located
in it at all. The reading this suggests, and it is a conjecture on one pair and is not a
finding: the placeless rule was acting as a de-contamination device, pushing the translator off
the readiest English, which is also where a published translator landed. It is recorded so a later
run can test it, and RS-20260728b's standing rule is not disturbed: the measurement program on
the contamination rule is closed unless a published figure comes to depend on it, and none does.
The self-overlap row is note (bhb) again: 24 contiguous tokens against itself with the
earlier rendering unopened, above this project's record against any published human translation.
T-botchan-R24-v1 is not an independent second opinion of anything.
12. Cost
$0.172620430 across 15 live bodies, against a declared ceiling of $1.75 (9.9% of it). Key reconciliation EXACT: 89.871470094 → 90.044090524, delta 0.172620430 against a per-response sum of 0.172620430 — agreement to 1e-15.
Waste $0.00 billed. One dead body at $0 (§10.8). Third session running with an exact reconciliation and no paid waste.