Repository path: workshop/experiments/E-20260810z-idiom-reach/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260810z-idiom-reach |
| status | frozen |
| created | 2026-08-10 |
| updated | 2026-08-10 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-idiom-reach.md, workshop/regimes/R24-source-first-located-low.md, workshop/regimes/R23-device-crossed-low.md, workshop/regimes/R22-placeless-low.md, workshop/translations/botchan/R24-v1/translation.md, workshop/translations/botchan/R22-v1/translation.md, workshop/experiments/E-20260810c-register-reach/design.md, wiki/findings/results/RS-20260810c-register-reach.md, wiki/findings/results/RS-20260809g-device-cross.md, framework/v0.2/README.md, config/models.md, config/budget.md, wiki/method-notes.md |
E-20260810z — an unexercised permission and an absent device: telling them apart
Frozen 2026-08-10 (S155) before dispatch. ARM-idiom-reach step 1. The translation limb
(T-botchan-R24-v1, its W5+ table and R24) was committed at 96ed0aa before this file
existed; code.py was written and self-tested before it too; this file was committed at
95db9a2 before the pre-run critic saw it.
AMENDED after the pre-run critic, before any generation call.
NEEDS-REDESIGN, 16 findings, 7 BLOCKING; 12 accepted, 4 accepted-in-part, 5 individual remedies overruled in writing. AmendmentsA1–A11are in force and are marked in place below. The adjudication iscritic.md; the pre-amendment text isgit show 95db9a2:.../design.md.
What the critic changed, in one paragraph. The design compared a newly generated source-first
+I arm against a −I arm generated at S150 in a different call, which conflates the permission
with independent-pass variability (BLOCKING 1, 2, 11): a contemporaneous source-first −I arm
Nsrc is added and the primary moves inside the resulting 2 × 2 (A1). The positive control
was an arm instructed to maximise the measured outcome, which validates nothing (BLOCKING 5, 6,
14): it is replaced by the lead's own frozen R22 and R24 English at each site, a documented
negative/positive pair (A4, A5). A loc answer was counted without checking the named marker
was in the text (13): it now must be (A6).
1. Question
framework/v0.2 §7.2 refuses a register-carriage recommendation for the third time, and
RS-20260810c §6 names the one thing that would move it:
is the located idiom unavailable in Japanese narration, or is it merely unreachable by minimal revision? The instrument that would separate them writes the
+Iarm from the source, not from the∅, and pays the tie-rate cost that the minimal-pair procedure was adopted to avoid.
This is that instrument.
When two independent hands are given the located-idiom permission and write from the Japanese rather than revising their own placeless English, does the English they produce become placeable — and does it become more placeable than the same permission buys inside a revision?
(The second half is the critic's A1: the question is only answerable as an interaction, against
a source-first −I arm dispatched in the same stage.)
2. What is carried from E-20260810c unchanged, and why that is the whole design
Everything except the one clause under test. The sites, the hands, the judges, the loc
prompt, the item-batching rule and the respelling protocol are the previous run's, and three of
the six model arms rated here are the previous run's frozen output, re-rated in this batch.
| carried | what |
|---|---|
| sites | the same 30, runs/sites.json of E-20260810c, verbatim, ids JA-01–JA-30. No new census. The 0.874 admission rate and the narrowed estimand of RS-20260810c §3 carry with them |
| hands | H1 = x-ai/grok-4.5, H2 = qwen/qwen3.7-max |
| judges | L1 = nvidia/nemotron-3-ultra-550b-a55b, L2 = z-ai/glm-5.2 — the same two as S145 and S150, which is what makes loc(B) a reproduction rather than a new measurement |
loc prompt |
LOC_HEAD verbatim, all three questions (err, add, loc) |
| batching | rating is split by site, with all arms of a site kept together in one call (RS-20260806g: loc and REG are batch-sensitive, and a difference is only read within one call) |
−S protocol |
code.py imports E-20260810c/code.py rather than copying it, so that "respelling" means here exactly what it meant there |
| note (bkr) | no content-parity check gates anything. err and add are asked inside the reproduced prompt and reported descriptively |
A21 |
where an arm's English at a site is byte-identical to that hand's ∅, that is recorded as a non-application and is not treated as a device |
The one change is the procedure by which the +I arm is written. In E-20260810c the A
cell was a minimal revision of the hand's own ∅. Here Asrc is the same clause, the same
model, the same temperature and the same item format, written from the Japanese with no English
in front of it.
3. What this run does NOT attempt, registered so a reader does not look for it
There is no REG stage and no register comparison. P1a — how much of English's downward
movement each device buys — is not asked here and is not withheld here; it is simply not this
run's question. §7.2's obstacle was the manipulation, and until a +I arm exists that actually
differs from its baseline there is nothing to compare. This run is a reach run. If it passes,
the register comparison is the successor and it needs a REG panel; if it fails, the register
comparison has no arm to run on.
No new census, no new sites, no new criterion. Reusing a frozen list removes a degree of freedom the lead would otherwise have after having translated this span twice.
4. Arms
Six model arms per hand, plus two lead instrument controls — fourteen English renderings per site.
| arm | switches | procedure | provenance |
|---|---|---|---|
∅ |
−I −S |
one pass, from the source, per site | frozen, E-20260810c/runs/gen-{H}-null.parsed.json |
A |
+I −S |
minimal revision of that hand's ∅ |
frozen, E-20260810c/runs/gen-{H}-A.parsed.json |
B |
−I +S |
minimal revision of that hand's ∅ |
frozen, E-20260810c/runs/gen-{H}-B.parsed.json |
Nsrc |
−I −S |
one pass, from the source, per site — A1, added by the critic |
new |
Asrc |
+I −S |
one pass, from the source, per site — R24's procedure |
new |
Lsrc |
+I −S, dose 2 |
one pass, from the source, preferring the placeable option wherever English offers one | new |
NLEAD |
−I −S |
the lead's T-botchan-R22-v1 English at this site — instrument control (A4) |
frozen, S150 |
PLEAD |
+I −S |
the lead's T-botchan-R24-v1 English at this site — instrument control (A4) |
frozen, 96ed0aa |
A1 — the design is a 2 × 2 and the primary lives inside it.
−I |
+I |
|
|---|---|---|
| minimal revision | ∅ (frozen) |
A (frozen) |
| source-first | Nsrc |
Asrc |
Nsrc and Asrc are dispatched in the same stage, to the same models, at the same temperature,
over the same items, in the same format, and their prompt strings differ in exactly one clause
— machine-checked before dispatch: the unified diff of the two is two lines and both are the
clause. Nothing else about them differs, which is what loc(Asrc) − loc(Nsrc) therefore isolates.
A4 — NLEAD and PLEAD are instrument controls and enter G3b and nothing else. They are
one rendering per site with no hand, drawn from two frozen lead translations of this exact span:
NLEAD written under a rule that forbade located means, PLEAD under one that permitted them,
with a per-site log naming the located item taken at 45 unprimed sites. The paragraph-to-site
alignment is the lead's, made once, and code.assert_lead_sites() proves every span is a
contiguous substring of the frozen translation — the boundaries are the lead's choice and not
one word is. Three of the thirty spans are byte-identical across the pair (JA-11, JA-14,
JA-27), which bounds G3b at 0.90 rather than 1.00 and is stated here rather than discovered.
Nsrc's prompt is E-20260810c's BASE + CLAUSE["null"] + TAIL, and Asrc's is the same
with CLAUSE["A"], both with the REVISE block absent. The REVISE block is what makes A a
revision of ∅; its absence is what makes Asrc a translation.
Lsrc's prompt is Asrc's with one sentence appended to the clause, printed here in full so
the dose is on the record:
At every site where such an option exists, PREFER it: where English offers both a placeless wording and one a reader would place, choose the one a reader would place.
Lsrc is an instrument dose, not a translation regime, and no artifact is filed for it. Its
role is narrowed by A6: it measures prompt-following reach — how far a hand told to
prefer placeable English moves the placement rate — and it is not offered as a measurement of
what the material affords. The critic's BLOCKING 5 is right that an arm instructed to maximise the
measured outcome cannot validate the instrument, and A4's PLEAD/NLEAD pair does that job
instead.
5. Roles, and no model does double duty
| role | slug | also used as |
|---|---|---|
| pre-run critic | openai/gpt-5.6-terra |
nothing else in this run |
hand H1 |
x-ai/grok-4.5 |
nothing else |
hand H2 |
qwen/qwen3.7-max |
nothing else |
judge L1 |
nvidia/nemotron-3-ultra-550b-a55b |
nothing else |
judge L2 |
z-ai/glm-5.2 |
nothing else |
The lead is not in the run. T-botchan-R24-v1 is rated by nobody, enters no figure, and is
reported only as a translator's-chair record beside the panel's (§9). The lead wrote it knowing the
arm's question, which disqualifies it from every primary more strongly than any overlap measurement
could.
6. The measure
loc, exactly as E-20260810c §7 defines it: for each ⟨site, hand, arm⟩ a judge returns a string
naming the marker that would make a reader place the English, or the empty string. A non-empty
string is a placement.
loc(arm, judge)= share of the 60 ⟨site, hand⟩ pairs at which that judge returned a non-emptylocstring for that arm.
Denominator 60 = 30 sites × 2 hands. Reported per judge and, in the per-hand table, per hand. Absolute levels are never pooled across judges; every difference is read within one judge.
7. Gates
| gate | bar | what it protects |
|---|---|---|
G1 −S purity |
Nsrc, Asrc and Lsrc contain 0 respellings (code.respellings) at ≥ 28 of 30 sites, each hand |
an arm that respells is not −I −S vs +I −S; the I/S confound R23 exists to prevent |
G2 ~~the procedure changed something~~ demoted to a sanity record, A3 |
no bar; divergence rates reported | the critic's BLOCKING 4 is right that byte divergence between two independently generated arms is guaranteed and shows nothing. F2 is deleted and there is no separate manipulation gate for P1 — see §12.7 |
G3 batch anchor and reproduction |
loc(B) ≥ 0.40 on both judges |
B is frozen and is not one of the arms being compared. S145 measured 0.5714 / 0.5357 and S150 0.5833 / 0.5667 on these two judges. A batch that loses this is not the batch those runs measured in. It no longer pretends to validate the instrument for P1's question (A5) |
G3b instrument sensitivity to a standard-spelled located idiom — A5, new |
loc(PLEAD) − loc(NLEAD) ≥ +0.20 on both judges |
can this instrument see a located idiom documented to be there, in text with no respelling in it? P1's whole question is standard-spelled located idiom, and G3 says nothing about it |
G4 returns |
100% of the primary-arm cells — Nsrc, Asrc, ∅, A × 2 hands × 2 judges × the complete-case sites — after at most one re-dispatch (A7) |
the critic's finding 10: missingness that is arm- or difficulty-correlated moves the primary |
~~G5~~ |
deleted (A5, A6) |
it was one quantity in two roles and could not corroborate anything (finding 15) |
Failure criteria, and what each one does
F1—G3fails on either judge →P1,P1b,P2andP3are all withheld. The batch'slocis not the instrument S145 and S150 measured with.F2— ~~G2~~ deleted (A3).F3—G1fails on a hand-arm → the failure is reported and the primary is withheld, not recomputed on the surviving hand.A7, from the critic's finding 9: a denominator that moves after an arm-specific failure is observed can only make a result look stronger.F4—G3bfails →P1's NULL reading is withheld and its positive reading is not. Asymmetric, and registered here rather than chosen later: an instrument that cannot see the device makes a null uninterpretable, while a positive result under a blind instrument is conservative. IfP1passes andG3bfails,P1stands and the result says the instrument control failed.F5—G4fails after one re-dispatch →P1is withheld and the run reports the returned cells with the shortfall and its distribution across arms, judges and sites named.
No bar here is loosened after it fires and none may be moved after any body is read. G3's
0.40 sits below both prior measurements of the same quantity on the same judges; P1's +0.20 is
E-20260810c's G2 bar unchanged.
8. Predictions
| id | statement | bar |
|---|---|---|
P1 |
primary, A1. The located permission, exercised from the source, produces placeable English — measured against a source-first −I arm dispatched in the same stage |
loc(Asrc) − loc(Nsrc) ≥ +0.20 on both judges |
P1b |
the same effect inside the revision procedure, on the frozen arms, in this batch | loc(A) − loc(∅) < +0.20 on both judges |
P1c |
the interaction, and it is what licenses any claim about the procedure (A1, critic 11): does the permission buy more from the source than from a revision? |
reported as [loc(Asrc) − loc(Nsrc)] − [loc(A) − loc(∅)], per judge. No claim that "S150's null was procedural" may rest on less than this |
P1d |
the pass-to-pass check: two source-first −I passes, one from S150 and one from today, agree |
reported as loc(Nsrc) − loc(∅), per judge, descriptive |
P2 |
prompt-following reach (A6, narrowed). A hand told to prefer placeable English moves the placement rate |
loc(Lsrc) − loc(Nsrc) ≥ +0.30 on both judges. Not a measurement of what the material affords |
P3 |
There is no cheap way down reproduces a third time on a frozen arm | loc(B) ≥ 0.40 on both judges (the same quantity as G3, declared) |
P4 |
registered mechanical corroborator, no jury, no gate (A3). Frozen located-lexicon hits rank Lsrc > Asrc > Nsrc, and Asrc > A |
reported on code.lexicon_strict() and code.lexicon(), both |
P5 |
descriptive, no gate: the markers the judges name for each arm, each checked mechanically for presence in the rated text (code.grounded) and for being a respelling |
reported as a table |
Every rate is reported twice (A6): over all non-empty loc calls, and over TEXT-GROUNDED calls
only — a marker counts as a placement only if code.grounded() finds it in the English the judge
was shown. Where the two readings differ in sign, the grounded one governs and the result says
so.
Supporting inference, registered now (A8). A two-sided exact site-level sign-flip
permutation over the 30 site-level differences (Asrc minus Nsrc placements, summed over the
two hands), per judge, α = 0.05. The site is the analysis unit; ⟨site, hand⟩ is the reporting unit.
It is support, not a gate. The "both judges" rule is a conjunctive requirement, which is
conservative, and is not offered as a multiplicity correction.
The primary is computed on the complete-case site set (A7) — sites at which every primary arm
returned from every hand and both judges rated every one — fixed before the data are read.
The add-excluded sensitivity of A2 is reported beside it.
What each outcome means, written before the data
P1 |
P2 |
G3b |
reading |
|---|---|---|---|
| pass | — | — | The device is reachable from the source. With P1c positive, S150's null was procedural and framework/v0.2 §7.2's stated reason must be replaced; with P1c near zero, the permission buys the same little either way and the procedure was never the obstacle |
| fail | pass | pass | The device is available and merely permitting it does not get it taken. A hand told to prefer placeable English produces it, and the instrument demonstrably reads standard-spelled located idiom; a hand merely permitted it does not, from the source any more than from a revision. The refusal stands and its reason moves off the procedure |
| fail | fail | pass | The instrument works and neither dose reaches the device on this material. The strongest available statement that the obstacle is the material and the language, not the design |
| fail | — | fail | F4 fires. P1's null is withheld. The instrument cannot see a located idiom that a frozen log says is there, and P4/P5 are all the run has |
All three are informative and the design is indifferent between them, which is the condition
R24 was written under and is repeated here.
9. The lead's limb, and what it is for
T-botchan-R24-v1 answers the availability question from inside a translator, on the same span,
under the same switch, and its W5+ table is a per-site record of TAKEN / AVAILABLE NOT TAKEN
/ NONE AVAILABLE. Its counts are computed by code.log_counts() and asserted by verify.py;
none is typed by hand (note (bey)).
It is not evidence about the panel's behaviour and not a control for anything. It is quoted
beside P2 because a translator saying a located option existed at this site and here it is and
a model straining for one are two readings of the same question, and where they disagree that is
worth seeing. The lead's two primed rows are excluded from every rate the log yields.
Contamination, measured after the freeze (materials/contamination.json,
tools/dependence_check.py):
| pair | n7 | n12 | n15 | longest run | verdict |
|---|---|---|---|---|---|
T-botchan-R24-v1 ~ Morri 1918 |
9 | 3 | 0 | 14 | DEPENDENT? |
T-botchan-R22-v1 ~ Morri 1918 (reference, S150) |
1 | 0 | 0 | 7 | clean |
T-botchan-R24-v1 ~ T-botchan-R22-v1 (self) |
275 | 92 | 50 | 24 | DEPENDENT? |
Two things are recorded here before any result is read. The +I rendering sits far closer to the
published translation than the −I rendering of the same span by the same translator did — and
the 14-token run is "i had studied for three years but to tell the truth i had no", which
contains no located idiom at all. And the lead matches itself at 24 contiguous tokens across two
sessions with the earlier rendering unopened, which is note (bhb) again and is why T-botchan-R24-v1
is not an independent second opinion of anything. Neither figure gates this run, because the lead's
artifact enters no figure in it.
10. Procedure
python3 code.py— self-tests pass. Done before this file was frozen.- Pre-run critic,
openai/gpt-5.6-terra, over this design,R24,code.pyand the frozenW5+table. Findings adjudicated incritic.md; every accepted amendment applied before any generation call; every overrule written down with its reason. - Generation — 6 calls: 2 hands × {
Nsrc,Asrc,Lsrc}, temperature 0.0, cap 6,000, one call per ⟨hand, arm⟩ over all 30 sites, JSON array out. - Mechanical checks
G1andG2computed fromcode.pybefore any rating call goes out;code.assert_lead_sites()must pass or nothing is dispatched. IfF3fires it fires here. - Rating — 4 blocks × 2 judges = 8 calls (
A9); 14 arms per site (6 model arms × 2 hands, plusNLEADandPLEAD), 420 items, ≤ 112 per call; item order shuffled across the whole block with seed 8155; all arms of a site in the same call (A10, and see §12.7). analyse.py→analysis/report.json.verify.pyrecomputes every reported number from the raw bodies and the frozen inputs, and asserts the frozen-input hashes.
Every raw body is written to runs/ before anything is computed from it; a dead body is rotated
into runs/discarded/ and never overwritten (note (bhd)).
11. Pre-flight cost estimate — built from max_tokens, not from expected output (note (abc))
Restated with arithmetic after A11 (critic finding 16), and after the critic itself was paid
for at $0.02704075 actual.
| stage | calls | cap | worst case from max_tokens |
|---|---|---|---|
| critic | 1 | 16,000 | $0.02704075 — spent |
| generation | 6 | 6,000 | $0.30 |
| rating | 8 | 16,000 | $0.56 |
| base total | 15 | $0.89 | |
| maximum permitted retry path — one full re-dispatch of generation and rating | 14 | +$0.86 | |
| ceiling | $1.75 |
UTC day 2026-08-10 stood at $2.236650490 of $5.00 across seven sessions; headroom $2.763349510, and $2.736308760 after the critic. The ceiling fits with $0.99 to spare. A run that does not fit is split, scaled down or deferred; nothing here needs to be.
12. Declared deviations and known defects of this design
- Two hands, both LLMs, one work, one author, one language pair, thirty sentences — S150's limit 2, carried whole, including that the estimand is criterion-positive narration sentences of this span.
P2/G5andP3/G3are each one quantity in two roles. Declared here rather than discovered later.G3is defensible becauseBis not one of the arms under comparison;G5is not independent ofP2, andF4is written to make that harmless rather than to hide it.Lsrcis a dose, not a regime. It is not a serious rendering policy and no claim about good translation rests on it.W8(not parodic) still binds throughBASE.- One post-freeze edit to
T-botchan-R24-v1, made before this design existed and recorded here: rowA62'sitemcell named the placeless wording before the located one, which would have put a placeless phrase into the frozen lexicon. The cell was reordered to name the located item first. No translation decision, code, count or coverage code changed;git diff 96ed0aais the check. locis a judgement about English by two models whose competence is asserted by nobody (A10). Tier D is NOT PASSED and every figure here is panel-perceived andprovisional.- The frozen
∅,AandBarms were generated in a different call from the one that rates them here. Their re-rating is a fresh measurement of frozen text, which is what makesP1bandP3reproductions; it is not a re-run of S150 and cannot detect a generation-side defect. - There is no independent manipulation gate on
P1, andA3says so rather than hiding it. Byte divergence between two independently generated arms is guaranteed and gates nothing; a blinded item-level coding of whether a located option was taken is a third panel stage the budget does not hold. The mechanical corroborators areP4(frozen lexicon, jury-free) andA6's grounding requirement; neither is a gate. - All fourteen arms of a site are rated in one call (
A10, critic BLOCKING 7). Arm labels are never shown and item order is shuffled across the whole block, but a judge could in principle read the arms comparatively. The batching isRS-20260806g's rule and is why differences are read within a call at all. The empirical answer is that this exact batching returnedloc(A) − loc(∅)= −0.0167 / 0.0000 at S150 — if comparative reading inflates differences, it inflated nothing there. It remains this run's largest instrument limitation. NLEADandPLEADare written by a translator who knew the arm's question, andPLEAD's contamination against Morri 1918 isDEPENDENT?(§9). Neither matters for their one role:G3basks only whether two judges can see a located idiom that a frozen log documents, in text with no respelling in it. They enter no comparison with any hand and no primary.