Repository path: workshop/experiments/E-20260808f-placeless/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260808f-placeless |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| links | wiki/arms/ARM-low-pole.md, wiki/findings/results/RS-20260808e-low-pole.md, workshop/experiments/E-20260808e-low-pole/design.md, workshop/regimes/R22-placeless-low.md, workshop/regimes/R21-vulgarisation.md, workshop/translations/malavoglia-i/R22-v1/translation.md, workshop/translations/malavoglia-i/R21-v1/translation.md, config/models.md, config/budget.md |
| senses | style-correspondence, naturalness |
| internal-judgment-only | true |
| provisional | true |
E-20260808f — the repaired gate, and whether English's low pole has a place in it
Frozen before dispatch of any call; ten amendments applied after the pre-run critic pass and
recorded in critic.md, each marked A<n> below. workshop/regimes/R22-placeless-low.md and
T-malavoglia-i-R22-v1 — including its translator's log and its dependence measurement — were
committed at df8bf58, before this file existed. This is ARM-low-pole step 2 and the arm's
last session inside its declared budget of 2.
1. What this run is for, and the sentence the subject rule asks for
E-20260808e built a census of what four published English hands do where their sources drop below
the neutral written register, and then withheld all three of its primaries because a
content-parity control it had built itself failed at 0.3902 against a registered bar of 0.25. The
diagnosis is on the result page (§6.1) and is not in doubt: the control asked the wrong question
— whether the published rendering and the low rendering state the same content — which cannot
separate the low arm damaging something from the published hand expanding, and these hands expand.
Two things follow, and they are this run's two stages.
- Stage A repairs the gate, source-relative and on every arm, so that criterion 2 of
ARM-low-polecan be decided rather than left open. - Stage B asks the question the withheld numbers point at. If the low pole is reachable in
English narration, the framework recommendation that has been waiting on this arm — "where the
source marks low register, do X" — needs to know what X costs. Both of this project's low
arms so far reached the pole with English located somewhere:
LOW-Awith Blimey, as hell, old girl; the lead'sR21with a wrong 'un, them that, dead spit, off his grandad. A translator who does not want a Sicilian fishing village speaking Lancashire cannot use either.
Subject-rule sentence (continue-prompt.md §4.5): this unit teaches whether English can be
lowered to a source's popular register without relocating the book to a particular English place,
which is the constraint any register-carriage recommendation has to satisfy. Stage A is a gate
inside the unit it blocks — method work done as a gate, never as the principal, which is exactly
what §4.5 licenses.
2. The wire between the limbs, in one sentence
The translation limb generates the problem the study limb measures: writing «I Malavoglia» twice under two rule sets that differ on exactly one dimension — whether located English is admissible — produced a log claiming that English's placeless low register exists but is made entirely of worn idiom, and stage B puts that claim to a hand that has never seen it.
3. Materials
Everything is stored from E-20260808e and nothing is re-rated that was rated there.
| what | where |
|---|---|
47 admitted sites, four cells, with src, pub, cell, type |
E-20260808e/runs/voted_sites.json |
LOW-A (mistral-medium-3-5), LOW-B (grok-4.5), 52 sites each |
E-20260808e/runs/low{A,B}_parsed.json |
V = lead R21 English at the 17 IT sites |
E-20260808e/arms/v_r21_sites.json |
| 5 planted content errors | E-20260808e/arms/planted.json |
W = lead R22 English at the 17 IT sites |
built here, arms/w_r22_sites.json, from the frozen artifact |
LOW-P = a placeless-constrained low arm |
generated here, stage A0 |
W is a labelled exhibit, never a licensed arm. T-malavoglia-i-R22-v1 declares
contamination: high — the lead had seen Craig 1890 at all seventeen IT sites before writing it —
and measures 9 / 0 / 0, longest run 8, clean against Craig, and 162 / 88 / 64, longest run
43 against the lead's own R21. The second figure is above note (bhb)'s standing 37 and is the
new high-water mark for the lead against itself. V and W are one translator's two passes,
read as a single-variable manipulation and never as two hands. No claim about what English can
reach is read off either. LOW-P carries that.
4. Stage A0 — the LOW-P arm
LOW-P is generated by mistralai/mistral-medium-3-5, the same model that wrote LOW-A, so
that the LOW-A~LOW-P contrast is the constraint and not the model. The prompt is
E-20260808e's low prompt byte-identical except for one paragraph, which is replaced by
R22's W5, stated plainly in English (amendment A10):
Where the source is popular speech, the English is popular speech. Build the lowness ONLY from means available in every standard variety of English: contraction, ellipsis, short clauses, plain common words, bare verbs. Do NOT use any word, idiom or grammatical form that a reader would place in a particular country, region, social class or period — no regional slang, no class-marked grammar, no eye-dialect, no nonstandard spelling. A reader in any English-speaking country must find nothing in your English that belongs somewhere else.
It sees the 17 IT sites only — both types, mixed, in list order, with no type label — and no English, no hypothesis, no mention of narration or dialogue, and nothing about this project.
5. Stage A — G3′, the repaired gate, and the locatability census
One judgement per ⟨site, arm⟩ pair against the source alone. No arm is labelled, no arm is placed beside another, and the checker is never told which English is published. Five calls, split by site so that every call carries a mix of arms:
| call | sites | arms | items |
|---|---|---|---|
IT-1 |
IT sites 1–9 | PUB LOW-A LOW-B LOW-P V W |
54 |
IT-2 |
IT sites 10–17 | the same six | 48 |
JA |
all 12 | PUB LOW-A LOW-B |
36 |
FR |
all 9 | PUB LOW-A LOW-B |
27 |
RU |
all 9 | PUB LOW-A LOW-B |
27 |
192 items + the 5 planted, shuffled within each call under a fixed seed. Three questions per item, all source-relative:
err— does this English omit or misstate something the source states? A few words, or the empty string.add— does this English add an element the source does not have — an extra fact, an intensifier, an epithet, an exclamation, an oath? A few words, or the empty string.loc— would a reader place this English in a particular country, region, social class or period of English? Name the marker, or the empty string.
Checker (amendment A3): qwen/qwen3.8-max — a lab new to this project, not in the panel,
not a seat, wrote no arm, and sat no S136 call. It replaces nvidia/nemotron-3-ultra-550b-a55b,
which answered a related question on 41 of these items at S136; the critic called that dependence
BLOCKING and the swap deletes it rather than mitigating it. Nemotron remains this run's
critic, a role no primary depends on. moonshotai/kimi-k3 is not dispatched — note
(bkh-corr): two finish_reason: length bodies with zero content in this project, $0.125778 and
$0.179433.
Because the replacement has no failure record here, dispatch order is a live probe (A3): the
RU call (27 items, smallest) goes first; only a valid three-field JSON body releases the other
four. A dead or malformed body is rotated to runs/discarded/ and the fallback is nemotron with
limitation 1 restored to the result verbatim. Nemotron's 4-of-4 planted-error record does not
transfer, and G3′a now has real work to do.
The gate, registered before dispatch
G3′ is not a weakened G3. It asks a stricter question (against the source, not against
another translation), of more arms (five, not one), and it adds a calibration the original
lacked. All four clauses must hold or the primaries stay withheld:
G3′a— planted errors. ≥ 4 of 5 named. Below that the call is uninformative.G3′b— the floor arms (amendmentA8, symmetric and stricter). BothLOW-AandLOW-Bmust clear. If either fails,P1andP3stay withheld andARM-low-polecloses with criterion 2 unmet. No survivor is promoted andFLOORis never redefined — the stored figures assume the per-site minimum of both arms, and redefining the floor would require recomputing them, which is the last degree of freedom this design has and it is hereby closed.G3′c— the calibration, relational (amendmentA1).err(PUB)is reported, and a floor arm clears iferr ≤ 0.25ORerr ≤ err(PUB). A published translation put through the identical question is the only way to know whether 0.25 measures the arms' damage or the checker's strictness; iferr(PUB) > 0.25, the absolute bar is declared not interpretable in isolation and the relational clause carries the gate alone. This is weaker than the clause as first frozen in exactly one case, whichcritic.md§A1 names; it is not weaker thanG3, the gate that fired, which had no calibration at all.G3′d— diagnostic, not a gate (amendmentA2). Iferr(LOW) < err(PUB) − 0.10, the discrepancy is reported and read against theaddrates. These published hands are documented to expand, so a more-faithful floor arm may be a true finding and nothing is withheld on it.
What G3′ passing releases, fixed here so there is no freedom left
If G3′ holds, P1, P2 and P3 are read exactly as RS-20260808e §8 computed them, from the
S136 stage-4 bodies, recomputed by verify.py. Nothing is re-rated. No number is recomputed with a
new choice. For the record, and so that reading them changes nothing:
P1—PUB − FLOOR≥ +0.75 per cell: IT +1.2549, JA +2.3611, FR +2.1852, RU +1.0741; pooled +1.6809. Four of four clear.P2—PUB(D) − PUB(N)≤ −0.75 in every eligible cell: IT −0.5463, JA −1.0556.P2FAILS on its registered criterion and is reported as failing.P3—FLOOR(N) − PUB(N)≤ −0.75 per cell with ≥3 N: IT −1.5000, JA −3.0000. Both clear.
Q0 — the qualifier on P3, registered, because it is the honest reading
A floor arm that reaches below the published hand by supplying an oath has not shown that the source's register is carriable; it has shown that English can be made vulgar. So:
Q0. RecomputeFLOOR(N) − PUB(N)on the N sites where the floor arm shows noadd. If the restricted figure keeps its sign and stays ≤ −0.75, explanation (a) is excluded on carried register. If it does not, the exclusion is qualified to "the low pole is reachable in narration only by supplying", and that qualification is the finding, stated in the result's first paragraph rather than in its limits.
6. Stage B — is the low pole placeless?
The 8 admitted IT N sites — narration, which is where the arm's whole question lives — with five arms rated in one call per seat:
| arm | what it is |
|---|---|
PUB |
Craig 1890. Anchors the scale and reproduces a known sign. |
LOW-A |
the stored unconstrained low arm. |
LOW-P |
the same model under the placeless constraint. The licensed contrast. |
V |
lead R21. Exhibit. |
W |
lead R22. Exhibit. |
40 items, 3 seats, interleaved and shuffled under a fixed seed; no seat sees an arm label, a
date, an author, a regime, or which item is published. Instrument: REG, reproduced word for
word from E-20260808e/run.py, which reproduced it from E-20260806b.
Seats (config/models.md): J1 = P1 openai/gpt-5.6-terra, J2 = P2
google/gemini-3.6-flash, J3 = P5 deepseek/deepseek-v4-pro. Judgment is not parallelised
within a call and no seat sees another's output. No seat wrote any arm in this run; LOW-B is
not in stage B, so P3/grok's absence from the seats is moot here and LOW-A's author (mistral) is
non-panel.
Gates and questions, frozen
G6— the manipulation took.loc(LOW-P) ≤ 0.25over its 17 audited sites, from stage A. If it fails, the placeless instruction did not produce placeless English andQ1is withheld.G7— it was not bought with content.err(LOW-P) ≤ 0.25, same bar and same call. If it fails,Q1is withheld.-
G8— the constraint bit at all.loc(LOW-A) − loc(LOW-P) ≥ 0.20. If the unconstrained arm was already placeless, there was no manipulation to test andQ1is descriptive only, with the reason stated. -
Q1— the primary, licensed, and able to fail.REG(LOW-P) − REG(LOW-A)at the 8 N sites ≥ +0.75: constraining a low arm to placeless English raises it. Support means the translator's logs are right and the low pole in English narration is bought with a place. Failure —|Q1| < 0.75— means placeless English reaches the same depth, the logs are wrong, and the framework can recommend carriage without relocation. Prediction:Q1holds. The rule set is indifferent and the failure is the more useful outcome for the framework. Q2— descriptive, contaminated, and diluted by a measured amount (amendmentA5).VandWare byte-identical at 3 of the 8 N sites —IT-03,IT-09,IT-12— which force a zero difference by construction.Q2is computed on the 5 differing sites and says so wherever it appears. The 3 identical sites are themselves a finding: at three of eight narration sites the unconstrained low rendering was already placeless and W5 had nothing to remove.Q3— descriptive, and it is what the framework actually needs.loc(PUB)across all 47 sites, split byREGsign and by period band (A6): pre-1900 {Craig 1890 IT, Garnett 1895 RU} versus post-1900 {the 1903 FR, Morri 1918 JA}. Period is perfectly confounded with the cell pair, so the split carries no inference and is reported so that the confound is named rather than hidden. Where a published hand did go below its source — the four sites in 47 — was its English located? If the published low is located, then the four instances are not a model a translator can copy; if it is placeless, they are.Q4— descriptive.add(PUB)versusadd(LOW-A),add(LOW-B),add(V),add(W).R22's W6 forbids supplying and the log says a human found it cheap to obey;E-20260808e§6.2 measured a model arm supplying at 0.3171. This is the first measurement of the same quantity on a published hand.
Failure criteria
F1—G3′aorG3′cfails → the whole stage-A call is uninformative; every primary stays withheld;Q0is not computed; stage B still runs (it does not depend onG3′) butQ1also needsG6/G7, which come from the same call, soQ1is withheld too and the run reports stage B descriptively.F2—G3′bfails on both arms →P1,P2,P3stay withheld;ARM-low-polecloses with criterion 2 unmet, and the arm page says so.F3—G6orG7fails →Q1withheld.F3′— the futility rule (amendmentA4).|Q1| < 0.30→ the result reads "inconclusive on power", never "the framework can recommend carriage without relocation." Only|Q1| ≥ 0.75in the predicted direction supportsQ1; the band between is reported as the band it is.F6— no backup checker (amendmentA9). IfG3′afails, no second checker is dispatched this session; the primaries stay withheld and the arm closes with criterion 2 unmet. Registering a backup would be an invitation to shop for a checker that passes.F4— a seat returns fewer than all its items → that seat is dropped from stage B and the reduced seat count is stated; below 2 seats stage B is descriptive only.
No gate is weakened after it fires. Every withheld quantity is reported as withheld.
7. Procedure
snapshot open—GET /api/v1/key.critic— one adversarial pass over this frozen file by a non-panel model; amendments recorded incritic.mdand applied before any other dispatch.lowp— stage A0.parity— stage A, five calls.rate— stage B, three calls.snapshot close;analyse.py;verify.pyrecomputing every reported number from the raw bodies, including the S136 figuresP1/P2/P3are read from.
Every raw body is written to runs/ before anything is computed from it; dead bodies go to
runs/discarded/ and are never overwritten (note (bhd) — violated by the lead at S136 and named
here so it is not violated twice).
8. Budget
Pre-flight worst case built from max_tokens, not from expected output (note (abc)):
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| critic | 1 | 16,000 | $0.06 |
| lowp | 1 | 6,000 | $0.05 |
| parity | 5 | 14,000 | $0.50 (qwen/qwen3.8-max, $2.00 / $6.00 per M) |
| rate | 3 | 9,000 | $0.25 |
| declared | $0.96 |
UTC day 2026-08-08 stands at $3.138678256 of $5.00 before this run, with $1.861321744 headroom. The declared worst case fits. Lead translation is free and is not ledgered (charter §3, A4).
A slug this run will not dispatch: moonshotai/kimi-k3, per note (bkh-corr), whose remedy
failed at S136 because its precondition was never written down. It is written down here.
9. Known limitations, written before the result
- ~~Stage A is not independent of the S136 parity call.~~ DELETED by amendment
A3— the checker was replaced with a model that sat no S136 call, rather than the dependence being argued away. What replaces it as a limitation: the new checker has no track record in this project at all, which is why theRUcall is dispatched first as a live probe and whyG3′ais doing real work. Q1rests on one cell, one source, eight narration sites, and one generating model. Exact paired permutation on 8 sites bottoms out at P = 0.0039; that is the most this design can say.VandWare one translator's two passes and share a 43-token run. Nothing about English is read off them.REGis a direction, not a magnitude, and is batch-sensitive by a factor of three (RS-20260806g). No stage-B number is compared with any stage-A or S136 number; stage B is read only as within-call contrasts.locis a single model's judgement of what "sounds like somewhere". It is the same kind of judgement the translator's log makes, taken from a party that does not know the hypothesis — an improvement on one hand's opinion, not an authority.- The seats are three LLMs,
G4was not re-run (stage B re-uses S136's screened seats on the same materials), and the critic's standing BLOCKING 8 — no human raters exist in this project — is unrepaired and is a property of the project. - Period is still confounded with hand across the four published hands, 1890–1918.