Repository path: workshop/experiments/E-20260811-floor/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260811-floor |
| status | frozen |
| created | 2026-08-11 |
| updated | 2026-08-11 |
| senses | style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-idiom-reach.md, wiki/findings/results/RS-20260810z-idiom-reach.md, workshop/translations/mikan/R06-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R22-placeless-low.md, workshop/translations/botchan/R22-v1/translation.md, workshop/translations/botchan/R24-v1/translation.md, framework/v0.2/README.md, config/models.md, config/budget.md, wiki/method-notes.md |
E-20260811-floor — how far into a page of English prose does the reader learn which country it comes from?
ARM-idiom-reach step 2, respecified by RS-20260810z §7.3. Frozen before the pre-run critic
saw it. The translation limb T-mikan-R06-v1 was frozen at commit 49e154a, before this design
existed, and its translator's log lists the choices its translator noticed making before any
detector in code.py had been run over it.
1. The question, and why it is not method work
framework/v0.2 §7 has spent five sessions measuring what a located device adds to a
placeless baseline. RS-20260810z found that the baseline is not placeless: two judges placed
the lead's R22 rendering — a rule whose entire content is use nothing a reader could place —
and named apologise among the markers. A spelling. Every difference §7 has reported is a
difference from a floor nobody measured.
The question. How much of ordinary, competent English narrative prose — written under no register instruction at all — is nationally located, how early, and can a reader tell?
One sentence on what this teaches about translating literature (continue-prompt.md §4.5):
it teaches whether "write English that no reader could place" names an act a translator can
actually perform, or an instruction that is void on the first page. The object is the English
language as a translator has to write it, not any project statistic. The corpus is twelve published
translators and authors whose national variety is documented; the project's own arms are reported
beside them and are not one side of the primary comparison (note (bly)).
2. Declared deviation from the arm's step-2 wording, with the reason
The arm page says step 2 is "the floor measurement — what share of ordinary competent English
narration two judges will place with no register manipulation at all — plus a loc instrument
shown to read standard-spelled located English."
The loc instrument is not used here, and NAT replaces it. loc asks would a reader place
this English? — a question with no documented right answer for any item, so the instrument can only
ever be checked against another instrument, which is what failed at S155 (G3b, one judge of two).
NAT asks was this written by a British or an American hand? — a question with a documented
right answer for every gate item, so the judges are validated inside this run against the
bibliographic record rather than against each other.
The cost, declared: the floor produced here is not on §7's loc scale and may not be
subtracted from any loc figure. What it licenses is a statement about placement by nation, which
is the channel RS-20260810z §1 actually caught the baseline leaking through.
3. Materials
3.1 The gate corpus — twelve hands, national variety documented
Ground truth GT is the orthographic tradition the hand wrote in, taken from the bibliographic
record (hand's nationality and career, corroborated by place of first publication), fixed before any
text was measured. BR includes the Commonwealth tradition.
| id | text | hand | first publication | GT |
|---|---|---|---|---|
BR1 |
Turgenev, "The Singers" | Constance Garnett (English, 1861–1946) | Heinemann, London 1897 | BR |
BR2 |
Chekhov, "Vanka" | Constance Garnett | Chatto & Windus, London 1922 | BR |
BR3 |
Gogol, "The Mantle" | Claud Field (British, 1863–1941) | London 1916 | BR |
BR4 |
Chekhov, "The Bet" | S. Koteliansky & J. M. Murry (Murry English, 1889–1957) | Maunsel, London 1915 | BR |
BR5 |
Szymański, "Kowalski the Carpenter" | Else Benecke & Marie Busch (British) | Milford/OUP, London 1921 | BR |
BR6 |
Just So Stories (original English) | Rudyard Kipling (English) | Macmillan, London 1902 | BR |
BR7 |
"Miss Brill" (original English) | Katherine Mansfield (NZ, London career) | Constable, London 1922 | BR |
US1 |
Turgenev, "The Singers" | Isabel F. Hapgood (American, 1851–1928) | Scribner, New York 1903 | US |
US2 |
Gogol, "The Cloak" | Isabel F. Hapgood | Crowell, New York 1886 | US |
US3 |
Brazilian Tales | Isaac Goldberg (American, 1887–1938) | Four Seas, Boston 1921 | US |
US4 |
"The Black Cat" (original English) | Edgar Allan Poe (American) | Philadelphia 1843 | US |
US5 |
Akutagawa, "The Spider's Thread" | Glenn W. Shaw (American, 1886–1961) | Hokuseido, Tokyo 1930 | US |
Seven BR, five US; three of the twelve are originally English rather than translated, so the
floor is not an artefact of translation. BR1/US1 are the same Turgenev text in two hands —
the matched pair, which makes content unable to explain any difference between them.
US5 is the one case where hand and publisher disagree (American hand, Tokyo publisher). GT is
scored on the hand, as declared above, and the disagreement is named here so a reader can discount
it; G1 is also reported with US5 dropped.
3.2 Reported beside the gate, with no ground truth and no gate
| id | text | what it is |
|---|---|---|
X1 |
Morri, Botchan ch. 1 (Tokyo 1918) | a Japanese hand writing English for a Japanese publisher — no native variety |
X2 |
T-botchan-R04-v1 |
lead, close translation, no register rule |
X3 |
T-botchan-R06-v1 |
lead, single pass, no register rule |
X4 |
T-botchan-R22-v1 |
lead, the placeless rule — the baseline §7 has measured against |
X5 |
T-botchan-R24-v1 |
lead, located idiom permitted |
X6 |
T-mikan-R06-v1 |
lead, this session, no register rule, log frozen at 49e154a |
3.3 Extraction, declared before measuring
Provenance headers, # comment lines and Project Gutenberg boilerplate are stripped by the rules in
run.py:body(). The measured span is the first 3,000 words of the body, or the whole body if
shorter. Lead artefacts are taken from the plain-text spans files already frozen by earlier
sessions (r04-spans.txt, r06-spans.txt) or from between ## The translation and the next ##
heading. No text was inspected for spelling before its entry in §3.1 was written.
4. The detector (code.py, frozen with this file)
Two inventories, written from reference knowledge of British/American splits, not from any corpus text.
ORTH— spelling proper. Both members of a pair are the same word, so a difference between two texts cannot be a difference of subject, period or register. Primary and gated.LEX— word choice and four morphological preferences (towards, learnt, fortnight, gotten). Reported, never gated: a nineteenth-century American may write railway without error.
The asymmetry rule. Each entry declares which side is diagnostic. -ise is British-marked and
-ize is not American-marked (Oxford spelling); sidewalk is American-marked and pavement is
not British-marked. One-sided entries contribute on one side only. A character span is claimed by
the first rule that matches it, so no token is counted twice.
Per text: n_marks, br, us, fmp (word index of the first mark), density (marks per
1,000 words), verdict (BR/US/TIE/NONE), purity (= max(br,us)/total).
5. Gates — checked before any prediction is read
| gate | bar | why |
|---|---|---|
G1 detector soundness, jury-free |
ORTH verdict on the first half of each gate text equals the verdict on the second half on ≥ 10 of 12, and ORTH purity ≥ 0.80 on ≥ 9 of 12 |
the detector must be measuring a stable property of the text rather than scattered noise. G1 fails → P1–P3 are WITHHELD. Restructured after the pre-freeze dry run; §10 says why, and what it replaced |
G1b detector soundness, jury-side |
token concordance M5 ≥ 0.50 — the share of the words judges quote as deciding a gate item that the detector independently marked in that same passage |
two instruments identifying the same tokens is the only external validation available. G1b fails → P1's claim is about the page and not about reading, and NAT/SWAP carry the reading claim alone |
G2 returns |
100% of NAT and SWAP cells returned, after at most one re-dispatch per dead body |
missingness correlated with item or judge moves every rate |
G3 swap validity |
every SWAP pair differs in exactly one character span and in nothing else; verified mechanically in verify.py |
a "one-letter" manipulation that moved anything else is not the manipulation |
G4 jury floor sanity |
on the SWAP stage, judges answer CANNOT TELL on < 0.90 of items |
a seat that never commits contributes nothing and is reported as abstaining rather than pooled |
6. Predictions, with bars fixed before the run
| # | prediction | bar |
|---|---|---|
P1 |
primary. The floor is immediate. Every gate text carries ≥ 1 ORTH mark, and the median fmp over the twelve is small |
median fmp ≤ 200 words and n_marks ≥ 1 on 12 of 12 |
P2 |
published prose is not merely marked but consistently marked | median ORTH purity over the gate texts ≥ 0.85 |
P3 |
the placeless rule bought nothing measurable. X4 (R22) carries ORTH marks at a density inside the published range |
X4.n_marks ≥ 1 and X4.density ≥ min(density) over the twelve gate texts |
M4 |
reported, not gated: does the mechanical verdict match the hand's documented nationality? Formerly G1; demoted in §10 |
reported as a 12-row table; every mismatch named and discussed |
P4 |
the translator does not notice. The detector finds more distinct mark types in X6 than the frozen log records its translator noticing |
detected distinct ORTH+LEX types in X6 > 13 (the log's count, frozen at 49e154a) — descriptive, no gate |
NAT1 |
the floor in a reader's units. Judges asked to name the variety do not mostly refuse | pooled CANNOT TELL rate over gate items ≤ 0.35 |
NAT2 |
and they are right | among non-CANNOT TELL answers on gate items, agreement with GT ≥ 0.70 on ≥ 2 of 3 judges |
NAT3 |
the matched pair: BR1 and US1 are the same Turgenev in two hands |
reported as a 2 × 3 table; descriptive |
SWAP |
one letter is enough. Flipping a single orthographic mark flips the verdict | verdict changes between forms on ≥ 0.50 of the 16 pairs, on ≥ 2 of 3 judges |
Both directions are informative and the design is indifferent. If P1 fails, the floor is not
immediate and §7's placeless baseline is defensible as written. If NAT1 fails — judges mostly
answer CANNOT TELL — then the marks the detector finds are below a reader's threshold, and the
framework's baseline survives at the level of reading even though it fails at the level of the page.
That would be the more interesting result and it is registered as such.
7. Procedure
corpus— buildmaterials/corpus.jsonfrom §3 by the declared extraction rule; run the detector; writeanalysis/mech.json. No API.items— buildNATitems (2 per text × 18 texts = 36; the first qualifying window of 3–6 sentences and 60–120 words starting at word 200, and the first starting at word 1,000; for texts under 1,200 words, at word 100 and at the midpoint; for the matched pairBR1/US1, passage 1 is the first qualifying window from word 1 in both, so the two are the same Turgenev) andSWAPitems (sentences of 15–45 words containing exactly one invertible ORTH mark and no LEX mark; up to two candidates per gate text, then round-robin across all twelve texts to 16, so every gate text contributes one before any contributes two — without the round robin the first six texts fill the stage and the American-published texts never appear). Fixed shuffle seed 8156. Arm and source labels are never shown to a judge.critic— one independent pre-run critic pass over this file andcode.py(openai/gpt-5.6-terra), adjudicated in writing before any judge is dispatched.nat— 3 blocks × 12 items × 3 judges = 9 calls.swap— 2 forms × 3 judges = 6 calls, 16 items each. FormV1shows pairs 1–8 as published and 9–16 swapped;V2the complement. No judge ever sees both members of a pair in the same call, and the two calls are independent dispatches.verify— recompute every reported number from the raw bodies; mutation tests on the detector;G3checked character by character.
Judges. L1 = nvidia/nemotron-3-ultra-550b-a55b and L2 = z-ai/glm-5.2 — the two seats
RS-20260810z used, so the floor is read by the hands that read S155's arms; L3 =
google/gemini-3.6-flash, a third seat added because S155 showed two seats can disagree about the
object. The critic model is not a judge.
Judgement is not parallelised; every dispatch is sequential and every raw body is written to
runs/ before anything is computed from it. A dead body is rotated into runs/discarded/ and never
overwritten (note (bhd)).
8. Pre-flight cost estimate, built from max_tokens (note (abc))
| stage | calls | prompt (est.) | cap | worst case |
|---|---|---|---|---|
| critic | 1 | 13,000 | 6,000 | $0.049 |
nat |
9 | 2,300 | 1,400 | $0.25 |
swap |
6 | 1,400 | 1,400 | $0.15 |
| re-dispatch headroom | ≤ 4 | — | — | $0.10 |
Declared ceiling: $0.75. Today's UTC ledger (2026-08-11) is empty; the cap is $5.00.
9. What this design does not do
- No
locfigure and no comparison withRS-20260810z's numbers. §2. - No claim about register. Nation-marking and register-marking are different channels; this run measures the first and says nothing about the second.
- The publisher, not the translator, often fixes the spelling. House style is real and is not
controlled for. It does not weaken
P1— a text located by its publisher is still located, and a translator who wants placeless English cannot get it from a publisher either — but it does meanG1's success is evidence about texts, not about translators' habits. - Twelve texts is a small corpus and it is English-only on the target side.
fmpmedians are reported with the full ordered list, not as point estimates. X6'sP4count depends on the log being honest about what its translator noticed. The log was frozen beforecode.pywas written, which is the only guarantee available; the lead wrote both, which is stated rather than defended.
10. What was known before the freeze — declared, not tidied away
This design was written in full, including every bar in §6, before run.py existed and before any
number had been computed. The mechanical stage was then run as a dry run before this file was
committed, and three things happened that a reader is entitled to know about.
- Two rules in
code.pywere false positives on their face and were deleted. The-our/-orstem list containedterr, which fired on terror — a word with no British variant — and the-isestem list containedsurprandadvert, which fired on surprise and advertise, neither of which has an-izeform in any variety. Both deletions reduce the mark count and pushfmplater, which makesP1andP3harder, so the correction runs against this session's expectation and not with it. A negative-control list of 45 non-split words (error, horror, mirror, doctor, author, actor, sailor, tailor, emperor, exercise, enterprise, compromise, promise, precise, otherwise, story, glory, memory, factory, history…) now returns zero marks and is re-checked inverify.py. - Missing splits were deliberately NOT added. The same negative-control pass showed the
inventory has no entry for parlour/parlor, and it plainly lacks others (judgement,
skilful, instalment, fulfil). Adding them would have raised sensitivity in the direction
that helps
P1, after the fact, so none was added. The consequence, stated as a bias direction:n_marksanddensityare lower bounds andfmpis an upper bound. Every figure this run reports is conservative with respect toP1. G1was restructured, and this is the substantive change. As written,G1required the detector to recover each hand's documented nationality. The dry run showed that Isabel Hapgood — an American, published by Scribner in New York — wrote neighbourhood, grey, colour, favourite, woollen, rumours, theatre and honourable, and the same for her 1886 Crowell Cloak. Either the American hand wrote in the British tradition or the scans are of a British-market printing; the design cannot tell which, and neither reading is a detector failure. The originalG1was therefore not a test of the detector at all but a test of an assumption about hands that this session had no right to make. It is demoted toM4, reported with every mismatch named, and replaced by two gates that do not depend on the bibliographic record: internal split-half stability, and concordance with the words the judges themselves quote. No bar in §6 was changed, no text was added to or removed from §3, and the jury limb had not been dispatched when this was written.
What this costs. The mechanical limb of this run is not blind — its numbers were seen before
this file was committed, and it should be read as a measurement with a pre-written analysis plan
rather than as a pre-registered test. The jury limb (NAT, SWAP, G1b, M5) is
pre-registered: not one of its calls had been made when the freeze commit was taken.