Repository path: workshop/experiments/E-20260808e-low-pole/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260808e-low-pole |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| links | wiki/arms/ARM-low-pole.md, workshop/experiments/E-20260808e-low-pole/site-criterion.md, wiki/findings/results/RS-20260806g-negative-pole.md, wiki/findings/results/RS-20260806b-berman-occurrence.md, wiki/findings/results/RS-20260807b-register-collapse.md, workshop/translations/malavoglia-i/R06-v1/translation.md, workshop/translations/malavoglia-i/R21-v1/translation.md, wiki/base/anchors/A-morri-botchan/A-morri-botchan.md, config/models.md, config/budget.md |
| senses | style-correspondence, naturalness |
| internal-judgment-only | true |
| provisional | true |
E-20260808e — the low pole: does a published translator ever go down, and does it depend on where?
Frozen before dispatch of any call. T-malavoglia-i-R06-v1, T-malavoglia-i-R21-v1 and
site-criterion.md were frozen and committed at 422ecff, before Craig 1890 was opened and
before any site was cut. List A (build_sites.py) was built after that commit, from the sources
alone, and is frozen with this file. Amendments after the pre-run critic pass are recorded in
critic.md and marked A<n> here.
1. The question
Across three language pairs and seven hands the project has recorded no instance of a published
English translator lowering the register where the source lowers it (ARM-low-pole §The standing
observation). Two arms instructed to go down reached −0.917 and −1.500 on the same instrument.
Two explanations have never been separated:
- (a) the English-language explanation —
RS-20260806g§2, written from inside the translating: English's raising shelf is deep and available in narration; its lowering shelf is almost entirely dialogue-shaped. If the source's marked sites are narration, there was nothing to reach for. - (b) the translator explanation — decorum, period, audience, programme.
Where a source marks its language below the neutral written register, does a published English hand reproduce the lowering — and does the answer depend on whether the site is narration or direct speech?
Subject-rule sentence (continue-prompt.md §4.5): this unit teaches what published literary
translators do with a source's popular speech, and whether a property of the English language
rather than of the translators explains their doing nothing. The object is translation practice
and the resources of English; no project statistic, rater or published figure is the subject.
2. Materials
Four cells, four language pairs, four published hands, all public domain, original and translation freely reachable and read whole in the spans used (charter §7, A8).
| cell | source | published hand | stored at |
|---|---|---|---|
| IT | Verga, «I Malavoglia» (1881), Ch. I opening, 609 words | Mary A. Craig, 1890 | arms/source_it.txt, arms/craig1890_span.txt |
| JA | Sōseki, 「坊っちゃん」 (1906), Ch. 1 opening | Yasotarō Morri, 1918 | wiki/base/anchors/A-morri-botchan/ |
| FR | Maupassant, «Le Petit Fût» (1884), whole, 1,871 words | anon., The Works of Guy de Maupassant vol. 1, 1903 | arms/petitfut_span_{fr,en}.txt |
| RU | Turgenev, «Бирюк» (1848), the hut scene, 848 words | Constance Garnett, 1895 | arms/biryuk_span_{ru,en}.txt |
Three of the four hands are new to the project. The JA cell re-uses the anchor
A-morri-botchan's stored texts, and its site list is new.
2.1 The lead's translations, and what the contamination gate did to them
The lead rendered the IT span twice — R06 close and R21 vulgarisation — both frozen and
committed before Craig 1890 was opened. tools/dependence_check.py, run immediately after that
commit and before this design existed (runs/dependence.json):
| pair | shared 7-grams | 12-grams | 15-grams | longest run | verdict |
|---|---|---|---|---|---|
Craig 1890 ~ lead R06 |
44 | 6 | 2 | 16 | DEPENDENT? |
Craig 1890 ~ lead R21 |
6 | 0 | 0 | 8 | clean |
lead R06 ~ lead R21 |
20 | 9 | 6 | 20 | DEPENDENT? |
Consequence, taken before any site was cut and not revisited afterwards: the lead's R06
rendering is removed from this experiment entirely. It is not an arm, it is not rated, and no
figure is read on it. A close rendering that shares a 16-token run and six twelve-grams with the
published hand it would be compared against cannot be an independent contemporary hand, whatever
the explanation for the overlap. (The 16-token run is a literal list — had always had boats on the
water and tiles in the sun now at trezza there — and the rule is measured overlap, not the excuse.)
The lead's R21 arm is clean against Craig (6 / 0 / run 8) and is retained, as a manipulation
and not as a translator: it exists to show what the English at these sites can be made to do. It
matches the lead's own R06 at 20 contiguous tokens, which is note (bhb)'s figure and is expected;
R06 is not in the run, so nothing is read across that pair.
3. Sites
Selected against site-criterion.md, frozen before any published English was opened. A site is a
stretch of source text marked below the neutral written literary register of its own language and
period, by at least one of four device classes (LEX, IDIOM, MORPH, SYN), and is typed N
(narration, including free indirect discourse) or D (direct speech) from the source's own
punctuation and framing.
- List A — the lead's, 41 sites,
build_sites.py, frozen with this file. Every source string and every published-English string in it is asserted by the script to occur in the stored text; the script exits non-zero otherwise. Distribution: IT N=10 D=4 · JA N=4 D=4 · FR N=0 D=9 · RU N=1 D=9. Pooled N=15, D=26. - List B — an independent census (
censusstage), a non-panel model given each source span, the criterion verbatim, and nothing else — no English, no hypothesis, no mention of translation.
The primary is read on A ∩ B — the sites the independent census also identifies — with
A \ B and |B \ A| both reported. If |A ∩ B| < 15, the primary falls back to List A and the
selection bias is stated in the result's first paragraph. Declared here, before dispatch.
That FR has no N sites is a fact about Maupassant, recorded rather than repaired: his narration
is standard written French. It means the moderator in §6 P2 is carried by IT and JA.
4. Arms
| arm | what it is | items |
|---|---|---|
PUB |
the published English at the site | 41 |
LOW |
an independent non-panel model's rendering of the same source site, instructed to pitch the English at or below the source's own register with the content unchanged. The reachability floor. Not the lead. | 41 |
V |
the lead's R21 English at the IT sites only |
14 |
96 items. Within each cell the arms are interleaved and shuffled by a fixed seed; no seat sees an arm label, a date, an author, a regime, or which item is published.
5. Instrument
REG, reproduced word for word from E-20260806b/build_payloads.py via E-20260806g — the
scale S118 and S124 both ran, so that this run's figures are comparable to theirs:
REG +3 the English is much higher in register than the source · 0 the same · −3 much lower.
Each seat is given, per item: the source site in its original language, the English at that
site, and one line of surrounding source context. It returns an integer −3…+3 and a why
line naming the English word or phrase its code rests on.
Seats (config/models.md): J1 = P1 openai/gpt-5.6-terra, J2 = P2
google/gemini-3.6-flash, J3 = P5 deepseek/deepseek-v4-pro. One call per cell per seat, 12
bodies. Judgment is not parallelised across seats within a call and no seat sees another's
output.
6. Gates, predictions, failure criteria — frozen
Gates
G1— one-sidedness (F1′re-test). NegativeREGcodes must be ≥ 10% of all returned cells. S118 measured 2.1% and 4.8% on neighbouring scales and withheld two of three primaries on exactly this. IfG1fails, every primary is withheld.G2— the reachability floor. PooledLOWmeanREG≤ −0.75, and negative in ≥ 3 of 4 cells. IfG2fails, no census zero can be read as a refusal andP1/P2are withheld.G3— content parity. An independent non-panel call (a different model from the one that wroteLOW) judges each ⟨source site,PUB,LOW⟩ triple for propositional equivalence, and is given 5 planted content errors among the items. It must name ≥ 4 of 5 planted errors. IfLOWis judged non-equivalent at more than 25% of sites,G2is not usable andP3is withheld — a low score would then be measuring damage.G4— seat competence. 8-item screen, two per language, one factual comprehension question per source site, no English shown. A seat scoring < 6 of 8 is dropped; a cell with fewer than 3 seats is descriptive only.G5— moderator cell size. IfPUBon the primary site list has < 6 N sites or < 6 D sites,P2is withheld.
Predictions
P1— the census claim, under test and able to fail.PUBdoes not go below its source: pooledPUBmeanREG≥ 0, and no single hand's mean below −0.25. This is the standing seven-hand observation extended to four new sites-and-hands; a failure is a finding.P2— the moderator, and the arm's own question.PUB's meanREGat D sites is at least 0.75 scale points below its mean at N sites, and the direction holds in both cells that carry ≥ 3 of each type (IT, JA).P3— the craft claim, tested as a fact about English.LOWmeanREGat N sites ≤ −0.75. IfLOWreaches the low pole at D and not at N, explanation (a) is supported and the published hands' zeros in narration are partly the language's. IfLOWreaches it at both, (a) is not available and the zeros are the translators'.P4— descriptive, no gate.VagainstLOWat the same fourteen IT sites: two low renderings of one span by two different hands under two different instructions.
Failure criteria
F1—G1fails → all primaries withheld.F2—G2fails →P1,P2withheld.F3—G3's planted-error catch < 4 of 5 → the parity call is uninformative,G3cannot clearP3, andP3is reported with the confound unexcluded.F4—G5fails →P2withheld.F5— any cell with fewer than 3 competent seats is descriptive only and is excluded from every pooled figure, with the exclusion stated.
No gate is weakened after it fires. Every withheld primary is reported as withheld.
7. Procedure
snapshot open—GET /api/v1/key.critic— one adversarial pass over this frozen file by a non-panel model; amendments recorded incritic.mdand applied before any other dispatch.census— List B, non-panel, source spans only.low— theLOWarm, non-panel, source sites only, register instruction, content held.screen—G4, 3 seats × 8 items.rate— 4 cells × 3 seats,REG+why.parity—G3, a different non-panel model, with 5 planted errors.snapshot close;analyse.py;verify.pyrecomputing every reported number from the raw bodies.
Every raw body is written to runs/ before anything is computed from it; dead bodies go to
runs/discarded/ and are never overwritten (note (bhd)).
8. Budget
Pre-flight worst case built from max_tokens, not from expected output (note (abc)):
| stage | max_tokens |
worst case |
|---|---|---|
| critic | 16,000 | $0.06 |
| census | 10,000 | $0.07 |
| low | 8,000 | $0.05 |
| screen (3) | 2,000 | $0.05 |
| rate (12) | 6,000 | $0.72 |
| parity | 8,000 | $0.05 |
| declared | $1.20 |
UTC day 2026-08-08 stands at $2.281624846 of $5.00 before this run; declared worst case fits the $2.718375154 headroom. Lead translation is free and is not ledgered (charter §3, A4).
9. Known limitations, written before the result
REGis a direction, not a magnitude.RS-20260806gestablished that the same texts move by a factor of three in negative-code rate depending on what else is in the batch. No figure here is an absolute.- The moderator is carried by two cells. FR has no N sites and RU has one.
P2rests on IT and JA, and a failure there is a failure on two works. - Period is confounded with hand. Three of four published hands are 1890–1918. A Victorian register norm and a translator's decorum cannot be separated by this design.
LOWis a 2026 model told to go low. That it can does not establish that a human translator in 1890 could have reached the same place in the same English.- The lead knew the hypothesis while translating
R21, andR21's log states a conclusion about narration before any measurement.Vis descriptive only for that reason, andP3is read onLOW, which the lead did not write.