Repository path: workshop/experiments/E-20260806g-berman-negative-pole/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260806g-berman-negative-pole |
| status | frozen |
| created | 2026-08-06 |
| updated | 2026-08-06 |
| links | wiki/arms/ARM-berman-occurrence.md, wiki/base/sources/S-berman-tendances.md, wiki/findings/results/RS-20260806b-berman-occurrence.md, workshop/experiments/E-20260806b-berman-occurrence/design.md, workshop/regimes/R21-vulgarisation.md, workshop/translations/cavalleria-rusticana/R21-v1/translation.md, workshop/translations/cavalleria-rusticana/R08-v1/translation.md, config/models.md, wiki/goodness-senses.md |
| senses | style-correspondence, cultural-mediation, naturalness |
| internal-judgment-only | true |
| provisional | true |
E-20260806g — Berman's downward pole: can a translator go the other way, and can a reader tell?
Frozen before dispatch of any call. The R21 rule set and
T-cavalleria-rusticana-R21-v1's translator's log were frozen before this file was written, and
the contamination measurement (§2.2) was run before any site was cut. Amendments after the
pre-run critic pass are recorded in critic.md and marked A<n> here.
1. The question
Berman's fifth and sixth deforming tendencies are one axis with two poles: ennoblissement, the raising of a source's register, and vulgarisation, the false popular speech — « pseudo-argot » — a translator reaches for when the source speaks a vernacular his own language has no equivalent for. Both are deformations; Berman condemns them together.
RS-20260806b measured that axis on five English renderings of Verga's Sicilian and found the
downward pole empty. Strettell 1893 at +0.833, Dole 1896 at +0.583, the foreignising
control at +0.500, the unbriefed 2026 machine at −0.083. On a signed −3…+3 scale, over 186
cells, the negative half was used 29 times in all — and on the neighbouring EXPL and POP scales, 4
times and 9 times, which fired the run's own one-sidedness gate F1′ and withheld two of its three
primaries.
That leaves two questions genuinely open, and they are the same question asked of the craft and of the reading:
Is there anywhere below Verga for an English translator to go — and if a translator goes there on purpose, does a reader register it as lower, or merely as different?
Subject-rule sentence (continue-prompt.md §4.5): this unit teaches whether the downward
deformation a major tradition names is something a translator can actually perform on a popular
source, what performing it costs the story, and whether the direction of a register movement is
visible to a reader at all — or whether any departure from a source's register reads as departure
upward. The objects measured are four English renderings of Verga and one newly written one. The
F1′ re-test in P4 is a gate inside this unit, not its subject: the run would be worth doing
if the scale's behaviour were already settled, because Berman's two-sided claim would still be
untested on any corpus.
Why it is not a re-run. E-20260806b asked do the tendencies occur in what translators
published. This asks is the catalogue's own second pole reachable, and it answers it by making
the missing arm rather than by looking harder at the arms that existed. The new arm is the unit's
translation limb; nothing about it was recoverable from the S118 corpus.
2. Materials
Source. As E-20260806b: Verga, «Cavalleria rusticana», Vita dei campi (1880; 1881 text,
it.wikisource.org), ¶50–81 of 81, 740 Italian tokens. Public domain (Verga d. 1922). arms/IT.txt
is byte-identical to the S118 file.
Four arms, at arms/*.txt:
| arm | rendering | date | role here |
|---|---|---|---|
| V | the lead, R21 vulgarisation, T-cavalleria-rusticana-R21-v1 |
2026 | the new arm; the primary rests on it. Rules frozen before translating, log frozen before this file |
| X | the lead, R08 resistancy, T-cavalleria-rusticana-R08-v1 |
2026 | the arm that read +0.500 on REG at S118 while counter-programmed — the confound this run tries to resolve |
| S | Alma Strettell, 1893, two-scan verified | 1893 | the ennoblement reading being compared against (+0.833 at S118) |
| D | Nathan Haskell Dole, 1896, Gutenberg #37979 | 1896 | as S (+0.583 at S118) |
Two arms of E-20260806b are deliberately dropped. L (the lead's R04 close rendering) is
excluded because §2.2 measures V as dependent on it and on nothing else; keeping both would put
two renderings by the same hand at 9 shared 12-grams into one comparison. P (the unbriefed
deepseek-v4-pro rendering) is dropped so that P5 is free to take a coding seat — a model may
not judge a text it produced (charter §5), and dropping the arm buys a seat that costs a twentieth
of the one it replaces.
2.1 Copy-text
S and D are the S118 files unchanged, with that run's verification (two-scan diff, three
resolved OCR defects, Strettell's paragraph division unrecoverable) carrying over verbatim as a
declared limit. V was extracted from the frozen translation.md by
analysis/extract_arm.py, which re-flows the hard-wrapped paragraphs and asserts a 31-paragraph
alignment with the Italian; the assertion passes.
2.2 Contamination, measured before any site was cut (CLAUDE.md standing rule)
tools/dependence_check.py over all six English renderings of this locus, run after the R21
log was frozen and before this design existed (runs/dependence.json, cells in
runs/dependence_cells.json):
| pair | shared 7-grams | shared 12-grams | longest run | verdict |
|---|---|---|---|---|
| V ~ L — the new arm against the lead's own close rendering | 54 | 9 | 17 | DEPENDENT? |
| V ~ P | 35 | 0 | 11 | clean |
| V ~ D | 13 | 0 | 10 | clean |
| V ~ S | 10 | 0 | 11 | clean |
| V ~ X — two lead renderings, opposite register programmes | 4 | 0 | 8 | clean |
| S ~ D (the published pair, carried over) | 46 | 11 | 17 | DEPENDENT? |
| L ~ P (carried over) | 205 | 90 | 30 | DEPENDENT? |
Four things follow and bind what this run may claim.
- V is clean against every arm the run uses. Zero shared 12-grams against S, D and X; the longest run against any of them is 11 tokens, below this project's own recall floor for a constrained passage (notes (bez), (bgf), (bgi)). Whatever V scores, it does not score it by resembling Strettell, Dole or the resistancy arm.
- V's single dependency is on the lead's own prior rendering, at 9 shared 12-grams and a
17-token run — higher than V's overlap with any published human. This is note (bhb) landing
for the third time and it is why arm L is not in this run. The
contamination: highdeclaration onT-cavalleria-rusticana-R21-v1is this measurement, not a guess. - V ~ X is the lowest-overlap pair in the whole table — 4 shared 7-grams, longest run 8. Two renderings by the same hand, of the same 740 words, in the same repository, under opposite register programmes, are further apart than two published human translators three years and one publishing world apart (46 / 11 / 17). Reported as a fact about what a regime change does to a lead rendering; no gate depends on it, and it is not offered as evidence that either arm is good.
- The run is not licensed to call any pair independent on the tool's verdict, which flags nine
of fifteen pairs including one that is independent by construction. The zeros are the informative
cells; the
DEPENDENT?labels are not.
3. The three coded questions — unchanged, verbatim, from E-20260806b
EXPL, REG and POP are reproduced word for word from E-20260806b/build_payloads.py,
including the signing (positive = Berman's predicted direction), the 0 is a real answer sentence,
and the Use negative values wherever they apply sentence. Changing the instrument while testing
what the instrument can register would make the result uninterpretable, so nothing in the prompt
is touched except the number of versions.
The scales, for reading this file without the other open: EXPL +3 much more explicit / −3 leaves considerably more unsaid · REG +3 much higher / −3 much lower · POP +3 popular marks effaced almost entirely / −3 marked more strongly than the Italian.
4. Sites — the same twelve, nine theory-chosen and three systematically sampled
analysis/sites.json, carried over from E-20260806b with the V spans added by
analysis/build_sites.py and machine-verified: every V span is a verbatim substring of arms/V.txt
(12 of 12) and covers the same paragraph extent as the Italian span (12 of 12).
Keeping N1–N3, the paragraphs chosen at the 1/4, 2/4 and 3/4 positions without inspecting their
content, preserves S118's anti-circularity control: A3 there was the amendment that changed the
run, and dropping it now would quietly re-introduce the bias it was written against.
5. The machine limb (no API call, analysis/machine.py)
M1–M4 recomputed with V added, on the whole locus. Registered before dispatch this time, so
unlike S118 (A8 there) M1 may carry a gate:
- M1 — expansion. English tokens ÷ 740.
- M2 — sentences.
- M3 — type–token ratio, and the
adornare/renderenetworks. - M4 — retained source-culture tokens.
Registered: R21's V3 (contract) was measured at 1.0446× and MISSED — the frozen log records the
miss. M1's role here is F6 below, and the "nobody wrote English shorter than the Italian" finding
of S118 §6.1 is extended, not tested, by an arm that tried to.
6. Procedure
- V produced first, unbriefed by any other arm, log frozen (§2, and
translation.md). - Contamination measured (§2.2). Sites cut. This design frozen.
- Independent pre-run critic, one non-panel model, against this file. Amendments written here before dispatch.
- Per site the seat receives the Italian span and the English versions, labelled
V1… in an order shuffled per site bymd5(site_id), mapping inanalysis/keymap.json, shown to no seat. Provenance, date, regime and authorship are withheld from every seat. - REPEAT control: at S3 and S7, arm V is included twice under two labels (five
versions at those sites). A seat coding identical text differently is measuring noise. (S118 put
the duplicate on the arm its primary used; so does this one.) The raw
md5draw placed the twins adjacent at S7 and one apart at S3, which would have made the control trivial — a seat shown the identical text twice in a row is not being tested on reproducibility. A minimum separation of 2 is therefore imposed by deterministic rotation (build_payloads.py), giving separations of 2 and 4, which is exactly what S118's unconstrained draw happened to produce. The constraint makes the two runs' controls comparable; it does not raise the standard, and it was applied before any dispatch. - Three seats, one call per (site, seat):
P1openai/gpt-5.6-terra·P3x-ai/grok-4.5·P5deepseek/deepseek-v4-pro. 36 calls.P2(google/gemini-3.6-flash) is dropped on this project's own record — it needed a 6,000-token cap at S118 for hidden reasoning and was 80% of the seat spend for 33% of the cells at S120 — andP5replaces it, which is possible only because arm P was dropped. P1 and P3 are shared with S118, which is what licensesF5. - Output format, line guard and
###END###terminator exactly as S118. - Cell value = median of the three seats; arm means pool the twelve sites.
Judgment is not parallelised. The seats are coders of a source–target relation, not a jury.
No goodness sense is scored, no quality claim is made, no arm is ranked against another, and Tier D
is NOT PASSED — every sentence of the result is provisional.
7. Predictions, registered
P1— the primary. The downward pole is reachable and readable. Arm V scores < 0 on REG, with ≥ 8 of 12 site-medians negative, and a REG mean ≥ 1.00 scale point below the pooled S/D mean. (One scale point is S118's declared resolution,A10there; a smaller gap is not reported as a size.)P2— the negative half of EXPL is reachable. Arm V scores < 0 on EXPL.R21's V4 (withhold, do not repair) and V3 (contract) are the rules that should produce it. This is the direct test of the scale thatF1′fired hardest on at S118 (4 negative codes of 186, 2.1%).P3— vulgarisation replaces a vernacular, it does not restore one. Arm V scores ≥ 0 on POP. This is registered against the naive expectation: a low, slangy English looks like it should read as more marked in popular speech, i.e. negative. Berman's position is the opposite — pseudo-argot belongs to the same family as effacement. If V comes out clearly negative on POP, Berman's separation of tendency 6 from tendencies 10–12 is not supported by how these readers code, and the result says so.P4—F1′re-tested (gate, see §8). Negative-code share recomputed per scale with V in the line-up; the S118 threshold of 5% of coded cells stands unchanged.P5— descriptive, no gate: does REG measure direction or departure? Pre-committed both ways, so that neither outcome can be storied afterwards. If V reads negative on REG while X reads positive (X was +0.500 at S118 while counter-programmed), then REG's positive pole is occupied by two different things — raising and marking as foreign — andRS-20260806b's limit 5 is a real confound with a real axis behind it. If V reads positive on REG too, REG is measuring departure from the source's register in either direction, and limit 5 stops being a caveat and becomes a fatal objection to the ennoblement claim, which this run would then withdraw.
8. Failure criteria, registered — what withholds what
F1(gate). REPEAT: if the mean absolute difference between the two V labels at S3 and S7, over three questions and three seats, exceeds 0.50 scale points, all primaries are withheld as coder noise. (S118 recorded 0.000 on this control; a large value here would be new information about the instrument and would stop the run.)F2(gate). If fewer than two of three seats return a usable code in more than 20% of (site, arm, question) cells, primaries are withheld.F3. A seat returning the identical value for every arm at 8 or more of 12 sites on a question is degenerate on that question: reported in the headline, its codes excluded from that question.F4(gate onP1) — the badness confound, and the criterion this design most needs. A coder might score arm V low on REG because it is worse English, not because it is lower English. Every WHY line attached to a |value| ≥ 2 code on arm V is classified against a rubric frozen here:register— names a word, phrase, idiom, oath or address form and its level ("mum", "mates", "Christ almighty", contractions, "thanks a lot");error— names something as wrong, mistranslated, omitted, or a failure of accuracy;other— names something else, or restates the scale (vacuous), or names something not in the span (mislocated).
If error + other exceeds one quarter of arm V's extreme REG codes, P1 is withheld — the
coders would be reading badness rather than lowness. All WHY lines are published verbatim in
the result's provenance directory so the classification can be re-done by anyone; the lead's
classification is internal-judgment-only and the counts are reported with that flag on them.
- F5 (licence, not withholding). Reproduction. If S or D's REG mean in this run
differs from its S118 value by more than 0.75 scale points, the cross-run comparison is not
licensed: V is then reported against this run's S and D only, and no sentence compares any
figure here with a figure there.
- F6. If M1 does not make V the shortest English arm, R21's V3 was not executed at
all and the contraction half of P2's rationale is withdrawn (the withholding half stands).
- F7. If arm V's POP mean and arm X's POP mean have the same sign and differ by less
than 0.50, P3 is reported as inconclusive rather than confirmed: the two arms would be
indistinguishable on the scale that is supposed to separate their programmes.
9. Seats, cost, pre-flight
Worst case built from max_tokens, never from an assumed output length — note (abc), the one
estimate this project has overrun.
| stage | seat | calls | max_tokens | list in/out per M | worst case |
|---|---|---|---|---|---|
| 0 — independent pre-run critic | one non-panel model, nvidia/nemotron-3-ultra-550b-a55b |
1 | 16,000 | — | $0.10 |
2 — coding, P1 |
openai/gpt-5.6-terra |
12 | 2,500 | $1.00 / $6.00 | $0.19 |
2 — coding, P3 |
x-ai/grok-4.5 |
12 | 2,500 | $2.00 / $6.00 | $0.21 |
2 — coding, P5 |
deepseek/deepseek-v4-pro |
12 | 2,500 | $0.435 / $0.87 | $0.03 |
| — routing margin ×1.5 on stage 2 | $0.22 | ||||
| — retry reserve | $0.20 |
Declared worst case $0.95. Today's ledger (2026-08-06 UTC) stands at $3.442032216 of $5.00 across seven sessions; headroom $1.557967784. The run fits with $0.61 to spare.
Seat choice is a cost decision made from this project's own records, and it is stated because it
is also a design change: P2 is dropped and P5 added, so two of three seats are shared with
S118 and F5 is a convergent reproduction check across a partly-changed panel rather than a
strict one. The result reports it that way.
Reserve (note (bfc), declared before dispatch): qwen/qwen3.7-max, which produced no arm in
either experiment. Timeout 300s, halved from the inherited 900s per S122's finding.
Lead translation of arm V cost $0 and is not ledgered (charter §3, A4).
Amendments after the pre-run critic pass
The independent pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel, takes no coding seat,
$0.032205, provider Together, 108.6s) returned NEEDS-AMENDMENT, 10 findings, 5 BLOCKING.
Verbatim body in critic.md. Eight are accepted, one is accepted as a declared limit, and one
BLOCKING finding is rejected with the arithmetic. No coding call had been dispatched when these
were written.
A1 (finding 1, BLOCKING — accepted; the best finding in the pass). F4 as registered could
not fire. The prompt asks for one WHY line per version, naming what drove the coder's most
extreme of the three codes — so if a coder's most extreme code on arm V is EXPL or POP, its WHY
line says nothing about REG, and F4 has no evidence for exactly the codes it must classify. The
critic's first repair — three WHY lines per version — is refused, because §3 fixes the seat's
task verbatim from S118 and tripling the free-text output would change both the seat's behaviour and
the truncation risk that note (bhf) has fired on nine times. The second repair is taken and
tightened:
F4classifies only those arm-V WHY lines where REG is the coder's (co-)most-extreme code, which is computable from the three returned integers and is reported as a coverage figure.- New coverage gate. Arm V has 15 version-instances (12 sites + 3 repeats, per
A6) × 3 seats = 45. If REG is the most extreme code in fewer than 15 of those 45,F4cannot fire, andP1is reported with the badness confound explicitly UNRESOLVED — not cleared, not fatal, undecidable on this evidence and named as such in the headline.
A2 (finding 2, BLOCKING — accepted, and resolved at $0 with data already on disk). POP cannot
by itself separate pseudo-argot (Berman's tendency 6) from authentic vernacular preserved
(tendencies 10–12): both would push the code negative. A fourth scale is refused for the reason
in A1. The repair is that M4 is registered as the discriminator, and it is objective and
already computed: arm V retains zero source-culture tokens, the lowest count of any arm this
project has measured (S 3, D 22, X 27). So:
P3is demoted from a prediction to an exploratory comparison with its direction still registered.- A negative POP on arm V may never be reported as "the vernacular was restored." With M4 at zero, a negative POP can only mean the coders are registering English vernacular marking, not the source's — which is the finding, and it is reported as a fact about the scale and about reading, not as a success for the arm.
A3 (finding 3, BLOCKING — accepted, with a second classifier bought). A gate that withholds the
primary may not rest on one unreproducible lead judgment. F4's classification is therefore done
twice, independently, on the full set, not a sample: once by the lead against the exemplars now
frozen below, and once by an independent non-panel model that takes no coding seat, given the
same rubric, the same exemplars and the WHY lines alone. If the two classifications disagree on
more than 25% of lines, F4 does not fire and the confound is reported unresolved (as in A1).
Both classifications and every WHY line are published verbatim.
Frozen exemplars for the F4 rubric (the critic's "clumsy and unidiomatic" case is resolved
explicitly: it is error):
| category | fires on | exemplars |
|---|---|---|
register |
names a word, phrase, idiom, oath or address form and its level | "'mum', 'mates', 'thanks a lot' are colloquial" · "contractions throughout" · "'Christ almighty' is current profanity where the Italian is an oath" |
error |
names something wrong, mistranslated, omitted, inaccurate, clumsy or unidiomatic | "'handy' does not render 'bravi tiratori'" · "clumsy and unidiomatic" · "drops the honorific, losing information" |
other |
restates the scale (vacuous), or names something not in the span (mislocated), or neither of the above |
"much lower register" (no textual cause named) |
A4 (finding 6, NON-BLOCKING — accepted). The REPEAT control is underpowered at 2 sites (18
paired comparisons). A third repeat site, N2, is added — 27 paired comparisons, at the cost of
one extra version at one site. N2 keeps its own arm codes, so the anti-circularity control is
untouched.
A5 (finding 9, NON-BLOCKING — accepted). F3's degeneracy threshold drops from 8 of 12 to 6
of 12, and a variance check is added: a seat whose within-site variance across arms is < 0.25
on a question is flagged on that question and reported alongside.
A6 (finding 8, NON-BLOCKING — accepted). Pairwise inter-seat mean absolute difference is
reported per scale, as S118 did (EXPL 0.367, REG 0.522, POP 0.467 there), and a scale whose
inter-seat MAD exceeds 1.00 on a seven-point scale carries that figure in every sentence citing
it.
A7 (finding 10, NON-BLOCKING — accepted). F6 is decoupled from P2. P2's rationale
rests on V4 (withhold, do not repair) alone, which is the rule that bears on explicitness;
V3 (contract) is a separate rule about length and the two were run together in one sentence without
argument. F6 becomes a standalone descriptive check: was V3 executed at all? Already answered
by the machine limb before dispatch: V is the shortest English arm at 1.0446× against S 1.1851, D
1.1892, X 1.2054, so F6 does not fire — and V3's own target of ≤ 1.000 was still missed.
A8 (finding 5, BLOCKING — accepted). P5 is not a prediction and is relabelled. It has no
gate and cannot fail, because it is a registered interpretive rule: a pre-commitment to which
reading follows from which outcome, written down so neither can be storied afterwards. It is moved
out of §7 into §7a under that name and is not counted in the run's outcome.
A9 (finding 7 — accepted as a declared limit, not repaired). P5 deepseek/deepseek-v4-pro
produced arm P of E-20260806b, a rendering of this same passage. Arm P is not in this run,
so charter §5 is satisfied on its own terms; but the seat has rendered this locus and now codes four
other renderings of it, and that is a latent style-correlation confound. It is declared rather than
repaired: the alternatives cost four to fifteen times as much per call and the run's headroom is
$1.56. The result reports arm-V and arm-X figures per seat, so that a P5-specific effect would be
visible rather than pooled away.
REJECTED — finding 4 (BLOCKING). The critic reads the routing-margin row as the margined total and calls the declared worst case understated by $0.43. It is a line item added to the seat rows, not a replacement for them: $0.19 + $0.21 + $0.03 = $0.43; ×1.5 = $0.645; $0.645 − $0.43 = $0.215, the row as written. $0.645 + $0.10 critic + $0.20 retry = $0.945 ≈ $0.95, the declared figure. No error. (This is the second consecutive session in which a critic's arithmetic BLOCKING finding was rejected with two lines of algebra — recorded, because a critic that is right about design and wrong about sums is a pattern worth watching, not a reason to stop buying critics: findings 1, 2 and 3 above are worth many times the $0.032.)
Revised pre-flight. Coding 36 calls unchanged (the extra REPEAT version at N2 adds one
version to one site's payload, inside the existing caps): $0.645 margined. Plus the F4 second
classifier, 1 call, max_tokens 4,000, non-panel: $0.05. Plus $0.20 retry reserve. Revised
declared worst case $0.90 remaining, against $1.525762784 of headroom after the critic's
actual $0.032205.
A10 (mid-run, declared before it was applied). P5 deepseek/deepseek-v4-pro returned
finish_reason: length with zero content characters and ~9,000 tokens of hidden reasoning on two
consecutive attempts at site N2 — provider GMICloud then SiliconFlow — while the same
slug at provider DigitalOcean had returned a clean 16-line body for $0.00094221 at N1.
This is note (bhf) with the provider, not the cap, as the variable, and it is exactly S122's
finding. Per (bhf) rule (iii) the cap is not raised a second time; instead routing is pinned:
call.py now sends provider: {"only": ["DigitalOcean"], "allow_fallbacks": false} for that slug
alone. The run was stopped and restarted rather than allowed to spend two doomed attempts at
every remaining site; the 5 completed .done.json bodies are re-used without re-billing, which is
what that cache is for. The two failed bodies, $0.0129131, bought nothing and are ledgered as
waste.