Repository path: workshop/experiments/E-20260806b-berman-occurrence/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260806b-berman-occurrence |
| status | frozen |
| created | 2026-08-06 |
| updated | 2026-08-06 |
| links | wiki/arms/ARM-berman-occurrence.md, wiki/base/sources/S-berman-tendances.md, workshop/translations/cavalleria-rusticana/R04-v1/translation.md, workshop/translations/cavalleria-rusticana/R08-v1/translation.md, workshop/translations/cavalleria-rusticana/R06-v1/translation.md, config/models.md, wiki/goodness-senses.md |
| senses | naturalness, style-correspondence, cultural-mediation, accuracy |
| internal-judgment-only | true |
| provisional | true |
E-20260806b — do Berman's deforming tendencies occur?
Frozen before dispatch of any coding call. Amendments after the pre-run critic pass are recorded
in critic.md and marked A<n> here.
1. The question
S-berman-tendances §What this cannot ground, written S025 and never acted on, says:
It is not evidence that the twelve tendencies occur. Berman asserts them and illustrates each with one or two examples chosen to show it. There is no corpus, no counting, no control. Cited as a taxonomy, never as a measurement.
This experiment makes the measurement the source lacks, on one passage, in one pair.
The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5): this unit teaches
whether the deformations a major tradition says translation inflicts on literary prose are actually
present in published literary translations, and whether they belong to translation as such or to the
period that produced them. The objects measured are four English renderings of an Italian story;
none of the project's own statistics, raters or published figures is the subject.
Berman's own restriction is the second half of the question. He claims universals — « des universaux de la déformation inhérents au traduire comme tel » — and says the classical belles infidèles merely coincided with them, « cette coïncidence est momentanée ». If the tendencies are universals, a 2026 rendering should carry them as a 1893 rendering does. If they are the norms of Anglophone translating in the 1890s, it should not.
2. Materials
Source. Giovanni Verga, «Cavalleria rusticana», Vita dei campi (1880; 1881 text from
it.wikisource.org, retrieved 2026-08-06). Verga d. 1922 — public domain. Locus ¶50–81 of 81,
740 Italian tokens (744 whitespace-separated words): Alfio's return, Santa's denunciation, the kiss of challenge, the duel.
Stored at arms/IT.txt.
Five arms, all stored at arms/*.txt:
| arm | rendering | date | provenance |
|---|---|---|---|
| S | Alma Strettell, Cavalleria Rusticana and Other Tales of Sicilian Peasant Life (Unwin, Pseudonym Library) | 1893 | archive.org scan, two-scan verified (see §2.1); 877 words |
| D | Nathan Haskell Dole, Under the Shadow of Etna (Joseph Knight, 1896) | 1896 | Project Gutenberg #37979, proofread text; 870 words |
| P | panel seat P5 (deepseek/deepseek-v4-pro), single pass, unbriefed |
2026 | run_translate.py; the prompt names no theory, no other rendering and no style |
| L | the lead, R04 (close, source-only), T-cavalleria-rusticana-R04-v1 |
2026 | log frozen before this design was written |
| X | the lead, R08 (resistancy), T-cavalleria-rusticana-R08-v1 |
2026 | the registered negative control; R08 was frozen at S048 from Venuti's account of his own practice, not written for this run |
All four English translators of record are out of copyright (Strettell fl. 1890s; Dole d. 1935; both volumes pre-1929).
2.1 Copy-text verification (gate, done before the design was written)
Strettell's text exists only as OCR. Two independent archive.org scans were fetched and diffed
word by word. All 54 differences were quote-mark OCR, hyphenation across the narrow column, or
scan-2 "Digitized by Google" boilerplate, except three, each resolved against the second scan
and recorded here: soldz,t → soldi; fora → for a; "lama dead man → "I am a dead
man. Two publisher footnotes were removed as apparatus, not translation. Strettell's paragraph
division is not recoverable from either scan and is not used; the arm is presented as continuous
prose. This is a declared limit and it is why tendency 1 (rationalisation) is measured at the
sentence level and not the paragraph level.
2.2 Contamination, measured before the design (CLAUDE.md standing rule)
tools/dependence_check.py, all four English arms pairwise, after the two lead logs were frozen
(runs/dependence.json):
| pair | shared 7-grams | shared 12-grams | longest run | tool verdict |
|---|---|---|---|---|
| L ~ P (two 2026 machine renderings, different labs) | 205 | 90 | 30 tokens | DEPENDENT? |
| L ~ S | 80 | 23 | 22 tokens | DEPENDENT? |
| P ~ S | 62 | 19 | 22 tokens | DEPENDENT? |
| L ~ D | 77 | 16 | 17 tokens | DEPENDENT? |
| P ~ D | 74 | 14 | 18 tokens | DEPENDENT? |
| S ~ D (two published translators, 1893 vs 1896) | 46 | 11 | 17 tokens | DEPENDENT? |
| L ~ X | 75 | 9 | 15 tokens | DEPENDENT? |
| P ~ X | 38 | 1 | 12 tokens | DEPENDENT? |
| X ~ S | 18 | 0 | 9 tokens | clean |
| X ~ D | 20 | 0 | 10 tokens | clean |
(The P rows were added after arm P was produced and before any coding call was dispatched; nothing below was changed to fit them.)
Four things follow and are binding on what this run may claim.
- Arm L is contaminated and is excluded from the primary. It sits above the published pair's own figure against Strettell. The standing rule — the lead is never the independent third translator where a design's validity turns on independence from any published rendering — applies, and the period-versus-universal contrast (P3 below) is therefore registered on arm P alone. L is reported as a labelled arm because its translator's log is on record, which is what P4 needs, and for nothing else.
- The negative control is clean at 0 shared 12-grams against both published arms. Whatever X scores, it does not score it by resembling Strettell or Dole.
- The published pair is not demonstrably independent of each other — 11 shared 12-grams, a 17-token run, three years apart, both in London/Boston literary publishing. Dole may have known Strettell. This is a limit on P1, which pools them: two arms that share a source of English are less than two independent observations. Reported, not repaired.
- The tool's verdict does not discriminate on this material, and the run must not lean on it.
Nine of ten pairs are flagged
DEPENDENT?, including the one pair that is independent by construction if any pair is (S ~ D). On a 740-token passage this constrained, a close rendering produces long shared runs whoever makes it — which is the recall-floor finding notes (bez), (bgf) and (bgi) already established, met again here. What the numbers therefore support is a ranking, not a verdict: L and P sit above the published pair's own figure, X sits below it at zero. And the direction of the residual worry runs against P3, not for it — if arm P is soaked in Victorian English it will read more like Strettell and Dole, which is the direction that makes P3 fail. Contamination on P is therefore the conservative error here, and that is why P3 is registered on P despite it. - The largest overlap in the table is between the two machine renderings — 90 shared 12-grams and a 30-token run between L and P, eight times the published pair's figure, across two different labs with no shared prompt. This is not a nuisance parameter of this design; it is a fact about machine literary translation that the design did not set out to measure, and the result page reports it as such.
3. The three coded questions
Berman's own catalogue is, by S-berman-tendances's reading, about five independent operations
wearing twelve names (2 and 3 are corollaries of 1; 7 overlaps 1; 10, 11 and 12 are one domain at
three grain sizes; 5 and 6 are a pair). The coding follows that reading rather than the numbering.
Two of the operations are measured by machine at $0 (§5); three are put to seats, each on a
signed −3…+3 scale so that a null and a reversal are both expressible, and each signed so that
positive is Berman's predicted direction:
- EXPL — clarification and expansion (tendencies 2, 3). "Compared with the Italian at this site, does the English make explicit something the Italian leaves unsaid or indefinite — a subject, an object, an antecedent, a relation, a reason? +3 much more explicit · 0 leaves as much unsaid as the Italian · −3 leaves considerably more unsaid than the Italian."
- REG — ennoblement and vulgarisation (tendencies 5, 6). "Compared with the Italian at this site, is the English pitched higher — more literary, more polished, more decorous — or the same, or lower — plainer, rougher, more common? +3 much higher · 0 same · −3 much lower."
- POP — destruction of vernacular networks, of idioms, and effacement of superimposed languages (tendencies 10, 11, 12). "The Italian here carries marks of popular Sicilian speech: forms of address, oaths, proverbs, set idioms, dialect words. Compared with the Italian, does the English efface them — replacing them with ordinary standard English or with an English idiom of its own — or keep them, or mark them more strongly than the Italian does? +3 effaced almost entirely · 0 kept as they are · −3 marked more strongly than the Italian."
Each seat also returns, per version, one WHY line of ≤ 25 words naming what drove its most extreme of the three codes. The WHY lines are evidence, not decoration: a code with no locatable cause in the text is what the verification pass looks for.
4. The nine sites
Cut in analysis/sites.json, anchored per arm and machine-verified to align (45 spans, 0 failures).
| id | what the Italian does there |
|---|---|
| S1 | «vi adorna la casa» — Santa's venomous idiom, first half of a network closed at S9 |
| S2 | «Santo diavolone! … non vi lascierò gli occhi per piangere! a voi e a tutto il vostro parentado!» — an oath English does not have, an idiom, and a social unit (parentado) |
| S3 | «— Va bene … grazie tante.» — eight words of lethal understatement |
| S4 | «adesso che era tornato il gatto … smaltiva l'uggia all'osteria» — an inverted proverb and a dialect-flavoured verb |
| S5 | «Avete comandi da darmi? / Nessuna preghiera» — a fixed courtesy formula, and a polysemy (preghiera = request and prayer) |
| S6 | «tutte le avemarie che potevano capirvi» — a construction whose antecedent the Italian does not supply |
| S7 | «la mia vecchia … la mia vecchierella» + «come un cane» — a diminutive gradient English has no morphology for |
| S8 | «bravi tiratori» (the weapon unnamed) + «come la rese, la rese buona» (repetition) |
| S9 | «E tre! … la casa che tu m'hai adornato … cadde come un masso» — the network closing, and a simile |
5. The machine limb (no API call; analysis/machine.py)
Registered before the coding run, computed on the whole locus, not the sites:
- M1 — expansion (tendency 3): English tokens ÷ 740 Italian tokens, per arm.
- M2 — rationalisation at the sentence level (tendency 1): sentences per 100 words, and the arm's sentence count against the Italian's.
- M3 — quantitative impoverishment (tendency 4) and destruction of networks (tendency 9): type–token ratio; and how many distinct English words each arm uses for the two occurrences of «adornare» and the three of «rendere». Berman's network claim predicts more than one.
- M4 — effacement of the vernacular (tendencies 11, 12): count of retained Italian-culture tokens
(
compare/compar,gnà,fra,Canziria,fichidindia,Gesummaria,avemarie,soldi,osteria,Pasqua,diavolone) per arm. This is an objective index of one of the three coded questions and is the cross-check on POP.
6. Procedure
- Arm P produced first, unbriefed, before this design was frozen (§2). Arms L and X and their logs frozen before that.
- Per site, the seat receives the Italian span and the five English versions, labelled
V1…V5in an order shuffled per site by a deterministic seed (md5(site_id)), the mapping stored inanalysis/keymap.jsonand not shown to any seat. - REPEAT control: at S3 and S7 arm L is included twice, under two different version labels (six versions at those sites). A seat that codes the identical text differently is measuring noise.
- Three seats, one call per (site, seat): P1
openai/gpt-5.6-terra, P2google/gemini-3.6-flash, P3x-ai/grok-4.5. 27 calls. No seat produced any arm (charter §5); arm P came from P5, which takes no coding seat. - Output format is strict lines
V<n>|EXPL|<int>,V<n>|REG|<int>,V<n>|POP|<int>,V<n>|WHY|<text>, terminated by###END###. The runner enforces the line count and the terminator as seat failures (notes (b), (bgw)). - Cell value = median of the three seats. Arm means pool the nine sites.
Judgment is not parallelised. The seats are used here as coders of a source–target relation,
not as a jury: nothing in this run is a quality judgment, no sense is scored, and Tier D is NOT
PASSED, so every sentence of the result is provisional.
7. Predictions, registered
- P1 — occurrence (primary). Pooled over nine sites and three seats, the two published human arms S and D score > 0 on all three questions, and ≥ +0.75 on at least two of the three, with the same sign in each arm taken separately.
- P2 — the negative control (gate). Arm X scores < 0 on all three questions and ≤ −0.75 on at least two.
- P3 — period or universal. Arm P (2026, unbriefed, no knowledge of this project) scores lower than the mean of S and D by ≥ 0.50 on at least two of the three questions. A null here — P level with the Victorians — is a result for Berman's universality claim, and is to be reported as such rather than as a failure.
- P4 — the translator's log (descriptive). At S8 the lead's frozen log records supplying "with the knife", a noun the Italian withholds; at S6 it records refusing the parallel supply; at S7 it records losing the diminutive. Arm L's EXPL code should be positive at S8, at or below 0 at S6, and its POP code positive at S7. This is not a test of Berman; it is a test of whether a translator's own account of where he deformed is recoverable by a reader of the English.
8. Failure criteria, registered — what withholds what
- F1 (gate). If X does not score negative on at least two of three questions, the scale cannot register the anti-Berman direction, and P1 and P3 are withheld entirely.
- F2 (gate). REPEAT: if the mean absolute difference between the two arm-L labels at S3 and S7, over the three questions and three seats, exceeds 0.50 scale points, all primaries are withheld as coder noise.
- F3. If fewer than two of three seats return a usable code in more than 20% of (site, arm, question) cells, primaries are withheld.
- F4. A seat that returns the identical value for every arm at 7 or more of 9 sites on a
question is degenerate on that question; its codes are excluded from that question and the
exclusion is reported in the headline. (This has fired before:
RS-20260806-same-manhad a seat return the same integer on every text it saw.) - F5. If S and D disagree in sign on a question, P1 is not claimed for that question — a tendency present in one Victorian and absent in the other is not a universal.
- F6. If M4 (machine, objective) and POP (coded) disagree in their arm ordering, the POP result is reported as unsupported by its own objective cross-check.
9. Seats, cost, pre-flight
| stage | seats | calls | max_tokens | worst case | reserve |
|---|---|---|---|---|---|
| 0 — independent pre-run critic | one non-panel model | 1 | 16,000 | $0.30 | — |
| 1 — arm P | P5, reserve kimi-k3 | 1–2 | 8,000 | $0.15 | kimi-k3 |
| 2 — coding | P1, P2, P3 | 27 | 2,500 | $0.85 | P5 |
| — retries | — | — | — | $0.25 | — |
Declared worst case $1.55. Today's ledger (2026-08-06 UTC) stood at $0.614681164 of $5.00
before this session; headroom $4.385. The estimate is built from max_tokens, not from an
assumed output length (note (abc)), with a 1.5× routing margin on the coding stage (note on P5
routing, config/models.md). Lead translation of three arms cost $0 and is not ledgered.
Amendments after the pre-run critic pass
The independent pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel, $0.0214474) returned
NEEDS-REDESIGN, 15 findings, 6 BLOCKING. Verbatim body in critic.md. Eleven amendments were
accepted, one finding was rejected with evidence, and three were accepted as declared limits that
cannot be repaired inside this run. No coding call had been dispatched when these were written.
A1 (finding 2, BLOCKING — accepted). F1 as registered was incoherent: it withheld the primary
if the resistancy arm X failed to score negative — but if Berman is right that these are
universals of translating as such, X should score positive too, so the gate would have fired
exactly when the hypothesis was supported. F1 is replaced by F1′, which tests the instrument
and not the hypothesis: if the negative half of a scale is used in fewer than 5% of all coded
cells on that question, that scale is one-sided and every directional claim on it is withheld.
X's own sign is now evidence, not a gate — X scoring positive is reported as support for
universality.
A2 (finding 3, BLOCKING — accepted). REG merges ennoblement and vulgarisation, which
Berman lists separately, so an arm doing both at different sites averages to zero and reads as
having neither. The scale is kept (splitting it would double the output per call and the failure
rate with it), but P1 on REG is now claimable on either of two grounds — mean ≥ +0.75 or ≥ 8 of
12 site-medians positive — and the site-level sign distribution is reported for every arm so that
a both-directions pattern is visible instead of cancelled.
A3 (finding 4, BLOCKING — accepted; this is the amendment that changed the run). The nine sites
were chosen because the Italian presents Berman's material there, so P1 was close to guaranteed.
Three further sites, N1–N3, are added by systematic sample: of the 31 locus paragraphs, the
20 not touched by S1–S9 were listed in order and the paragraphs at the 1/4, 2/4 and 3/4
positions taken — ¶12, ¶17, ¶25 — with no inspection of their content before selection. New
registered prediction P5: if the theory-chosen sites score materially higher than the
systematically-sampled ones (a gap ≥ 0.75 on two of three questions), P1 is reported as an
opportunity-site result only and not as a claim about the passage. Note what the sample already
shows and what it does not: all three sampled paragraphs turned out to contain vernacular markers
too, which is a fact about Verga's prose and is not evidence that the selection was unbiased.
The run is now 12 sites × 3 seats = 36 calls.
A4 (finding 7, NON-BLOCKING — accepted). The REPEAT duplicate is moved from arm L, which
§2.2 excludes from the primary, to arm P, which the primary uses. F2 now measures noise on an
arm whose codes are load-bearing.
A5 (finding 11 — accepted). F5 relaxed: a sign disagreement between S and D withholds P1 on
that question only if both arms' magnitudes exceed 0.5. Below that, the disagreement is reported
as inconclusive rather than treated as a gate.
A6 (finding 12 — accepted). The WHY verification is now a defined step, not a gesture. Every
WHY line attached to a code of |value| ≥ 2 is read against that arm's span for that site, and
recorded as located (it names something actually in that span), mislocated (it names something
that is not there, or that is in a different arm's span), or vacuous (it restates the scale). The
counts are reported in the result, and a question on which more than a quarter of extreme codes are
mislocated or vacuous carries that figure in every sentence that cites it.
A7 (finding 14 — accepted). P4 is relabelled a method-validation check, not a prediction about Berman, and is not pooled with P1–P3 or counted in the run's outcome.
A8 (finding 6 — accepted as a limit, not as a repair). The critic asked for registered
predictions on the machine limb M1–M4. They cannot honestly be registered: M1–M4 were computed
before this amendment was written. They are therefore declared descriptive and exploratory,
they are excluded from the run's registered outcome, and no gate depends on them. F6 (the M4/POP
cross-check) survives only as a reported disagreement, never as a withholding rule. (One
defect the machine limb found in itself is recorded because it nearly produced a false positive: the
sentence splitter did not split on a terminator followed by a closing quote, which made the
quotation-mark arms S, D and P appear to have 26–29 sentences against the Italian's 41 — a textbook
"rationalisation" result that was an artefact of punctuation style. Corrected before the coding run;
all five arms then read 40–43.)
A9 (findings 1 and 5, BLOCKING — accepted, and NOT repairable in this run). P3 cannot separate period from species. Arm P is a machine and arms S and D are humans, so any P-versus-Victorian difference is confounded between 2026 versus 1890s and machine versus human. The obvious repair — a third human arm from a later period — was attempted: D. H. Lawrence's 1928 rendering was searched for on archive.org and Project Gutenberg and is not freely reachable (charter §7, A8 forbids buying it and this project does not ask Tom for a text it can work around). P3 is therefore reframed and demoted: it is reported as period-or-species, the two are named as inseparable on this evidence, and no claim about Berman's universality thesis is drawn from it in either direction. What P3 can still do is falsify a strong reading — if arm P scores level with the Victorians, then whatever makes the Victorians deform is not something a 2026 machine avoids.
A10 (finding 15 — accepted as a declared limit). 12 sites × 5 arms × 3 seats is small. The run is powered for large effects only; the result states that a difference below roughly one scale point is not distinguishable from noise here, and no magnitude below that is reported as a size.
A11 (finding 13 — accepted, trivially). Tier D is defined in PROJECT.md §5 and its state in
config/models.md; the pointer is added. The critic had no access to either.
REJECTED — finding 10. The critic asserts that gpt-5.6-terra, gemini-3.6-flash, grok-4.5,
deepseek-v4-pro and kimi-k3 "do not exist as of the design date" and that the run "cannot be
executed as specified". This is the critic's training horizon, not a defect in the design: all five
slugs are recorded in config/models.md with prices read from the API, and arm P of this very
experiment was produced by deepseek/deepseek-v4-pro before the critic was called, billed at
$0.00270396 via StreamLake, with the raw body on disk at
runs/P_translate_deepseek-v4-pro_try1.raw. No substitution is made.
Findings 8 and 9 (NON-BLOCKING) — partially accepted. F4 is kept but demoted to a reported
diagnostic rather than an automatic exclusion, with inter-seat agreement reported alongside it;
M4 is additionally reported per site as well as over the whole locus.
Revised pre-flight. Coding is now 36 calls (12 sites × 3 seats) at max_tokens 2,500:
worst case $0.98 with the 1.5× routing margin, plus $0.25 retry reserve. Revised declared
worst case for the session: $1.71, against $4.385 of headroom at session start and
$0.0241514 already spent (arm P $0.00270396, the killed kimi attempt $0.049833, critic $0.0214474 —
$0.073985 actual so far).
A12 (mid-run, declared before it was applied). P2 (google/gemini-3.6-flash) returned
finish_reason: length twice at max_tokens 2,500 — it carries heavy hidden reasoning
(config/models.md). Its cap is raised to 6,000; P1 and P3 are unchanged at 2,500. In the same
patch the stage reserve was changed off deepseek/deepseek-v4-pro, which is the model that
produced arm P: a fall-through to it would have had a model code its own translation, against
charter §5. The reserve is now qwen/qwen3.7-max, which produced nothing in this experiment. The
four bodies already cached when this was found were audited: all three seats, no reserve body among
them. Revised worst case for the coding stage: $1.42.