Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260806g-berman-negative-pole/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260806g-berman-negative-pole
statusfrozen
created2026-08-06
updated2026-08-06
linkswiki/arms/ARM-berman-occurrence.md, wiki/base/sources/S-berman-tendances.md, wiki/findings/results/RS-20260806b-berman-occurrence.md, workshop/experiments/E-20260806b-berman-occurrence/design.md, workshop/regimes/R21-vulgarisation.md, workshop/translations/cavalleria-rusticana/R21-v1/translation.md, workshop/translations/cavalleria-rusticana/R08-v1/translation.md, config/models.md, wiki/goodness-senses.md
sensesstyle-correspondence, cultural-mediation, naturalness
internal-judgment-onlytrue
provisionaltrue

E-20260806g — Berman's downward pole: can a translator go the other way, and can a reader tell?

Frozen before dispatch of any call. The R21 rule set and T-cavalleria-rusticana-R21-v1's translator's log were frozen before this file was written, and the contamination measurement (§2.2) was run before any site was cut. Amendments after the pre-run critic pass are recorded in critic.md and marked A<n> here.

1. The question

Berman's fifth and sixth deforming tendencies are one axis with two poles: ennoblissement, the raising of a source's register, and vulgarisation, the false popular speech — « pseudo-argot » — a translator reaches for when the source speaks a vernacular his own language has no equivalent for. Both are deformations; Berman condemns them together.

RS-20260806b measured that axis on five English renderings of Verga's Sicilian and found the downward pole empty. Strettell 1893 at +0.833, Dole 1896 at +0.583, the foreignising control at +0.500, the unbriefed 2026 machine at −0.083. On a signed −3…+3 scale, over 186 cells, the negative half was used 29 times in all — and on the neighbouring EXPL and POP scales, 4 times and 9 times, which fired the run's own one-sidedness gate F1′ and withheld two of its three primaries.

That leaves two questions genuinely open, and they are the same question asked of the craft and of the reading:

Is there anywhere below Verga for an English translator to go — and if a translator goes there on purpose, does a reader register it as lower, or merely as different?

Subject-rule sentence (continue-prompt.md §4.5): this unit teaches whether the downward deformation a major tradition names is something a translator can actually perform on a popular source, what performing it costs the story, and whether the direction of a register movement is visible to a reader at all — or whether any departure from a source's register reads as departure upward. The objects measured are four English renderings of Verga and one newly written one. The F1′ re-test in P4 is a gate inside this unit, not its subject: the run would be worth doing if the scale's behaviour were already settled, because Berman's two-sided claim would still be untested on any corpus.

Why it is not a re-run. E-20260806b asked do the tendencies occur in what translators published. This asks is the catalogue's own second pole reachable, and it answers it by making the missing arm rather than by looking harder at the arms that existed. The new arm is the unit's translation limb; nothing about it was recoverable from the S118 corpus.

2. Materials

Source. As E-20260806b: Verga, «Cavalleria rusticana», Vita dei campi (1880; 1881 text, it.wikisource.org), ¶50–81 of 81, 740 Italian tokens. Public domain (Verga d. 1922). arms/IT.txt is byte-identical to the S118 file.

Four arms, at arms/*.txt:

arm rendering date role here
V the lead, R21 vulgarisation, T-cavalleria-rusticana-R21-v1 2026 the new arm; the primary rests on it. Rules frozen before translating, log frozen before this file
X the lead, R08 resistancy, T-cavalleria-rusticana-R08-v1 2026 the arm that read +0.500 on REG at S118 while counter-programmed — the confound this run tries to resolve
S Alma Strettell, 1893, two-scan verified 1893 the ennoblement reading being compared against (+0.833 at S118)
D Nathan Haskell Dole, 1896, Gutenberg #37979 1896 as S (+0.583 at S118)

Two arms of E-20260806b are deliberately dropped. L (the lead's R04 close rendering) is excluded because §2.2 measures V as dependent on it and on nothing else; keeping both would put two renderings by the same hand at 9 shared 12-grams into one comparison. P (the unbriefed deepseek-v4-pro rendering) is dropped so that P5 is free to take a coding seat — a model may not judge a text it produced (charter §5), and dropping the arm buys a seat that costs a twentieth of the one it replaces.

2.1 Copy-text

S and D are the S118 files unchanged, with that run's verification (two-scan diff, three resolved OCR defects, Strettell's paragraph division unrecoverable) carrying over verbatim as a declared limit. V was extracted from the frozen translation.md by analysis/extract_arm.py, which re-flows the hard-wrapped paragraphs and asserts a 31-paragraph alignment with the Italian; the assertion passes.

2.2 Contamination, measured before any site was cut (CLAUDE.md standing rule)

tools/dependence_check.py over all six English renderings of this locus, run after the R21 log was frozen and before this design existed (runs/dependence.json, cells in runs/dependence_cells.json):

pair shared 7-grams shared 12-grams longest run verdict
V ~ L — the new arm against the lead's own close rendering 54 9 17 DEPENDENT?
V ~ P 35 0 11 clean
V ~ D 13 0 10 clean
V ~ S 10 0 11 clean
V ~ X — two lead renderings, opposite register programmes 4 0 8 clean
S ~ D (the published pair, carried over) 46 11 17 DEPENDENT?
L ~ P (carried over) 205 90 30 DEPENDENT?

Four things follow and bind what this run may claim.

  1. V is clean against every arm the run uses. Zero shared 12-grams against S, D and X; the longest run against any of them is 11 tokens, below this project's own recall floor for a constrained passage (notes (bez), (bgf), (bgi)). Whatever V scores, it does not score it by resembling Strettell, Dole or the resistancy arm.
  2. V's single dependency is on the lead's own prior rendering, at 9 shared 12-grams and a 17-token run — higher than V's overlap with any published human. This is note (bhb) landing for the third time and it is why arm L is not in this run. The contamination: high declaration on T-cavalleria-rusticana-R21-v1 is this measurement, not a guess.
  3. V ~ X is the lowest-overlap pair in the whole table — 4 shared 7-grams, longest run 8. Two renderings by the same hand, of the same 740 words, in the same repository, under opposite register programmes, are further apart than two published human translators three years and one publishing world apart (46 / 11 / 17). Reported as a fact about what a regime change does to a lead rendering; no gate depends on it, and it is not offered as evidence that either arm is good.
  4. The run is not licensed to call any pair independent on the tool's verdict, which flags nine of fifteen pairs including one that is independent by construction. The zeros are the informative cells; the DEPENDENT? labels are not.

3. The three coded questions — unchanged, verbatim, from E-20260806b

EXPL, REG and POP are reproduced word for word from E-20260806b/build_payloads.py, including the signing (positive = Berman's predicted direction), the 0 is a real answer sentence, and the Use negative values wherever they apply sentence. Changing the instrument while testing what the instrument can register would make the result uninterpretable, so nothing in the prompt is touched except the number of versions.

The scales, for reading this file without the other open: EXPL +3 much more explicit / −3 leaves considerably more unsaid · REG +3 much higher / −3 much lower · POP +3 popular marks effaced almost entirely / −3 marked more strongly than the Italian.

4. Sites — the same twelve, nine theory-chosen and three systematically sampled

analysis/sites.json, carried over from E-20260806b with the V spans added by analysis/build_sites.py and machine-verified: every V span is a verbatim substring of arms/V.txt (12 of 12) and covers the same paragraph extent as the Italian span (12 of 12).

Keeping N1–N3, the paragraphs chosen at the 1/4, 2/4 and 3/4 positions without inspecting their content, preserves S118's anti-circularity control: A3 there was the amendment that changed the run, and dropping it now would quietly re-introduce the bias it was written against.

5. The machine limb (no API call, analysis/machine.py)

M1–M4 recomputed with V added, on the whole locus. Registered before dispatch this time, so unlike S118 (A8 there) M1 may carry a gate:

Registered: R21's V3 (contract) was measured at 1.0446× and MISSED — the frozen log records the miss. M1's role here is F6 below, and the "nobody wrote English shorter than the Italian" finding of S118 §6.1 is extended, not tested, by an arm that tried to.

6. Procedure

  1. V produced first, unbriefed by any other arm, log frozen (§2, and translation.md).
  2. Contamination measured (§2.2). Sites cut. This design frozen.
  3. Independent pre-run critic, one non-panel model, against this file. Amendments written here before dispatch.
  4. Per site the seat receives the Italian span and the English versions, labelled V1… in an order shuffled per site by md5(site_id), mapping in analysis/keymap.json, shown to no seat. Provenance, date, regime and authorship are withheld from every seat.
  5. REPEAT control: at S3 and S7, arm V is included twice under two labels (five versions at those sites). A seat coding identical text differently is measuring noise. (S118 put the duplicate on the arm its primary used; so does this one.) The raw md5 draw placed the twins adjacent at S7 and one apart at S3, which would have made the control trivial — a seat shown the identical text twice in a row is not being tested on reproducibility. A minimum separation of 2 is therefore imposed by deterministic rotation (build_payloads.py), giving separations of 2 and 4, which is exactly what S118's unconstrained draw happened to produce. The constraint makes the two runs' controls comparable; it does not raise the standard, and it was applied before any dispatch.
  6. Three seats, one call per (site, seat): P1 openai/gpt-5.6-terra · P3 x-ai/grok-4.5 · P5 deepseek/deepseek-v4-pro. 36 calls. P2 (google/gemini-3.6-flash) is dropped on this project's own record — it needed a 6,000-token cap at S118 for hidden reasoning and was 80% of the seat spend for 33% of the cells at S120 — and P5 replaces it, which is possible only because arm P was dropped. P1 and P3 are shared with S118, which is what licenses F5.
  7. Output format, line guard and ###END### terminator exactly as S118.
  8. Cell value = median of the three seats; arm means pool the twelve sites.

Judgment is not parallelised. The seats are coders of a source–target relation, not a jury. No goodness sense is scored, no quality claim is made, no arm is ranked against another, and Tier D is NOT PASSED — every sentence of the result is provisional.

7. Predictions, registered

8. Failure criteria, registered — what withholds what

If error + other exceeds one quarter of arm V's extreme REG codes, P1 is withheld — the coders would be reading badness rather than lowness. All WHY lines are published verbatim in the result's provenance directory so the classification can be re-done by anyone; the lead's classification is internal-judgment-only and the counts are reported with that flag on them. - F5 (licence, not withholding). Reproduction. If S or D's REG mean in this run differs from its S118 value by more than 0.75 scale points, the cross-run comparison is not licensed: V is then reported against this run's S and D only, and no sentence compares any figure here with a figure there. - F6. If M1 does not make V the shortest English arm, R21's V3 was not executed at all and the contraction half of P2's rationale is withdrawn (the withholding half stands). - F7. If arm V's POP mean and arm X's POP mean have the same sign and differ by less than 0.50, P3 is reported as inconclusive rather than confirmed: the two arms would be indistinguishable on the scale that is supposed to separate their programmes.

9. Seats, cost, pre-flight

Worst case built from max_tokens, never from an assumed output length — note (abc), the one estimate this project has overrun.

stage seat calls max_tokens list in/out per M worst case
0 — independent pre-run critic one non-panel model, nvidia/nemotron-3-ultra-550b-a55b 1 16,000 — $0.10
2 — coding, P1 openai/gpt-5.6-terra 12 2,500 $1.00 / $6.00 $0.19
2 — coding, P3 x-ai/grok-4.5 12 2,500 $2.00 / $6.00 $0.21
2 — coding, P5 deepseek/deepseek-v4-pro 12 2,500 $0.435 / $0.87 $0.03
— routing margin ×1.5 on stage 2 $0.22
— retry reserve $0.20

Declared worst case $0.95. Today's ledger (2026-08-06 UTC) stands at $3.442032216 of $5.00 across seven sessions; headroom $1.557967784. The run fits with $0.61 to spare.

Seat choice is a cost decision made from this project's own records, and it is stated because it is also a design change: P2 is dropped and P5 added, so two of three seats are shared with S118 and F5 is a convergent reproduction check across a partly-changed panel rather than a strict one. The result reports it that way.

Reserve (note (bfc), declared before dispatch): qwen/qwen3.7-max, which produced no arm in either experiment. Timeout 300s, halved from the inherited 900s per S122's finding.

Lead translation of arm V cost $0 and is not ledgered (charter §3, A4).


Amendments after the pre-run critic pass

The independent pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel, takes no coding seat, $0.032205, provider Together, 108.6s) returned NEEDS-AMENDMENT, 10 findings, 5 BLOCKING. Verbatim body in critic.md. Eight are accepted, one is accepted as a declared limit, and one BLOCKING finding is rejected with the arithmetic. No coding call had been dispatched when these were written.

A1 (finding 1, BLOCKING — accepted; the best finding in the pass). F4 as registered could not fire. The prompt asks for one WHY line per version, naming what drove the coder's most extreme of the three codes — so if a coder's most extreme code on arm V is EXPL or POP, its WHY line says nothing about REG, and F4 has no evidence for exactly the codes it must classify. The critic's first repair — three WHY lines per version — is refused, because §3 fixes the seat's task verbatim from S118 and tripling the free-text output would change both the seat's behaviour and the truncation risk that note (bhf) has fired on nine times. The second repair is taken and tightened:

A2 (finding 2, BLOCKING — accepted, and resolved at $0 with data already on disk). POP cannot by itself separate pseudo-argot (Berman's tendency 6) from authentic vernacular preserved (tendencies 10–12): both would push the code negative. A fourth scale is refused for the reason in A1. The repair is that M4 is registered as the discriminator, and it is objective and already computed: arm V retains zero source-culture tokens, the lowest count of any arm this project has measured (S 3, D 22, X 27). So:

A3 (finding 3, BLOCKING — accepted, with a second classifier bought). A gate that withholds the primary may not rest on one unreproducible lead judgment. F4's classification is therefore done twice, independently, on the full set, not a sample: once by the lead against the exemplars now frozen below, and once by an independent non-panel model that takes no coding seat, given the same rubric, the same exemplars and the WHY lines alone. If the two classifications disagree on more than 25% of lines, F4 does not fire and the confound is reported unresolved (as in A1). Both classifications and every WHY line are published verbatim.

Frozen exemplars for the F4 rubric (the critic's "clumsy and unidiomatic" case is resolved explicitly: it is error):

category fires on exemplars
register names a word, phrase, idiom, oath or address form and its level "'mum', 'mates', 'thanks a lot' are colloquial" · "contractions throughout" · "'Christ almighty' is current profanity where the Italian is an oath"
error names something wrong, mistranslated, omitted, inaccurate, clumsy or unidiomatic "'handy' does not render 'bravi tiratori'" · "clumsy and unidiomatic" · "drops the honorific, losing information"
other restates the scale (vacuous), or names something not in the span (mislocated), or neither of the above "much lower register" (no textual cause named)

A4 (finding 6, NON-BLOCKING — accepted). The REPEAT control is underpowered at 2 sites (18 paired comparisons). A third repeat site, N2, is added — 27 paired comparisons, at the cost of one extra version at one site. N2 keeps its own arm codes, so the anti-circularity control is untouched.

A5 (finding 9, NON-BLOCKING — accepted). F3's degeneracy threshold drops from 8 of 12 to 6 of 12, and a variance check is added: a seat whose within-site variance across arms is < 0.25 on a question is flagged on that question and reported alongside.

A6 (finding 8, NON-BLOCKING — accepted). Pairwise inter-seat mean absolute difference is reported per scale, as S118 did (EXPL 0.367, REG 0.522, POP 0.467 there), and a scale whose inter-seat MAD exceeds 1.00 on a seven-point scale carries that figure in every sentence citing it.

A7 (finding 10, NON-BLOCKING — accepted). F6 is decoupled from P2. P2's rationale rests on V4 (withhold, do not repair) alone, which is the rule that bears on explicitness; V3 (contract) is a separate rule about length and the two were run together in one sentence without argument. F6 becomes a standalone descriptive check: was V3 executed at all? Already answered by the machine limb before dispatch: V is the shortest English arm at 1.0446× against S 1.1851, D 1.1892, X 1.2054, so F6 does not fire — and V3's own target of ≤ 1.000 was still missed.

A8 (finding 5, BLOCKING — accepted). P5 is not a prediction and is relabelled. It has no gate and cannot fail, because it is a registered interpretive rule: a pre-commitment to which reading follows from which outcome, written down so neither can be storied afterwards. It is moved out of §7 into §7a under that name and is not counted in the run's outcome.

A9 (finding 7 — accepted as a declared limit, not repaired). P5 deepseek/deepseek-v4-pro produced arm P of E-20260806b, a rendering of this same passage. Arm P is not in this run, so charter §5 is satisfied on its own terms; but the seat has rendered this locus and now codes four other renderings of it, and that is a latent style-correlation confound. It is declared rather than repaired: the alternatives cost four to fifteen times as much per call and the run's headroom is $1.56. The result reports arm-V and arm-X figures per seat, so that a P5-specific effect would be visible rather than pooled away.

REJECTED — finding 4 (BLOCKING). The critic reads the routing-margin row as the margined total and calls the declared worst case understated by $0.43. It is a line item added to the seat rows, not a replacement for them: $0.19 + $0.21 + $0.03 = $0.43; ×1.5 = $0.645; $0.645 − $0.43 = $0.215, the row as written. $0.645 + $0.10 critic + $0.20 retry = $0.945 ≈ $0.95, the declared figure. No error. (This is the second consecutive session in which a critic's arithmetic BLOCKING finding was rejected with two lines of algebra — recorded, because a critic that is right about design and wrong about sums is a pattern worth watching, not a reason to stop buying critics: findings 1, 2 and 3 above are worth many times the $0.032.)

Revised pre-flight. Coding 36 calls unchanged (the extra REPEAT version at N2 adds one version to one site's payload, inside the existing caps): $0.645 margined. Plus the F4 second classifier, 1 call, max_tokens 4,000, non-panel: $0.05. Plus $0.20 retry reserve. Revised declared worst case $0.90 remaining, against $1.525762784 of headroom after the critic's actual $0.032205.

A10 (mid-run, declared before it was applied). P5 deepseek/deepseek-v4-pro returned finish_reason: length with zero content characters and ~9,000 tokens of hidden reasoning on two consecutive attempts at site N2 — provider GMICloud then SiliconFlow — while the same slug at provider DigitalOcean had returned a clean 16-line body for $0.00094221 at N1. This is note (bhf) with the provider, not the cap, as the variable, and it is exactly S122's finding. Per (bhf) rule (iii) the cap is not raised a second time; instead routing is pinned: call.py now sends provider: {"only": ["DigitalOcean"], "allow_fallbacks": false} for that slug alone. The run was stopped and restarted rather than allowed to spend two doomed attempts at every remaining site; the 5 completed .done.json bodies are re-used without re-billing, which is what that cache is for. The two failed bodies, $0.0129131, bought nothing and are ledgered as waste.