Repository path: wiki/findings/results/RS-20260728e-venuti-specifiability.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260728e-venuti-specifiability |
| status | active |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| links | workshop/experiments/E-20260728e-venuti-specifiability/design.md, workshop/regimes/R07-fluency.md, workshop/regimes/R08-resistancy.md, workshop/translations/osso-di-morto/R07-v1/translation.md, workshop/translations/osso-di-morto/R08-v1/translation.md, wiki/base/sources/S-venuti-invisibility.md, wiki/method-notes.md, config/models.md |
| senses | naturalness, style-correspondence, voice, cultural-mediation |
| internal-judgment-only | true |
| provisional | true |
RS-20260728e — the domestication↔foreignization axis is not measurable by rule coverage, and the reason is that two readers do not agree on what a rule decides
Design frozen at 9846fbd before either translation existed. Five predictions registered; four failed. Then the design's own reliability check fired, and it changes which numbers may be reported at all.
1. The headline
| measure | lead's coding | independent classifier's coding |
|---|---|---|
cov(R07) — fluency rules' decision rate |
6/44 = 0.136 | 15/44 = 0.341 |
cov(R08) — resistancy rules' decision rate |
39/45 = 0.867 | 18/45 = 0.400 |
Agreement between the two codings is 0.483 — below the 0.60 threshold the design pre-registered as its reliability failure criterion. Cohen's κ = 0.173 (p₀ 0.483, pₑ 0.375), which is the range conventionally called slight: the two readers are barely above chance on a three-way classification, applying the same written definitions to the same 89 sites.
Under the frozen rule, the classifier's numbers become the headline and the lead's are reported as unreliable. But the more useful reading is that neither set of numbers is a measurement. If two careful readers cannot agree on whether a written rule decided a given choice, then "how far along the domestication↔foreignization axis is this translation" is not a quantity that this route can deliver — and that is a stronger and more robust conclusion than either coverage rate.
2. Predictions, as registered
| prediction | lead's coding | classifier's coding | verdict | |
|---|---|---|---|---|
| P1 | cov(R07) > cov(R08) |
0.136 < 0.867 | 0.341 < 0.400 | FAILS on both codings |
| P2 | cov(R08) < 0.50 |
0.867 | 0.400 | fails on the lead's, holds on the classifier's |
| P3 | sep > 0.50 |
0.951 | — | holds (0.927 excluding the one orthography-only difference) |
| P4 | int(R08) ≥ D(R08) |
1 vs 39 | 1 vs 18 | FAILS on both |
| P5 | R07 longer than R08 | 373 vs 408 words | — | FAILS |
Four of five failed, and P1 failed on both codings. P1 was the session's central expectation and it was built on Venuti's own statement that his method "cannot be calculated before the translation process is begun." The prediction was that the fluency pole would be the better-specified one because he states it as a checklist. It is not. On the lead's coding the resistancy rules decide six times as often; on the classifier's they still decide more often.
3. Why P1 failed, and it is not a fact about Venuti
The mechanism is visible in the rule text and was written into the R08 log while translating, before any of this was computed:
A rule that says do what the source does decides almost any site, because the source has already decided it. R2 ("follow the source's clause order and period length") and R6 ("calque where the source's figure is not English") are cited at 34 of the 39 sites the lead coded D under R08. The fluency set has no such rule: F6 ("idiomatic syntax before close syntax") and F1 ("current, not archaic") exclude options without naming a target, which is why they leave two live options standing at 19 sites.
So the asymmetry is not foreignization is better specified than domestication. It is foreignization can be written as a pointer to an object that exists, and domestication cannot. "Be idiomatic" names no text; "do what the Italian does" names the Italian. This inverts the shape of the Tymoczko objection the experiment set out to test: the criteria that are missing are the domesticating ones.
4. The disagreement is not noise — it is one specific dispute, and it is the same dispute
| lead → classifier | count |
|---|---|
| D → D | 20 |
| D → S | 14 |
| D → P | 11 |
| P → D | 12 |
| P → P | 21 |
| P → S | 3 |
| S → D | 1 |
| S → P | 5 |
| S → S | 2 |
All 14 D→S disagreements are in R08, and 10 of the 14 cite only R2 and/or R6. The dispute is exactly one question: does a general "follow the source" rule name the feature at a site — a doublet, a plural, a causal conjunction — or does it merely apply to everything and therefore name nothing? The lead answered yes; the classifier answered no, consistently, fourteen times.
That is a defensible disagreement, and no wording in the frozen spec settles it. The coding scheme has a hole where its most consequential distinction should be, and both readers fell into it in opposite directions. In the other direction the classifier is more willing than the lead to call an exclusion-rule decisive (12 P→D and 1 S→D, almost all in R07) — so it is not simply a stricter reader; it is applying a different, unwritten theory of what "names the feature" means.
A second, smaller hole is on record from before the classifier ran. R08 site 36 («rotella» → "knee-wheel") is a rule conflict: R6 requires the calque, R9 forbids the parodic, and no live option satisfies both. The frozen D/P/S scheme has no code for that, so the site was recorded as P and the gap flagged in the log. The classifier — which was told to code conflicts as P! — coded it D and flagged no conflict anywhere in 89 sites.
5. What did work
sep = 0.951. At 39 of 41 shared sites the two renderings differ in wording; the two agree only on "a certain Federico M." and on the decision to repeat the framing phrase. The two rule sets do produce two texts, so the design did not fail into one translation written twice, and the coverage finding is about two real conditions.- The length result is clean and goes against the project. The resistant rendering is the longer one — 408 words to the fluent 373, against a 362-word source. The project has spent two sessions on the fact that its jury prefers longer texts (
RS-20260728c), and the working assumption in that work was that fluency's supplied subjects and connectives are what inflate length. Here the opposite holds: R5 ("do not supply what the source withholds") removes three subject pronouns, and R2/R3/R6 add far more than that back by keeping doublets, litotes and periphrases the fluent version compresses. Any future evaluation contrasting a domesticating with a foreignizing rendering must treat length as a live confound in the foreignizing direction. - R10 never fired — no site was decided by "abusive fidelity," the rule the regime's own limitation note predicted would absorb anything. R1 fired once. The resistancy rule set was assembled from Venuti's practice on a fragmenting modernist poet; on periodic 1869 prose, three of its ten rules were inert.
6. What this changes in the project
S-venuti-invisibility's standing recommendation is withdrawn. That page has said since S006 that the project should import domestication/foreignization "as a scored axis / pair of tendencies, never as a binary label." The scored-axis half is now unsupported by the project's own attempt to build the scale: two readers with the same rules and the same text return κ = 0.17. The page is amended to say the pair is usable as descriptive vocabulary for a strategy a translator declares, and not as a quantity a third party scores off the text.- It does not license the opposite claim. n = 1 passage, one language pair, one translator, one classifier, one pair of rule sets assembled by the party being tested. This is a case demonstration with a number attached, exactly as the design says.
- The target-culture half of Venuti's theory was never in the design. For Venuti foreignization is "a strategic construction whose value is contingent on the current target-language situation" (ch. 1, p. 20). A rule set operating on the source-to-target mapping cannot see that, so this is evidence about the discursive half only — and a defender of the axis can say, correctly, that the half that was tested is the half Venuti says is not where the strategy lives.
7. Verification and provenance
- 46 automated checks, 0 failures (
analyse.py): every logged site accounted for by the alignment; both logs' numbering contiguous; every quoted rendering in the alignment table verified to occur in the corresponding translation body. The logs' own hand-written tally lines were recomputed, not trusted — the R07 line said 7/28/9 and was wrong; the table says 6/30/8, and the file now records the correction. - Classifier.
openai/gpt-5.6-terra(P1), providerOpenAI,finish_reason: stop, in 6,726 / out 5,494 (4,840 reasoning),temperature: 0. $0.051713, against a $0.105 worst case built frommax_tokens6000 at list out-price plus the prompt at list in-price (note (abc)) — 49% of worst case, the highest fraction in five sessions. No fall-through was needed. The prompt was built by stripping the code and rules columns out of the two logs programmatically, so the lead's assignments could not leak; the classifier saw sites, live options and chosen wordings only. - The key-usage cross-check returned a delta of exactly 0.000000 against a per-request cost of $0.051713, on two snapshots plus a third taken later in the session. This is a new failure mode: the previous three inexact cross-checks were gaps attributed to concurrent non-project use, and this is the endpoint not moving at all. The per-request figure is what is ledgered, per
config/budget.md's stated method. Note (bcw) records it. - No pre-run critic pass, in declared departure from
continue-prompt.md§6; the design says so and gives the reason (§What this cannot establish). Given that the session's reliability check then fired, the departure looks worse in hindsight than it did in prospect and is recorded that way — a critic might have found the D/P/S hole in §4 before the codes were assigned rather than after.
Every judgment in this page that is not a computed number is internal-judgment-only; the classifier is a non-Anthropic panel model and the panel is not calibrated (config/models.md), so its coding carries no evidential authority either — it is a second reader, not an arbiter. provisional: true per charter §2.4.