Repository path: workshop/experiments/E-20260802f-licensed-strangeness/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260802f-licensed-strangeness |
| status | frozen |
| created | 2026-08-02 |
| updated | 2026-08-02 |
| senses | naturalness, style-correspondence, cultural-mediation, accuracy |
| internal-judgment-only | true |
| provisional | true |
| purpose | Evaluation-design study. Reader assumed: a source-blind juror of the kind charter §5 gives this project. Utility: to decide whether a proposed repair to the `naturalness` sense is scoreable at all. |
| links | wiki/arms/ARM-typology-derivation.md, wiki/goodness-senses.md, wiki/findings/sense-dossier.md, wiki/base/sources/S-venuti-invisibility.md, wiki/base/anchors/A-homer-sarpedon/A-homer-sarpedon.md, wiki/findings/results/RS-20260802b-pole-uniformity.md, workshop/translations/iliad-sarpedon/R07-v1/translation.md, config/models.md, config/budget.md |
E-20260802f — can a source-blind reader tell licensed strangeness from incompetence?
ARM-typology-derivation step 3, remaining half. The step owes the
naturalness-against-Venuti motion. This is the experiment that decides which motion to open.
1. The question, and why it is not the one the arm expected
wiki/goodness-senses.md defines naturalness as:
The translation reads as fluent, idiomatic prose of the target language; nothing rings as translationese unless the source rings strange in the same place.
The bolded clause is the sense's entire defence against the Venuti challenge recorded on the same
entry — that the wording is Nida's "complete naturalness of expression" and that ch. 1 of The
Translator's Invisibility is an argument against it (S-venuti-invisibility). The clause says: we
do not penalise strangeness that the source put there.
Reading A-homer-sarpedon closely enough to build this experiment showed that the clause protects
a narrower phenomenon than the challenge is about. The clause protects mimetic strangeness —
target-language strangeness matched to source-language strangeness in the same place. Venuti's
foreignization need not be mimetic: it is pressure on target-language values to register the
source's difference from the target, and that difference is often located where the source is
ordinary in its own language. Close-corsleted and wheat-producing are strange in English; whether
they register anything strange in Greek is the question.
And that question is contested in this very anchor, by two of its own six translators, in print
(amendment A6). Arnold's position is that Homer is noble, rapid and not quaint, so Newman's
compounds are English oddity registering nothing; Newman's is that Homer was archaic and peculiar
to later Greek ears, so his compounds are mimetic after all and the clause protects them. Whichever
is right, the consequence for the sense is the same and is worse than either party's: the clause's
applicability at a site turns on a contested judgement about the source's own markedness — a
judgement the project's jury cannot make, because it is not given the source, and which two
professional Hellenists could not settle between them in three published volumes. Under Arnold's
reading, every one of Newman's nine foreignizing cells of seventeen is an unprotected naturalness
failure (RS-20260802b §3); under Newman's, none of them is. The project's instrument would
return Arnold's verdict, for Arnold's reasons, without knowing it had taken a side.
So the motion's real question is not virtue or norm. It is: what should the clause condition
on, if not source-strangeness? The obvious repair is to condition on licence — strangeness
that registers a feature of the source the target lacks is not penalised — and the obvious
objection to the repair is that the licence is invisible in the scored object. A jury scores the
English. The licence lives in the Greek. If licensed strangeness and unforced bad English are
indistinguishable from the target text alone, the repair is unscoreable by the source-blind jury
this project's charter gives it (charter §5, and wiki/goodness-senses.md voice: "scored by
readers who cannot read the source").
That is what this experiment measures, and it is the fact that selects the motion's options.
The unit's one-sentence wire. The translation limb builds four renderings of one passage whose
strangeness is licensed, unlicensed, or absent; the study limb's blind rating decides whether the
proposed repair to naturalness — condition on licence rather than on source-strangeness — is
scoreable at all, and therefore which option the decision page carries as its provisional default.
Subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5): this unit teaches whether
the criterion that protects source-motivated strangeness in a literary translation can be applied by
a reader of the translation, or whether "reads naturally" is unconditionally a demand of the
receiving language. That is about evaluating translations, not about this project's apparatus.
2. Materials
Source. Homer, Iliad XII.310–327, Sarpedon to Glaucus. Greek from the Monro–Allen OCT via
Perseus, already stored at E-20260802b-pole-uniformity/materials/source-grc.json. Public domain.
Base rendering. T-iliad-sarpedon-R07-v1, the lead's fluency rendering, frozen at S088 before
any published English of this locus had been read, with a 20-decision log naming the live options
at every site. It is arm A verbatim and is not edited.
Four arms, all filed at materials/arms.md and frozen by commit before dispatch:
| arm | what it is | edits |
|---|---|---|
| A — FLUENT | T-iliad-sarpedon-R07-v1 verbatim |
0 |
| B — LICENSED | A, with strangeness introduced at five sites chosen by the record, not by the lead's taste | 5 |
| C — UNLICENSED | A, with strangeness of the same five kinds introduced at five stretches where the Greek is plain and nothing is at stake | 5 |
| D — POSITIVE | A, with three Greek words left untranslated (temenos, pepon, Kēres) | 3 |
How B's five sites were chosen, and why this matters. They are the sites of
E-20260802b-pole-uniformity's twenty-one where the six published translators, across 287 years and
two opposed manifestos, converged on the domesticating option — counted over all three
independent codings (lead, P1, P2), 18 cells each:
| site | Greek | feature | D cells of 18 |
|---|---|---|---|
| S06 | τέμενος | the granted precinct as an institution | 18 |
| S07 | νεμόμεσθα | graze/possess polysemy | 17 |
| S13 | κοιρανέουσιν … Λυκίην κάτα | the κάτα + accusative government | 16 |
| S15 | ὦ πέπον | the address form ripe / mellow one | 16 |
| S17 | anacoluthon at 322–324 | the broken construction | 16 |
These are the places RS-20260802b §6 called sites where the axis has no live foreignizing option
at all in practice. B tests that inference directly: if the lead can write English that
registers the feature at all five, then the unanimity is a norm being observed, not a limit being
hit — which is Venuti's claim, on this project's own material.
C's five stretches are matched to B's five in kind — a coined compound, a doubled
unresolved construction, an odd prepositional government, a calque-looking epithet, a broken
construction — and are placed where the Greek is plain and no site in E-20260802b's book falls.
No C edit changes the propositional content: C is unmotivated style damage, not an accuracy
error, because an accuracy error would be a different finding.
S17 is form-constrained in E-20260802b's sense — verse metre constrained four of the six
published renderings there. It is not constrained for the lead's prose arms. Declared, not hidden.
3. Seats and conditions
Panel roles per config/models.md; all non-Anthropic (charter §5, D-20260723-03).
- Pre-run critic: P2
google/gemini-3.6-flash— a seat taking no part in the rating stage. - Rating seats: P1
openai/gpt-5.6-terra, P3x-ai/grok-4.5, P5deepseek/deepseek-v4-pro. Three seats, stateless calls, never shown this design, the hypothesis, the arm labels, or which points belong to which arm.
Two conditions, same three seats, six calls:
- Condition N (no source) — the four renderings only. This is the project's actual jury situation.
- Condition S (source given) — the same four renderings plus the Greek text of XII.310–327, line-numbered. Dispatched after every condition-N body has returned, so nothing can leak backward; each call is stateless, so nothing leaks forward either.
Arm order is fixed and identical in both conditions: T1 = C, T2 = A, T3 = D, T4 = B.
Order is therefore confounded with nothing that varies, and cancels in P3's difference-in-difference.
There is no counterbalancing rotation — declared as a limit, not glossed: any order effect on
P1 and P2 within condition N is unmeasured.
What each seat is asked, per call:
- For each of the four renderings, a naturalness rating 1–5 on the wording of the project's own sense, supplied verbatim minus its conditional clause (so the seat is not handed the hypothesis): "reads as fluent, idiomatic prose of the target language; nothing rings as translationese."
- For each of the 13 marked points (5 in B, 5 in C, 3 in D), presented shuffled in a frozen
order with its arm identified only as
T1–T4:O— this reads as a deliberate attempt to carry over something the original does that English does not — orF— this reads as an unforced failure of the translator's English, with nothing behind it. (Wording per amendmentA5: the earlier form asked a source-blind seat to assert what the original contains, which it cannot see. Identical in both conditions, so the conditions differ only in source access.)
4. Predictions, registered
Let ATTR(x) = proportion of arm x's marked points coded O, pooled over three seats.
Let NAT(x) = mean naturalness rating of arm x, pooled over three seats.
Condition N:
- P1 — distinguishability.
ATTR(B) ≥ 0.60andATTR(B) − ATTR(C) ≥ 0.30. - P2 — the clause is inert on the score.
|NAT(B) − NAT(C)| ≤ 0.50, with both belowNAT(A).
Condition S:
- P3 — source access rescues the clause.
[NAT(B) − NAT(C)]_S − [NAT(B) − NAT(C)]_N ≥ 0.50.
Failure criteria, which withhold rather than reinterpret:
- FC1 — instrument check. If
NAT(A)is not the highest ofNAT(A),NAT(B),NAT(C)in condition N, the naturalness item is not measuring naturalness and P2 and P3 are withheld. - FC2 — positive control. If
ATTR(D) < 0.80in condition N, the attribution item cannot fire even where the original is physically present in the English, and P1 is withheld. (AmendmentA5: passingFC2establishes only that the item can fire where the Greek word is physically in the English. It does not establish that attribution works on syntax, which is where three of B's five points live. The control is easy by construction and its claim is correspondingly narrow.) - FC3 — matching check, and it is the one that can kill this design. B and C must be matched in
edit count (5 = 5, by construction) and in size. If the count of words changed between A and B
and between A and C differs by more than 30% of the larger, the arms are not matched and
P2 is withheld — because an unmatched pair confounds licence with amount of damage.
Computed by
analysis/score.pyfrom the frozen arms, and reported whatever it says. - FC4 — reproducibility. If the three seats do not agree on the sign of
ATTR(B) − ATTR(C), P1 is reported as unresolved and nothing is built on it.
5. What each outcome licenses, written before the run
This is the part that selects the motion, and it is fixed here so that the decision page's provisional default is not chosen after seeing the numbers.
| outcome | what it means | motion's provisional default |
|---|---|---|
| P1 fails | licensed and unforced strangeness are indistinguishable from the English alone | the licence repair is unscoreable by a source-blind jury → default D (score naturalness only against a declared purpose) |
| P1 holds, P2 holds | readers can see the licence and penalise the strangeness anyway | Venuti's structural point, measured on the project's own instrument → default C (split floor from assimilation) |
P1 holds, P2 fails with NAT(B) − NAT(C) > 0.50 |
the clause is already doing work without the source | → default B (rewrite the clause to name licence rather than source-strangeness) |
P1 holds, P2 fails with NAT(C) − NAT(B) > 0.50 |
licensed strangeness reads as less natural than matched damage with nothing behind it — the strongest form of the Venuti point | the licence repair cannot help → default D |
| P3 holds | the clause is operable only with source access, which this jury does not have | recorded on whichever default the above selects; it is a finding about jury design as much as about the sense |
(The two-sided split is amendment A4: P2 is an absolute-value criterion and §5 originally mapped
both of its failure directions to one motion. That was a logic error and the critic found it.)
Scope, narrowed by amendment A6. This run tests five manufactured edits on one nineteen-line
speech in one language pair. It settles nothing about foreignization theory in general, and nothing
about any translator's practice. What it can settle is whether this project's naturalness
clause can be applied by this project's source-blind jury.
No outcome licenses a change to wiki/goodness-senses.md this session. A motion is opened here
and ratified by a later session (charter §8). The session that opens a motion never ratifies it.
6. Verification
analysis/verify.py recomputes every number reported on the result page from the raw stored bodies,
imports nothing from tools/, and is checked by mutation tests each of which asserts that the bytes
on disk changed (note (bgu)) and restores every mutated file (note (bhd)).
7. Cost
Pre-flight, built from max_tokens and the worst plausible provider, not from an expected
answer length (note (abc); note (bhq), which fired at S091 for exactly the opposite reason —
seven of eighteen bodies spent their whole budget on hidden reasoning and returned nothing).
| item | seat | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic | P2 | 24,000 | $0.19 |
| condition N × 3 | P1, P3, P5 | 16,000 | $0.38 |
| condition S × 3 | P1, P3, P5 | 16,000 | $0.38 |
| retry margin (note (bhq)) | $0.25 | ||
| declared worst case | $1.20 |
Headroom at session open: $2.118063511 of the $5.00 UTC-day cap, six sessions already spent today. The run fits. All translation, all arm construction, all scoring and all verification are lead work at $0 (charter §3, A4).