Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260807f-carriage-or-strangeness/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260807f-carriage-or-strangeness
statusfrozen
created2026-08-07
updated2026-08-07
trackT2
sensesstyle-correspondence, perceived-source-carriage, accuracy, naturalness
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-sense-overlap.md, wiki/goodness-senses.md, workshop/translations/monelle-paroles/R04-v1/translation.md, workshop/translations/monelle-paroles/device-census.md, workshop/translations/monelle-paroles/contamination.md, workshop/experiments/E-20260807f-carriage-or-strangeness/materials/arms.md, workshop/experiments/E-20260802f-licensed-strangeness/design.md, wiki/findings/results/RS-20260802f-licensed-strangeness.md, wiki/findings/results/RS-20260807e-sense-tradeoff.md, config/models.md, workshop/regimes/R04-lead-close.md, workshop/regimes/R14-matched-flattening.md

E-20260807f — is perceived-source-carriage about the source, or about the strangeness?

Frozen before dispatch. No API call had been made when this page was first committed. Amendments made after the critic pass are numbered A<n> at the foot and dated; nothing above them is edited in place.

1. Question

wiki/goodness-senses.md carries two senses that both go up when a translation is pulled toward its source, and it has never asked whether they are two things:

RS-20260807e (S129) is why the question is live now: the foreignizing arm came out highest of seven arms on both, 6.000 and 6.722, while accuracy did not move at all. The two senses moved together and nothing in that design could separate them.

The question, in one sentence: does a reader's sense that a translation is carrying its source over track the source's actual formal features, or does it track markedness in the English?

Subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5): this unit teaches whether the impression that a translation "brings the reader to the author" is bought by carrying the author's forms or bought by being odd — which is the empirical content of the foreignizing claim, and which decides whether two of the eight senses of good are one sense under two names. That is about evaluating translations, not about this project's apparatus.

What is new against E-20260802f (S092), which is the predecessor. That run manipulated licensed against unlicensed strangeness at site level, on Homeric verse, and measured attribution of individual departures: real licences attributed at 15 of 15, manufactured oddity at 0.60 without the source and 0.40 with it. perceived-source-carriage did not exist as a sense then — it was created out of that run by D-20260802-13. This run asks the question in the sense's own scoring form (a 1–7 whole-rendering score), crossed with a second factor the predecessor never had (whether the source's forms are carried at all), on prose, in a new language pair, alongside style-correspondence on the same items.

2. Materials

Source. Marcel Schwob, «Paroles de Monelle», from Le Livre de Monelle (Paris, 1894) — the four consecutive litanies de la destruction / de la formation / des dieux / des moments. Public domain, Project Gutenberg #53374, read in French. 500 words, segmented into six segments (../../translations/monelle-paroles/source-fr-segments.txt). The project's first work in French symbolist prose and its first Schwob.

Why this passage. It is a text whose difficulty is almost entirely formal. Its eighteen catalogued devices (../../translations/monelle-paroles/device-census.md) — anaphora, jussive chains, a six-fold epigram template each member of which is a figura etymologica, verset paragraphing, unreduced keyword chains — can each be removed without changing what a single sentence asserts. A passage where form and content can be varied independently is what a 2 × 2 on this question needs, and this project has not had one.

Base rendering. T-monelle-paroles-R04-v1, the lead's close translation, with R06-v1 frozen first per R04 §2a. Both committed at fcdd9c2 before any comparator was read and before this design existed.

Contamination: suspected, measured, with a declared priming event — ../../translations/monelle-paroles/contamination.md. Nine fragmentary OCR lines of the Meloney 1929 English were seen before translating; measured overlap 27 shared 7-grams, 5 twelve-grams, longest run 16 tokens, against two independent published pairs at 6 / 0 / 10 and 23 / 6 / 17. The design's validity does not turn on independence from Meloney — every contrast is within-lead and matched — and §7 carries the one place it does bite.

The four arms (materials/arms.md, frozen by commit before dispatch), a 2 × 2:

arm FORM STRANGE words
CC carried none 519
FC flattened none 591
CS carried added 515
FS flattened added 581

CC is the R04 translation verbatim. FC applies eighteen flattening operators, one per catalogued device, in the spirit of R14. CS applies eighteen unlicensed markedness edits — six odd prepositional governments, six coined hyphenated compounds, six inversions — placed where the source is formally plain and of kinds that answer to nothing in the French. A calque would be source-licensed and is therefore excluded by construction. FS applies the same eighteen edits at the same propositional sites in the flattened wording.

Nothing in the 36 edits changes what the text says. That is the binding rule and gate G1 checks it.

3. Seats and blinding

Three seats, non-Anthropic, the Tier D jurors of S020/S034/S086 so the instrument is extended rather than replaced (config/models.md panel v1): J1 = P1, J2 = P2, J3 = P5. Slugs logged in runs/ as provenance.

Two stages, because the two senses are asked under different information by their own definitions.

Stage B's source-absent condition is required by perceived-source-carriage's own entry (the score must state whether jurors had the source text) and is the condition in which the sense is most at risk of being pure markedness-attribution — which is what this run is for. Stage A is dispatched first; Stage B's prompt never mentions that a source exists.

4. The senses, quoted verbatim to the seats

The strings sent are the first paragraph of each entry in wiki/goodness-senses.md, with the naturalness string in its revised, target-text-only form (D-20260802-13; RS-20260805e requires that a design say which string it used) and its register anchor named as unmarked / literary-contemporary.

5. Gates, checked in this order, all registered before dispatch

6. Registered predictions

Every one of these is registered in the direction that would flatter the two senses. The lead's expectation is that all four hold.

Descriptive, not a prediction: the item-level Pearson correlation between the two senses across all 24 items, reported with its factor decomposition. A value ≥ 0.70 with both factors varying is reported as the two senses are behaving as one on this material, and is not a threshold anything turns on.

7. Known confounds and limits, declared before the run

  1. Length is confounded with FORM. Flattening supplies the connectives English wants and the flattened arms are 12–14% longer (591 and 581 against 519 and 515). Every effect attributed to FORM is also an effect of 70 words. Reported, not repaired; the analysis reports the four cell means so a length reading stays available.
  2. Explicitation. The flattening operator supplies causal connectives (since, because) where the French is asyndetic. The relations are implied by the source, but making them explicit is the nearest the operator comes to touching content, and G1 is the check on it.
  3. The lead wrote all four arms, chose the passage, catalogued the devices and registered the predictions. G5 puts the device catalogue outside the lead; nothing puts the edits outside it.
  4. One passage, one author, one language pair, one hand. Nothing here generalises past French symbolist prose without replication.
  5. The contamination bite, from contamination.md §4: to the extent CC is remembered Meloney, the CARRIED pole is less purely the lead's craft than the design's language implies. Bounded at five twelve-grams in 519 words.
  6. Tier D is NOT PASSED and perceived-source-carriage has never been through Tier D at all (config/models.md). Every sentence of the result is provisional.

8. Failure criteria

9. Analysis

analyse.py computes every reported number from runs/*.json and nothing is transcribed by hand. verify.py imports nothing from analyse.py, recomputes every figure independently from the raw bodies, re-derives the item→arm mapping from arms.md rather than from labels.json, runs the G4 leak screen and the G1 parity arithmetic, and includes mutation tests that corrupt a stored score and assert the checks fire.

Exact permutation tests where the cell counts permit: the FORM and STRANGE main effects are tested by exhaustively relabelling the 24 items' factor assignment within seat, reporting the exact two-sided P.

10. Pre-flight budget

UTC day 2026-08-07; config/budget.md shows $2.783760945 of $5.00 spent across S125–S129, so headroom is $2.216239055. Declared ceiling for this run: $0.90.

stage calls max_tokens worst case
0 — pre-run critic (non-panel) 1 16,000 $0.06
G5 — device census (non-panel, French only) 1 4,000 $0.02
A — source present, 3 seats × 24 items 3 10,000 $0.30
B — source absent, 3 seats × 24 items 3 10,000 $0.26
parity check (non-panel, 24 pairs) 1 6,000 $0.04
re-dispatch allowance (F2) ≤ 3 14,000 $0.20
total $0.88

Worst cases are built from max_tokens at the most expensive plausible provider, per note (abc). Snapshots either side; per-response usage.cost is primary and the key delta is the cross-check.

Amendments

Pre-run critic: nvidia/nemotron-3-ultra-550b-a55b (non-panel), stop, 22,122 characters, $0.0297342, provider Together. Verdict NEEDS-REDESIGN, 8 BLOCKING. Six accepted in whole or in part, two remedies overruled with reasons. No scoring call had been dispatched.

Re-registered gates and predictions (superseding §5 and §6 where they conflict)