Repository path: workshop/experiments/E-20260807f-carriage-or-strangeness/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260807f-carriage-or-strangeness |
| status | frozen |
| created | 2026-08-07 |
| updated | 2026-08-07 |
| track | T2 |
| senses | style-correspondence, perceived-source-carriage, accuracy, naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-sense-overlap.md, wiki/goodness-senses.md, workshop/translations/monelle-paroles/R04-v1/translation.md, workshop/translations/monelle-paroles/device-census.md, workshop/translations/monelle-paroles/contamination.md, workshop/experiments/E-20260807f-carriage-or-strangeness/materials/arms.md, workshop/experiments/E-20260802f-licensed-strangeness/design.md, wiki/findings/results/RS-20260802f-licensed-strangeness.md, wiki/findings/results/RS-20260807e-sense-tradeoff.md, config/models.md, workshop/regimes/R04-lead-close.md, workshop/regimes/R14-matched-flattening.md |
E-20260807f — is perceived-source-carriage about the source, or about the strangeness?
Frozen before dispatch. No API call had been made when this page was first committed. Amendments
made after the critic pass are numbered A<n> at the foot and dated; nothing above them is edited in
place.
1. Question
wiki/goodness-senses.md carries two senses that both go up when a translation is pulled toward its
source, and it has never asked whether they are two things:
style-correspondence— marked formal features of the source receive functional equivalents rather than silent flattening. A verified source–target relation.perceived-source-carriage— a marked departure from target-language norm reads as a deliberate attempt to carry over a feature or effect of the source. A reader's attribution.
RS-20260807e (S129) is why the question is live now: the foreignizing arm came out highest of
seven arms on both, 6.000 and 6.722, while accuracy did not move at all. The two senses moved
together and nothing in that design could separate them.
The question, in one sentence: does a reader's sense that a translation is carrying its source over track the source's actual formal features, or does it track markedness in the English?
Subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5): this unit teaches whether
the impression that a translation "brings the reader to the author" is bought by carrying the
author's forms or bought by being odd — which is the empirical content of the foreignizing claim, and
which decides whether two of the eight senses of good are one sense under two names. That is about
evaluating translations, not about this project's apparatus.
What is new against E-20260802f (S092), which is the predecessor. That run manipulated licensed
against unlicensed strangeness at site level, on Homeric verse, and measured attribution of
individual departures: real licences attributed at 15 of 15, manufactured oddity at 0.60 without the
source and 0.40 with it. perceived-source-carriage did not exist as a sense then — it was created
out of that run by D-20260802-13. This run asks the question in the sense's own scoring form
(a 1–7 whole-rendering score), crossed with a second factor the predecessor never had (whether
the source's forms are carried at all), on prose, in a new language pair, alongside
style-correspondence on the same items.
2. Materials
Source. Marcel Schwob, «Paroles de Monelle», from Le Livre de Monelle (Paris, 1894) — the four
consecutive litanies de la destruction / de la formation / des dieux / des moments. Public
domain, Project Gutenberg #53374, read in French. 500 words, segmented into six segments
(../../translations/monelle-paroles/source-fr-segments.txt). The project's first work in French
symbolist prose and its first Schwob.
Why this passage. It is a text whose difficulty is almost entirely formal. Its eighteen
catalogued devices (../../translations/monelle-paroles/device-census.md) — anaphora, jussive
chains, a six-fold epigram template each member of which is a figura etymologica, verset
paragraphing, unreduced keyword chains — can each be removed without changing what a single sentence
asserts. A passage where form and content can be varied independently is what a 2 × 2 on this
question needs, and this project has not had one.
Base rendering. T-monelle-paroles-R04-v1, the lead's close translation, with R06-v1 frozen
first per R04 §2a. Both committed at fcdd9c2 before any comparator was read and before this
design existed.
Contamination: suspected, measured, with a declared priming event —
../../translations/monelle-paroles/contamination.md. Nine fragmentary OCR lines of the Meloney 1929
English were seen before translating; measured overlap 27 shared 7-grams, 5 twelve-grams, longest run
16 tokens, against two independent published pairs at 6 / 0 / 10 and 23 / 6 / 17. The design's
validity does not turn on independence from Meloney — every contrast is within-lead and matched —
and §7 carries the one place it does bite.
The four arms (materials/arms.md, frozen by commit before dispatch), a 2 × 2:
| arm | FORM |
STRANGE |
words |
|---|---|---|---|
| CC | carried | none | 519 |
| FC | flattened | none | 591 |
| CS | carried | added | 515 |
| FS | flattened | added | 581 |
CC is the R04 translation verbatim. FC applies eighteen flattening operators, one per
catalogued device, in the spirit of R14. CS applies eighteen unlicensed markedness edits —
six odd prepositional governments, six coined hyphenated compounds, six inversions — placed where the
source is formally plain and of kinds that answer to nothing in the French. A calque would be
source-licensed and is therefore excluded by construction. FS applies the same eighteen edits at
the same propositional sites in the flattened wording.
Nothing in the 36 edits changes what the text says. That is the binding rule and gate G1
checks it.
3. Seats and blinding
Three seats, non-Anthropic, the Tier D jurors of S020/S034/S086 so the instrument is extended rather
than replaced (config/models.md panel v1): J1 = P1, J2 = P2, J3 = P5. Slugs logged in runs/
as provenance.
- 24 items, ids
I01–I24, assigned bysha256of the (segment, arm) pair so the id order carries no information about either factor.materials/labels.jsonis never sent anywhere. - Order shuffled per seat under a fixed seed, with no two items of the same segment adjacent.
- No arm labels, no regime names, no mention that arms, factors or an experiment exist. Seats are told the set contains renderings from different sources, in no order, and that the same French may appear more than once.
Two stages, because the two senses are asked under different information by their own definitions.
- Stage A — source present. Each seat sees the French segment and one English rendering, and
scores
style-correspondenceandaccuracy, integers 1–7. - Stage B — source absent. Each seat sees the English alone and scores
perceived-source-carriageandnaturalness, integers 1–7.
Stage B's source-absent condition is required by perceived-source-carriage's own entry (the score
must state whether jurors had the source text) and is the condition in which the sense is most at
risk of being pure markedness-attribution — which is what this run is for. Stage A is dispatched
first; Stage B's prompt never mentions that a source exists.
4. The senses, quoted verbatim to the seats
The strings sent are the first paragraph of each entry in wiki/goodness-senses.md, with the
naturalness string in its revised, target-text-only form (D-20260802-13; RS-20260805e
requires that a design say which string it used) and its register anchor named as
unmarked / literary-contemporary.
5. Gates, checked in this order, all registered before dispatch
G1— accuracy parity.max − minof the four arm means onaccuracy(Stage A) must be ≤ 1.00. Above that, the arms are not propositionally matched, every downstream contrast is confounded with content, and the primary is withheld.G2— theSTRANGEmanipulation took. Blindnaturalness(Stage B), pooledmean(CC, FC) − mean(CS, FS)must be ≥ 1.00. Below that, nothing about strangeness is licensed by this run.G3— theFORMmanipulation took.style-correspondence(Stage A), pooledmean(CC, CS) − mean(FC, FS)must be ≥ 1.00. Below that, either the flattening operator did not remove the devices or the sense cannot see them, and both readings are reported; the primary is withheld.G4— leak screen. No item text may contain an arm name, a device name, a translator's note, or the words arm, rendering set, experiment. Mechanical, inverify.py.G5— the device census is not the lead's word alone. One non-panel seat sees the six French segments and no English at all and is asked to name each segment's marked formal features. The independent list must cover ≥ 60% of the eighteen device families, scored per segment and pooled, by a coding rule fixed here: a device counts as covered if the independent answer names the same formal property of the same segment, whatever vocabulary it uses. Below 60%, theFORMfactor is the lead's stipulation and the result says so.
6. Registered predictions
Every one of these is registered in the direction that would flatter the two senses. The lead's expectation is that all four hold.
P1—style-correspondenceis discriminant. ItsFORMeffectmean(CC, CS) − mean(FC, FS)is ≥ 1.00, and itsSTRANGEeffect|mean(CS, FS) − mean(CC, FC)|is < 0.75. Fails if unlicensed strangeness moves the sense that is defined on source form by 0.75 or more.P2—perceived-source-carriageresponds to the source. ItsFORMeffect is ≥ 1.00. Fails if carrying eighteen of the source's own formal devices does not raise the sense that is named after carrying the source over.P3—perceived-source-carriageresponds to strangeness. ItsSTRANGEeffect is ≥ 1.00. Expected fromRS-20260802f's 0.40–0.60 attribution of manufactured oddity.P4— the decisive cell comparison.mean(CC) − mean(FS)onperceived-source-carriageis ≥ 1.00: a rendering that carries the source's forms and adds no gratuitous oddity reads as more source-carrying than one that carries none of them and is merely odd. Fails, and the failure is the finding, ifFS ≥ CC— the sense would then be bought by strangeness alone.
Descriptive, not a prediction: the item-level Pearson correlation between the two senses across all 24 items, reported with its factor decomposition. A value ≥ 0.70 with both factors varying is reported as the two senses are behaving as one on this material, and is not a threshold anything turns on.
7. Known confounds and limits, declared before the run
- Length is confounded with
FORM. Flattening supplies the connectives English wants and the flattened arms are 12–14% longer (591 and 581 against 519 and 515). Every effect attributed toFORMis also an effect of 70 words. Reported, not repaired; the analysis reports the four cell means so a length reading stays available. - Explicitation. The flattening operator supplies causal connectives (since, because) where
the French is asyndetic. The relations are implied by the source, but making them explicit is the
nearest the operator comes to touching content, and
G1is the check on it. - The lead wrote all four arms, chose the passage, catalogued the devices and registered the
predictions.
G5puts the device catalogue outside the lead; nothing puts the edits outside it. - One passage, one author, one language pair, one hand. Nothing here generalises past French symbolist prose without replication.
- The contamination bite, from
contamination.md§4: to the extentCCis remembered Meloney, theCARRIEDpole is less purely the lead's craft than the design's language implies. Bounded at five twelve-grams in 519 words. - Tier D is NOT PASSED and
perceived-source-carriagehas never been through Tier D at all (config/models.md). Every sentence of the result isprovisional.
8. Failure criteria
F1— any gateG1–G3fails → the primary (P2,P4) is withheld and the run reports what its gates measured.F2— any seat returns fewer than 24 rows in a stage → that block is re-dispatched once; if it fails again the seat is reported as incomplete and every figure is recomputed without it.F3—G5below 60% → theFORMfactor is reported as lead-defined andP1's discriminant claim is downgraded to a statement about the lead's own catalogue.F4— ifaccuracyandstyle-correspondencecorrelate above 0.90 across items, Stage A is reported as a single judgment wearing two names andP1is withheld.
9. Analysis
analyse.py computes every reported number from runs/*.json and nothing is transcribed by hand.
verify.py imports nothing from analyse.py, recomputes every figure independently from the raw
bodies, re-derives the item→arm mapping from arms.md rather than from labels.json, runs the G4
leak screen and the G1 parity arithmetic, and includes mutation tests that corrupt a stored score
and assert the checks fire.
Exact permutation tests where the cell counts permit: the FORM and STRANGE main effects are
tested by exhaustively relabelling the 24 items' factor assignment within seat, reporting the exact
two-sided P.
10. Pre-flight budget
UTC day 2026-08-07; config/budget.md shows $2.783760945 of $5.00 spent across S125–S129, so
headroom is $2.216239055. Declared ceiling for this run: $0.90.
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| 0 — pre-run critic (non-panel) | 1 | 16,000 | $0.06 |
| G5 — device census (non-panel, French only) | 1 | 4,000 | $0.02 |
| A — source present, 3 seats × 24 items | 3 | 10,000 | $0.30 |
| B — source absent, 3 seats × 24 items | 3 | 10,000 | $0.26 |
| parity check (non-panel, 24 pairs) | 1 | 6,000 | $0.04 |
re-dispatch allowance (F2) |
≤ 3 | 14,000 | $0.20 |
| total | $0.88 |
Worst cases are built from max_tokens at the most expensive plausible provider, per note (abc).
Snapshots either side; per-response usage.cost is primary and the key delta is the cross-check.
Amendments
Pre-run critic: nvidia/nemotron-3-ultra-550b-a55b (non-panel), stop, 22,122 characters,
$0.0297342, provider Together. Verdict NEEDS-REDESIGN, 8 BLOCKING. Six accepted in whole or in
part, two remedies overruled with reasons. No scoring call had been dispatched.
A1(from BLOCKING (d) — the best finding in the pass, accepted whole). The design measuredperceived-source-carriageonly source-absent, so itsFORMeffect could not be told apart from seeing the source suppresses markedness attribution. A third stage is added: Stage C, same 30 items, source present,perceived-source-carriagealone, dispatched last. The sense is now measured under both information conditions on the same items with the same seats, which is what the question needs and whatRS-20260802fdid at site level.A2(fromA1). Two predictions are added,P5andP6below.A3(from BLOCKING (c), accepted in part). The critic went through the eighteen strangeness edits and found all six odd prepositional government edits traceable to the French preposition at that site —saturated of←saturé de,upon them←sur eux,unto←àtwice. That is correct and it made a third of theSTRANGEfactor source-licensed, which is the one thing the factor may not be. All six are replaced by broken / unresolved constructions — a kind the passage contains nowhere, since its syntax completes every construction it starts, and the kindRS-20260802fmeasured as the worst case for false attribution (5 of 6 codings read a manufactured anacoluthon as faithful). The critic's verdicts on the compounds and the inversions are overruled: French has neither Germanic noun-compounding nor English predicate fronting, soskin-gatherers,up-buildings,between-thing ties,tomb-wailer,over-/under-goodnessand all six inversions are markedness English supplies out of its own resources. Its arguments there invent French forms not in the text (a cleft at S3 that Schwob does not write, asur-/sous-prefixation the passage does not use).A4(from the same finding, accepted). Edit 2 moves offdistortion-mirrors, which shadowed the ordinary Frenchmiroirs déformants, toshape-warping mirrors.A5(from BLOCKING (e), accepted). An accuracy-parity gate over arms built to be propositionally equivalent cannot fail, and a gate that cannot fail is not a gate. A fifth armWRONGis added —CCwith one content error per segment and no formal change — andG1is re-registered with a power condition.WRONGis not a cell of the 2 × 2 and enters noFORMorSTRANGEfigure. Items go from 24 to 30.A6(from BLOCKING (e) onG3's circularity, accepted).G3usedstyle-correspondence— one of the two target senses — as the sole manipulation check forFORM.G3bis added: a mechanical device count inverify.py, which counts anaphora openings, repeated keyword chains, paragraph-initial connectives, jussiveLetopenings, template members and figurae etymologicae in each arm and asserts the flattened arms carry strictly fewer. It uses no seat and no judgment.A7(from BLOCKING (b), accepted as reporting). Every effect is computed within segment (each segment supplies all five arms, so the design is paired) and the analysis reports the item-level correlation of word count with each sense, so a length reading of anyFORMeffect stays visible. The confound is not repaired and §7.1 stands.A8(from BLOCKING (g), accepted in part). Order is re-randomised per stage as well as per seat, so nothing correlates across stages. The remedy of adding 24 filler passages is overruled: it doubles the spend against a leak that lets a seat see there are variants but not which factor is which, and the factors are not nameable from the items alone. Declared as a limit, asRS-20260807edeclared the same one.A9(from BLOCKING (f), limit accepted, remedy overruled). An independent translator must produce at least one arm per cell would destroy the design: the arms differ only by a matched operator, and an independent hand would reintroduce every difference the matching removes.G5puts the device catalogue outside the lead,G3bputs the manipulation check outside judgment, and the parity call puts propositional equivalence outside the lead. Nothing puts the edits outside the lead and §7.3 says so.A10(from BLOCKING (h), accepted). Binding on the result page: its headline may say what this manipulation did on this material, and may not say that two senses are or are not one thing. One passage, one author, one hand, one pair.A11(budget). Stage C and the sixth item per stage raise the declared ceiling from $0.90 to $1.20. Headroom on the UTC day at the time of the amendment is $2.216239055; the revised worst case is $1.17.
Re-registered gates and predictions (superseding §5 and §6 where they conflict)
G1—accuracy(Stage A):max − minover the four 2 × 2 arms ≤ 1.00, andmean(four arms) − mean(WRONG) ≥ 1.00. The second clause is the power condition: if the gate cannot separate six planted content errors from none, its first clause is uninformative and the primary is withheld.G3b— mechanical device count,verify.py: for each of the six segments, the flattened arms must carry strictly fewer counted device tokens than the carried arms. A segment that fails is named and excluded from everyFORMfigure.P5—perceived-source-carriage'sFORMeffect is larger source-present (Stage C) than source-absent (Stage B), by ≥ 0.50. The seat that can check the French should be able to tell carried form from invented oddity; the seat that cannot, should not.P6—perceived-source-carriage'sSTRANGEeffect is smaller source-present than source-absent, by ≥ 0.50. This isRS-20260802f's 0.60-without / 0.40-with, in the sense's own scoring form. Fails if seeing the French does not reduce the credit that unlicensed oddity earns.