Repository path: workshop/experiments/E-20260811c-source-beliefs/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260811c-source-beliefs |
| status | frozen |
| created | 2026-08-11 |
| updated | 2026-08-11 |
| senses | perceived-source-carriage, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-source-beliefs.md, wiki/goodness-senses.md, workshop/translations/sensitiva-amorosa/R06-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R14-matched-flattening.md, wiki/method-notes.md |
E-20260811c — does a translation that carries the source's form tell its reader the truth about the source?
Frozen 2026-08-11 before dispatch, and AMENDED 2026-08-11 on the pre-run critic pass
(critic.md: NEEDS-REDESIGN, 20 findings, 17 BLOCKING; 14 accepted, 3 in part, 3 overruled).
Every amendment was applied before any judge, parity, audit or recognition call was dispatched.
Track T2, arm ARM-source-beliefs step 1.
Copy-text and stimuli, fixed by checksum (sha256, first 16): source-sv.txt 4f40602f727bc821
· arm-CARRY.txt 919aecd274fa3c1f · arm-FLAT.txt f2dcdbb3fd10ac7b · arm-ODD.txt
7d79d606c69855f7 (the last recomputed after amendment A-F06).
0. The subject-rule sentence (continue-prompt.md §4.5)
What this unit teaches about translating literature or evaluating translations: whether a translation that carries its source's marked form gives a reader true beliefs about a text they cannot read — the epistemic claim foreignizing translation makes for itself — and whether English that is merely marked, carrying nothing, gives them false ones with the same confidence. The question is about translations and their readers, not about this project's instruments.
1. The question, and why it is the one owed
wiki/goodness-senses.md §perceived-source-carriage now carries four measurements and they all
say the same thing from different angles: a score on that sense is a fact about the reader and
there is no reader-side question yet found that turns it into a fact about the text. Manufactured
oddity is attributed to the source at 0.40–0.60 (RS-20260802f); unlicensed oddity outscores real
carried form, +1.861 against +1.750 (RS-20260807f) and +2.286 against +1.810 (RS-20260808d); a
strict word-multiset scramble outscores five of six independent hands' foreignizing arms
(RS-20260809h); and asking the reader to point at the source locus rather than rate does not
repair it — a scramble drew attribution in 0.60 of cells beside a 0.90 hit rate on real licences
(RS-20260810d). Note (blr) follows: a design that wants to know whether a translation carries
a source feature must measure it on the two texts, not ask a reader.
Every one of those measures the sense's validity. None measures its consequence. S157
gave cultural-mediation a consequence measurement — what a culture-bound decision does to a
reader, rather than whether raters box it right — and this run is the same move for this sense.
The consequence in question is the one Venuti's argument turns on. A foreignizing translation is defended on the ground that it gives its reader access to a text they cannot read. That is a claim about beliefs, and beliefs about a text have a truth value that can be settled in the text. So:
Does a reader of a form-carrying translation end up believing more true things about the Swedish than a reader of a flattened one — and does a reader of a merely-marked translation end up believing false ones?
What the run can establish, narrowed on critic findings F12 and F17 before dispatch. The critic is right that both halves of that question are answered by one mechanism, and that the mechanism is not subtle: the reader projects the form of the English onto the original. Nothing here shows a reader performing calibrated inference about Swedish. So the claim the run is registered to support is exactly this and no more:
The reader projects the translation's form onto the original. Where that form came from the source, the projection yields TRUE beliefs about a text the reader cannot read. Where it came from the translator, the same projection yields FALSE ones. The difference between the two is a fact about the translator's hand, not about the reader's discrimination.
That is still the consequence the sense has never had, and it is still Venuti's claim under test — because the defence of foreignizing translation is not that readers discriminate, it is that the form they are given is the source's. It is weaker than the framing this design opened with, and the weaker version is the one every sentence of the result must keep to.
This does not violate note (blr). The reader is never asked whether something was carried. The reader is asked what the original is like; truth is settled by counting the Swedish, by script.
2. Materials
Source. Ola Hansson, Sensitiva amorosa (1887), section IX opening — 925 Swedish words in
six segments (125–190 words each), material/source-sv.txt. Litteraturbanken's proofread etext of
the first edition. Public domain. The project's third Swedish work and its first by a writer
other than Lagerlöf. The passage is a November park, a woman seen once, and a companion beginning to
tell a summer memory; it has no proper name in it except the town reduced to "H." and a quayside
nickname.
Why this source. Its marked devices are exactly the ones English prose style suppresses by default and are all countable: main clauses strung on commas, single sentences of 66–102 words, polysyndeton, immediate word-doubling, and zero connectives of cause or concession in 925 words.
The three arms. All written by the lead. Propositional content held fixed throughout.
| arm | what it is |
|---|---|
CARRY |
T-sensitiva-amorosa-R06-v1, frozen at commit e6b90bd with its translator's log, before this design existed. R06: source-only, single pass, no return pass |
FLAT |
R14 restricted to F1 (cadence levelling), F3 (repetition flattening) and F6 (connective explicitation), applied to CARRY. The restriction is declared and deliberate: FLAT alters only what bears on the class-A statements and leaves every other property byte-identical, so that class B stays clean |
ODD |
FLAT plus unlicensed markedness — markedness of kinds the Swedish does not license anywhere in the passage. Four operators, applied in every segment: O1 one declarative made exclamatory; O2 two words set in italics; O3 one aside set off by a dash; O4 three consecutive sentences begun with And |
No word-order inversion is used anywhere in ODD. Swedish is a V2 language, so a fronted
inversion in the English would be a calque of the source, not an unlicensed edit — the defect
RS-20260807f's critic caught in six of eighteen edits and RS-20260808d avoided by the same
reasoning.
3. The instrument
Thirteen statements about the Swedish original, put to a seat who sees one English arm and one
segment. Every ground truth is decided by census.py over the Swedish and by nothing else: the
lead annotates nothing. (RS-20260811b's realia spans were the lead's and an independent annotator
recovered them at 0.517; that failure mode is designed out here rather than measured again.)
Class A — device-linked (five). Properties of the Swedish for which CARRY preserves the
evidence in English and FLAT removes or reverses it.
| id | statement | predicate over the Swedish |
|---|---|---|
| A1 | "At two or more places, the original has a comma followed directly by a pronoun subject and a verb, with no conjunction in between (a comma splice)." | ≥ 2 matches of a closed pronoun × finite-verb pattern |
| A2 | "…at least one sentence of more than sixty words." | max sentence > 60 tokens |
| A3 | "…two or more explicit connectives of cause, concession or result — words corresponding to because, since, although, therefore, so that." | ≥ 2 hits in a closed 15-word Swedish list |
| A4 | "…a single sentence containing the word and three or more times." | ≥ 1 sentence with ≥ 3 och |
| A5 | "…repeats a word immediately, in the form X and X, or repeats the same three-word sequence twice inside one sentence." | ≥ 1 |
Every statement is worded as the literal predicate that decides it (critic amendment A-F09). The design makes no claim that these predicates are a correct linguistic analysis of Swedish — only that they are stipulated, public, deterministic and reproducible from the checksummed copy-text. Every claim the result makes is qualified as true of the Swedish under the published predicate.
Class B — false lures (four). Properties that are false of the Swedish and that ODD's four
operators install in the English. CARRY and FLAT should both answer them correctly.
| id | statement | predicate | operator that lures |
|---|---|---|---|
| B1 | "…at least one exclamation mark." | ≥ 1 ! |
O1 |
| B2 | "…sets at least one word in italics." | ≥ 1 | O2 |
| B3 | "…contains at least one dash." | ≥ 1 dash | O3 |
| B4 | "…three or more consecutive sentences begin with the same word." | longest run ≥ 3 | O4 |
Three class-B cells are source-POSITIVE and are excluded from P2's estimand (critic amendment
A-F04, registered before dispatch): B3 in S4 and S5 (the Swedish has dashes there) and B4 in
S6 (the Swedish does have a three-sentence run). In those cells ODD's edit reinforces a true
answer rather than luring a false one, and pooling them would let the lure class be scored on cells
that are not lures. They are reported separately.
Class C — content control (four), each also worded as its predicate: past-tense verbs outnumber present-tense verbs · contains at least one quotation mark · mentions the sea, a sound, a harbour, a quay, or waves · uses the second-person singular pronoun. All three arms carry identical evidence for these. This is the reading gate, and it is a floor, not a certification that the seats can infer anything subtler (critic amendment A-F11).
Ground truth, from the Swedish (T = true; census.py --truth):
seg A1 A2 A3 A4 A5 B1 B2 B3 B4 C1 C2 C3 C4
S1 T T F T T F F F F T F F F
S2 F F F T F F F F F T F T F
S3 T T F T F F F F F T T F F
S4 F T F T T F F T F F F F T
S5 F T F T F F F T F T T T T
S6 T T F T T F F F T T F F F
31 of 78 cells true; class A 17/30, class B 3/24, class C 11/24.
Registered exclusions from the primary estimand (critic amendment A-F16), fixed before
dispatch: (A4, S2) — FLAT retains polysyndeton there, so the arms do not differ;
(A5, S5) — CARRY and FLAT both contain "set stone and stone rubble", an X and X the
Swedish does not have, so both arms are equally misled; (B4, S6) — CARRY's longest same-word
run is 2 where the Swedish is 3, so the frozen translation lost the source's anaphora and no arm
supports the true answer. All three are reported alongside the primary; the third is not repaired,
because it is a finding about the translation rather than a defect in the stimulus.
3.5 The prompts, verbatim (critic amendment A-F13)
The statement prompt, used for both the AB block and the C block, temperature: 0:
You will be shown an English translation of a passage from a Swedish prose work of the 1880s.
You cannot see the Swedish original.
Below the passage are {n} statements ABOUT THE SWEDISH ORIGINAL. For each one, decide whether it
is TRUE or FALSE of the original, judging from the English translation in front of you. The
translation is your only evidence about how the original is written.
Answer about THIS passage, not from general knowledge of Swedish or of Swedish literature. If you
think you recognise the work, ignore what you know about it.
Reply with exactly one line per statement and nothing else, in this format:
<id>: TRUE|FALSE, <confidence>
where <confidence> is an integer 0-6 saying how sure you are (0 = a pure guess, 6 = certain).
STATEMENTS:
{statements}
THE TRANSLATION:
{text}
There is no INSUFFICIENT-EVIDENCE option, and the critic asked for one. It is declined: it
would make accuracy non-comparable across arms and duplicate the 0–6 confidence scale that already
carries the uncertainty. Declared as a limit. The parity, explicitation-audit and recognition
prompts are in run.py and are reproduced in the result page.
Statement order is counterbalanced (critic amendment A-F15): the fixed interleaved order
B3 A4 B1 A2 B4 A1 B2 A5 A3 on odd-numbered segments and its exact reverse on even-numbered ones,
identical for every arm, so order is crossed with segment and confounded with no arm.
4. The manipulation check (G4), computed before dispatch
census.py over the Swedish and over each English arm. Reported as part of the design, not as a
result, because a manipulation that did not take makes everything downstream unreadable.
| feature | Swedish S1–S6 | CARRY |
FLAT |
ODD |
|---|---|---|---|---|
| comma-splices | 2 0 3 1 1 4 | 2 0 2 2 1 4 | 0 1 0 1 1 0 | 0 0 0 1 1 0 |
| longest sentence | 80 40 67 102 66 71 | 90 46 82 109 94 73 | 28 38 29 31 40 23 | 28 38 29 31 40 23 |
| explicit connectives | 0 0 0 0 0 0 | 0 0 0 0 0 0 | 6 2 2 3 3 2 | 6 2 2 3 3 2 |
| polysyndeton sentences | 2 1 2 2 3 2 | 2 1 2 2 3 2 | 0 2 0 0 0 0 | 0 2 0 0 0 0 |
| repetition figures | 1 0 0 1 0 3 | 1 0 0 1 1 2 | 0 0 0 0 1 0 | 0 1 0 0 1 0 |
| exclamation marks | 0 ×6 | 0 ×6 | 0 ×6 | 1 ×6 |
| italics | 0 ×6 | 0 ×6 | 0 ×6 | 2 ×6 |
| dashes | 0 0 0 4 3 0 | 0 0 0 4 3 0 | 0 0 0 2 3 0 | 1 2 1 2 3 1 |
| longest same-word run | 1 1 2 1 1 3 | 1 1 2 1 1 2 | 1 2 2 1 2 2 | 3 ×6 |
CARRY reproduces the Swedish profile almost exactly — polysyndeton identically in all six
segments, comma-splices within one in five of six, every long sentence long. FLAT reverses every
class-A feature. ODD installs all four class-B lures in every segment and changes nothing in
class A.
Three declared imperfections are visible in that table; all three are handled by the registered
exclusions in §3 rather than tidied away. After critic amendment A-F06 the class-A feature vector
of ODD is identical to FLAT's in all six segments — before the amendment ODD S2 had lost a
comma-splice and gained a repetition figure, which would have made P4 untestable. verify.py
asserts the identity.
5. Procedure
- Recognition gate (
G3), onopenai/gpt-5.6-terra, which judges nothing in this run (critic amendment A-F01): the whole of each arm, three calls, "name the work or the author". - Parity control (
G2),moonshotai/kimi-k3, which also judges nothing else: 23 pairs —CARRY×FLAT,CARRY×ODDandFLAT×ODDon all six segments (18 real pairs), plus five pairs in which version B carries one planted content error. Instruction: ignore style, punctuation and emphasis; attend only to what is asserted. - Explicitation audit,
openai/gpt-5.6-terra, 12 pairs (CARRYagainst each of the other two, six segments): the opposite question — what relation does B state that A leaves implicit? — with a count and a list. This exists becauseG2's "ignore style" instruction suppresses exactly the defect F6 explicitation introduces (critic amendment A-F08). - Judge stage: three blind seats × 36 items = 108 cells. One arm, one segment, one call. Class A and class B are asked together (they are the two sides of one comparison and must face identical conditions); class C is a separate call.
Seats. J1 x-ai/grok-4.5, J2 google/gemini-3.6-flash (reasoning: low), J3
deepseek/deepseek-v4-pro; caps 1200 / 1200 / 3000, generous from the first dispatch — note (bmb)
and S157's 25 dead deepseek bodies. Critic, recognition and audit openai/gpt-5.6-terra; parity
moonshotai/kimi-k3. No seat both judges and gates. The lead judges nothing.
6. Predictions, registered
The unit of inference is the SEGMENT, not the segment × seat cell (critic amendment A-F02). Three seats answering the same six segments are not eighteen independent draws. The primary test is the exact one-sided sign test over the six segments, seats pooled within segment; the seat-by-segment pattern is reported as a consistency description and carries no inferential weight. Tie rule, fixed in advance: ties are discarded and the exact test is conditional on the number of non-tied segments m, so the attainable one-sided P values are 1/2^m for a clean sweep — 0.0156 at m = 6, 0.0313 at m = 5, 0.0625 at m = 4, i.e. a run in which fewer than five segments break the tie cannot reach P < 0.05 and the primary is then reported as not estimable.
P1 (primary). Class-A accuracy is higher for CARRY than for FLAT. Per segment, the
difference in proportion correct over the class-A statements in the primary estimand (five
statements, four in S2). Registered bar: mean difference ≥ +0.15 and exact one-sided sign test
P < 0.05 over the non-tied segments.
P2. Class-B accuracy is lower for ODD than for FLAT, over source-negative class-B
cells only (§3), same unit, same test, registered bar mean difference ≤ −0.15 and P < 0.05.
P3 (exploratory, not a test). Mean confidence on source-negative class-B items by arm,
regardless of correctness, with mean class-A confidence by arm beside it, resampled by segment.
The claim it speaks to: markedness buys confidence without buying truth. Registered as descriptive
(critic amendment A-F18).
P4. Class-A accuracy for ODD is not higher than for FLAT. ODD is now certified to
carry an identical class-A feature vector to FLAT, so a difference of ≥ +0.15 in either direction
is a defect in the instrument and must be reported as one.
What a positive P1 and a positive P2 together license, and nothing more: these three seats,
on these six fixed segments, projected the English's form onto the Swedish — truly where the form
came from the source, falsely where it came from the translator.
7. Gates and failure criteria, registered
G1reading gate. Class-C accuracy ≥ 0.80 in each arm, reported per seat as well; any seat below 0.70 is named in the result.F1: ifG1fails in any arm, the primary is withheld.G2parity. ≥ 13 of 18 real pairs EQUIVALENT and ≥ 4 of 5 planted pairs caught by name.F2: if the planted arm is not caught at ≥ 4 of 5,G2is uninformative and is reported as such, not as a pass.F3: if fewer than 13 of 18 real pairs are equivalent, the run is about damage and not about form, says so, and the primary is read on the certified subset. Either way the primary is reported both on all six segments and on the certified subset.G2bexplicitation. The audit's count is reported, not gated — a zero bar on an LLM adjudicator's free text is a coin-flip, not a bar (the critic's "require zero" is overruled). Segments where the audit finds a RELATION-ADDED form the secondary subset forP1.G3recognition. If ≥ 2 of 3 arms are named by work or author, the primary is withheld. This gates explicit recognition only and claims nothing about latent familiarity.G4manipulation. Registered condition:FLATreverses ≥ 3 of 5 class-A features in ≥ 4 of 6 segments;ODDinstalls all four class-B lures in ≥ 5 of 6; andODD's class-A vector equalsFLAT's in all six. All three hold (§4).F4. If more than 5% of judge cells fail to return a parseable answer, both the as-registered dataset and the re-dispatched one are reported in full, with every gate verdict given for both.F5. Any statement whose wording does not match its predicate is dropped before dispatch, not reinterpreted afterwards. Three were reworded on the critic's findings and none was dropped.
8. What this design cannot show, written before the run
- The registered conclusion limit, in the critic's words and adopted verbatim: "In these three prompted LLM systems, for these fixed Swedish–English stimuli, visible English form affected answers to explicit questions about source form." Any human-reader or theory-level claim needs a separate human, multi-text, multi-translator study.
- The estimand is an instructed task, not reading (
A18), and Tier D is NOT PASSED: no jury verdict in this project carries evidential weight. Everything here isinternal-judgment-onlyandprovisional. - Six segments of one passage by one author in one hand. Not a sample of literary texts. The estimand is these six fixed segments.
- A prior about Swedish prose can INTERACT with the arms, not merely dilute them — the design's
original claim that it "can only shrink" the difference is struck on critic finding F14. A
seat may read a long comma-linked English sentence as likely literal preservation and a
polished short-sentence English as the translator's modernisation, which would produce
P1with nothing at all transmitted about this passage. This is a live alternative explanation ofP1and the result must carry it. P2is surface-cue transfer and is registered as such (critic finding F12). It shows that a punctuation or emphasis feature of the translation is read back onto the original; it does not show that markedness in general induces false beliefs. The crossed, held-out operator design that would generalise it is named as this arm's successor.FLATstates relations the Swedish leaves implicit — that is what F6 explicitation is, andRS-20260808d's audit found the same at 4 of 7 segments. The explicitation audit counts it and the primary is read on the certified subset as well.- The lead wrote all three arms. The ground truth predicates, the parity certification, the explicitation audit and the recognition gate are outside the lead; the edits are not. The critic's remedy — multiple independently-authored realisations — is a larger experiment and is declined here rather than faked.
9. Budget
Declared ceiling $1.80, built from max_tokens and the worst plausible provider (note (abc);
config/models.md's routing caution), raised from $1.60 before dispatch because the critic's
amendments added 6 parity pairs and a 12-call explicitation audit. Pre-flight: critic 1 call at cap
14,000, actual $0.094221; recognition 3 calls at cap 900 (~$0.02); parity 23 calls at cap 2,000
(~$0.18); audit 12 calls at cap 2,500 (~$0.14); judge 108 calls at caps 1,200/1,200/3,000 (~$0.55).
Expected ~$0.99. UTC day 2026-08-11 stood at $1.285740037 of $5.00 before this session, so the
ceiling fits with $1.91 to spare.