Repository path: workshop/experiments/E-20260808d-carriage-decoupled/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260808d-carriage-decoupled |
| status | frozen |
| created | 2026-08-08 |
| updated | 2026-08-08 |
| track | T2 |
| senses | perceived-source-carriage, naturalness, style-correspondence, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-sense-overlap.md, wiki/findings/results/RS-20260807f-carriage-or-strangeness.md, wiki/findings/results/RS-20260802f-licensed-strangeness.md, wiki/goodness-senses.md, workshop/translations/ved-vejen/device-census.md, workshop/translations/ved-vejen/R06-v1/translation.md, workshop/translations/ved-vejen/R14-v1/translation.md, config/models.md |
E-20260808d — does perceived-source-carriage survive when form and markedness come apart?
ARM-sense-overlap step 2. Frozen before dispatch. Nothing below was written after a number
existed.
1. The question, and why it is not a writing step
Step 2 was scoped as write the verdict into wiki/goodness-senses.md. The verdict it has to carry
is RS-20260807f §3: blind perceived-source-carriage correlates with naturalness at −0.975,
and its responses to carried source form (+1.750) and to gratuitous oddity (+1.861) are
indistinguishable — i.e. the sense D-20260802-13 created in order to be separate from
naturalness is that material's naturalness with the sign flipped.
That page's own limit 2 forbids writing it as a standing relation. Carrying Schwob's litany into
English marks the English, so on that passage FORM and markedness are confounded by
construction, and the observed identity is exactly what a confound of that shape produces. §6.4 of
that page specifies the replication owed: a source whose marked forms can be carried into
unmarked English, plus an accuracy gate scored without the other senses in the same call.
This is that replication. A standing sentence in the project's controlled typology, cited by every
future evaluation, may not rest on one passage whose design defect is written on its own result page.
What this unit teaches about translating literature (subject rule, wiki/tracks.md): it decides
whether a reader who reports that a translation carries its source across is reporting anything
beyond this English is odd — which is the empirical content of the foreignizing claim, and is a
question about translations and their readers, not about this project's instruments.
2. Materials
- Source. Herman Bang, Ved Vejen (1886), ¶829–852 — 503 Danish words in 24 paragraphs, the
dawn movement.
workshop/translations/ved-vejen/source-unit.txt. The project's first Danish. - Why this passage decouples.
workshop/translations/ved-vejen/device-census.md§3, in full: the project's unmarked / literary-contemporary register anchorA-mchugh-presenceis catalogued as invisible free indirect discourse ·say-only dialogue tags · texture carried by named particulars · structural one-sentence paragraphs, and Bang 1886 has all four (devices A4, A5, B7, A1/A3). Carrying these devices should therefore cost nothing onnaturalness. That is not an assumption of the analysis — it is gateG1, and the experiment is void if it fails. - Census. Sixteen Class A form-only devices (no proposition, no image) and seven Class B
image-bearing features, made from the Danish alone and committed at
789ae0dbefore any English of this passage existed. - Segments. Seven,
S1–S7, 46–95 Danish words, mapped to source paragraphs inT-ved-vejen-R14-v1§4.
The four arms
| arm | what it is | words |
|---|---|---|
CARRY |
T-ved-vejen-R06-v1 — the lead's close rendering, all sixteen Class A devices carried. Frozen at 5999d6f |
563 |
FLAT |
T-ved-vejen-R14-v1 — CARRY with all sixteen Class A devices removed by the 24-site operator table. Every Class B string identical. Length ratio 1.041 |
586 |
STRANGE |
FLAT plus 19 unlicensed markedness edits that answer to nothing in the Danish. Positive control for the sense |
609 |
WRONG |
FLAT plus five real content errors (three polarity reversals, one state reversal, one quantity change), no formal change. Positive control for the accuracy scale |
585 |
CARRY vs FLAT is the primary contrast, and it is the contrast RS-20260807f could not make:
form manipulated, markedness held (that is G1), content held (that is G3), imagery held byte for
byte (that is the Class A / Class B split).
The STRANGE edits use no word-order inversion. Danish is a V2 language, so an English
fronted-adverbial inversion would be a calque of the source's ordinary syntax — precisely the
defect RS-20260807f's pre-run critic caught in six of its eighteen unlicensed edits (its BLOCKING
(c)). Every edit carries, in materials/arms.py, a written statement of what the Danish does at that
site; all nineteen read PLAIN.
3. Panel and stages
Seats J1 = P1, J2 = P2, J3 = P5 (config/models.md). Critic and the two controls are
non-panel. The lead judges nothing (charter §5).
| stage | what | source shown | senses in the call |
|---|---|---|---|
| 0 | pre-run critic over this frozen design | — | — |
| 1 | blind naturalness, dispatched first, before any seat sees Danish (S134 amendment A10) |
no | naturalness alone |
| 2 | Danish competence screen, 5 items | yes | — |
| 3 | accuracy alone — §6.4's requirement, so a within-call halo from the other senses is impossible |
yes | accuracy alone |
| 4 | style-correspondence + perceived-source-carriage |
yes | those two |
| C1 | independent propositional-parity call (non-panel): are the paired texts saying the same things, and where do they differ? | yes | — |
| C2 | independent device census from the Danish alone (non-panel) | yes | — |
28 items (4 arms × 7 segments) per seat per rating stage; 3 seats; 252 sense-cells plus 15 competence cells. Item ids are opaque hashes, order is seed-fixed and independently shuffled per seat, and arms are interleaved.
4. Gates — all declared here, none movable after a number exists
| id | statistic | bar | what it protects |
|---|---|---|---|
G1 |
|Δnaturalness(CARRY−FLAT)|, blind, seat-pooled segment means |
≤ 0.75 | the design. If the carried devices cost naturalness, form and markedness are confounded here as they were on Schwob, and the primary is withheld |
G2 |
Δstyle-correspondence(CARRY−FLAT) |
≥ 1.00 | the manipulation took — a form-sense sees the devices go |
G3a |
|Δaccuracy(CARRY−FLAT)|, scored alone |
≤ 1.00 | content parity |
G3b |
independent parity call: CARRY/FLAT pairs judged propositionally equivalent |
≥ 6 of 7 | content parity, outside the lead |
G3c |
same call: planted WRONG errors named |
≥ 4 of 5 | that G3b is not a rubber stamp |
G4 |
Δperceived-source-carriage(STRANGE−FLAT) |
≥ 1.00 | the sense is live here. A null on the primary is uninterpretable without it |
G5 |
Δaccuracy(FLAT−WRONG) |
≥ 1.00 | the accuracy scale is live |
G6 |
Danish gloss, per seat | ≥ 4 of 5 | the source-present stages mean anything |
G7 |
independent census names Class A devices | ≥ 0.60 of 16 | the census is not the lead's private list. Reported, non-gating |
5. Predictions, registered
The primary is read on seat-pooled segment means, n = 7, with the 21-cell form reported
alongside. Exact paired permutation over sign flips, 2⁷ = 128 relabelings, minimum attainable
P = 0.0078.
P1(superiority). Δperceived-source-carriage(CARRY−FLAT) ≥ 1.00 with exact P ≤ 0.05. Holds ⇒ the sense has a source-specific component that survives decoupling.E1(equivalence, absolute). The 90% permutation interval on that difference lies inside ±0.75. Holds ⇒ the sense is measurably indifferent to carried source form.E2(equivalence, relative). |Δ(CARRY−FLAT)| ≤ 0.33 × Δ(STRANGE−FLAT) on the same sense, same seats, same items.E2exists becauseE1alone has already failed once on seven segments —RS-20260808creturned a point estimate of exactly 0.000 and an interval of [−0.777, +0.666], missing ±0.75 by 0.027, note (bkm).E2is powered by the run's own manipulation instead of by an outside guess at what margin matters.P2. r(blindnaturalness,perceived-source-carriage) over the 28 items ≤ −0.80.RS-20260807fgot −0.975 on confounded material; this asks whether the identity survives decoupling.P3. Δperceived-source-carriage(STRANGE−FLAT) ≥ Δ(CARRY−FLAT) — unlicensed oddity earns at least as much source-carriage credit as really carried form.P4. Δaccuracy(CARRY−FLAT), scored alone, is smaller thanRS-20260807f's +0.917 measured with four senses in one call. Holds ⇒ that page'sG1failure was a within-call halo. Fails ⇒ form contaminatesaccuracyas a property of the sense, and everyaccuracyfigure in this repo carries it.
6. Failure criteria — declared, and not weakened after firing
F1.G1fails ⇒P1,E1,E2WITHHELD. The passage did not decouple and this run is a second measurement of the confound, not a test of the sense.F2.G4fails ⇒P1,E1,E2WITHHELD. A sense that does not move on this material cannot be shown not to move on form.F3.G3b< 6 of 7 orG3c< 4 of 5 ⇒ the arms are not content-matched; everything is descriptive andP4is void.F4.G6fails for ≥ 2 seats ⇒ stages 3 and 4 are uninterpretable.F5. Any seat returning < 80% of its cells is dropped and the loss is declared in the result; cell completeness is reported whatever it is.F6. IfE1andP1both fail andE2holds, the result is reported as equivalent by the run's own yardstick and not certified at ±0.75, in those words, and the typology text says so.
7. Known limits, written before the run
- The lead wrote all four arms and chose the passage, as at
RS-20260807f.G3b,G5andG7put parity, the error scale and the census outside the lead; nothing puts the 24 flattening edits or the 19 oddity edits outside the lead. - The lead knew the hypothesis while writing the operator tables — though not while translating
CARRY, whose log was frozen at5999d6fbefore this design existed. - Contamination is a reachability statement, not a measurement (
device-census.md§4): no English «Ved Vejen» is reachable. Both primary arms descend from one rendering, so any recall is shared and cannot produce the contrast. - One phrase escapes the manipulation by design:
half so of gold and half of rosesis a Class A doubling sitting inside Class B image B1, so it is held constant and survives intoFLAT. This works against the run's hypothesis, and is declared inT-ved-vejen-R14-v1§2. F6/F7add words. Connective explicitation and FID attribution supply subordinators and attributions that the Danish leaves implicit.G3bis the control that adjudicates them, andRS-20260807f§5 limit 4 is why this is stated in advance rather than discovered.- One passage, one author, one hand, one pair, three seats. Tier D is NOT PASSED and
perceived-source-carriagehas never been through Tier D at all. Every sentence of the result isprovisionalandinternal-judgment-only.
8. Pre-flight cost estimate
Built from max_tokens, not from expected output (note (abc)), and priced at the worst plausible
provider (config/models.md, S022 caution).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| 0 critic | 1 | 16,000 | $0.06 |
| 1 blind | 3 | 8,000 | $0.20 |
| 2 competence | 3 | 2,000 | $0.05 |
| 3 accuracy | 3 | 8,000 | $0.25 |
| 4 two senses | 3 | 10,000 | $0.35 |
| C1 parity | 1 | 8,000 | $0.05 |
| C2 census | 1 | 6,000 | $0.04 |
| declared ceiling | 15 | $1.00 |
Today's ledger (UTC 2026-08-08) stands at $1.932724933 of $5.00 across S132–S134; $3.067 headroom. $1.00 fits.
9. Amendments adopted from the pre-run critic, 2026-08-08, before any rating was dispatched
nvidia/nemotron-3-ultra-550b-a55b (non-panel), max_tokens 16,000, stop, 11,484 characters,
$0.0349038. Verdict NEEDS-REDESIGN, 5 BLOCKING, 8 ADVISORY. Eight findings accepted, four
overruled with written reasons. The critic's own summary of the design's largest weakness — "no
gate catches the confound in Finding 1" — is what amendments A1 and A2 answer.
Accepted
A1(BLOCKING 1, in part). Three flattening sites supplied a causal subordinator the Danish does not state (because×2,since×1). All three are replaced by bareand, inT-ved-vejen-R14-v1§5. Two other explicitations were checked against the Danish and kept because the Danish asserts the relation itself.A2(BLOCKING 1 + BLOCKING 5). The parity control now asks two questions, not one: propositional equivalence (gating,G3b), and separately enrichment — does either passage state outright a relation, cause, attribution or inference the other leaves for the reader to supply? The second is reported, not gating, because a positive answer is expected: the whole manipulation is about what is enacted versus what is stated, and a gate that fires by construction is not a gate. Its rate and direction are what the result must carry.A3(BLOCKING 4).G1is tightened from ≤ 0.75 to ≤ 0.50, and the design's claim is restated: carrying these devices costs less than half a scale point ofnaturalness, not "nothing". Justification, which the critic correctly said was missing: 0.50 is under a third of the −1.444 thatRS-20260807f'sFORMmanipulation cost on confounded material, and above the 0.333 thatRS-20260805eshowed is not distinguishable from unit-level noise on this instrument.A4(BLOCKING 2).P2is void as registered — 28 items are 7 segments under 4 conditions, so the correlation would report between-arm separation, not a within-segment relation.P2is replaced byP2′: the correlation between blindnaturalnessandperceived-source-carriagecomputed on segment-centred residuals (each item's score minus its segment's mean over the four arms), df = 20. The uncentred figure is reported alongside, labelled as the invalid one.A5(BLOCKING 3).P4's cross-experiment comparison is dropped. The halo question is now tested inside this run, on the critic's own proposed design: stage 5 has the same three seats score all four senses in one call on the same 28 items, andP4′compares Δaccuracy(CARRY−FLAT) scored alone against the same difference scored alongside three other senses, pairwise, same seats, same items.RS-20260807f's +0.917 becomes a descriptive reference and not a test.A6(ADVISORY 1).G3bis raised from ≥ 6 of 7 to 7 of 7.A7(ADVISORY 3). Power is stated rather than implied: withn= 7,P≤ 0.05 requires all seven segment differences to share a sign, andE1has already failed once on seven segments at a point estimate of exactly zero (note (bkm)).E2exists for that reason.A8(ADVISORY 4). Segments are 46–95 Danish words. Both the unweighted and the word-weighted pooling are reported, and a disagreement between them is a finding, not a choice.
Overruled, with reasons
- BLOCKING 1's main remedy — drop A4, A10, A12, A13, A14, A15, A16 from the manipulated set.
Overruled. The critic's ground is that these change narrative mode, discourse structure and
perspective. They do; narrative mode is form, and it is exactly the form this experiment
exists to manipulate. Adopting the remedy would leave a manipulation of punctuation, italics and
dialogue tags, which is not what
wiki/goodness-senses.mdmeans bystyle-correspondenceand is not what any translator of Bang decides. The enrichment charge inside the finding is real and is answered byA1andA2; the reclassification is not. - BLOCKING 5's remedy — move
style-correspondenceto a separate or blind call. Overruled on the blind half: this sense is defined as a source–target relation and scoring it without the source is incoherent. Answered better on the substance by amendmentA9: a mechanical, judgment-free manipulation check,G2b, computed byanalysis/checks.pyover the frozen texts before dispatch. It counts paragraph splits, one-sentence paragraphs, paragraph-initialAnd, barehe said, italic spans, four-dot ellipses and the doubling repetitions.CARRY45,FLAT7,STRANGE7,WRONG7 —CARRYstrictly higher in 7 of 7 segments, and the residue of 7 is the same in all three derived arms, so it cannot differentiate them. This is a description of frozen materials, not a threshold on an outcome, and it is a stronger manipulation check than any rater. - ADVISORY 2 — make
G7gating at ≥ 0.80. Overruled. A formal-device catalogue is not a closed set; an independent model naming fewer devices is weak evidence about the lead's list rather than strong evidence against it.G7stays reported and non-gating at 0.60, as atRS-20260807fG5. - ADVISORY 6 — read a positive
G4/P3as "the sense is oddity-sensitive, not source-sensitive". Not overruled so much as adopted as the result's interpretive frame: that is precisely the conclusion the run is built to license or refuse, and it is written into §5 of the result.
Revised ceiling: $1.40 (stage 5 adds 3 calls at max_tokens 12,000, worst case $0.40). Today's
headroom is $3.067.