Repository path: workshop/experiments/E-20260813h-carriage-or-elevation/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260813h-carriage-or-elevation |
| status | frozen |
| created | 2026-08-13 |
| updated | 2026-08-13 |
| senses | perceived-source-carriage, style-correspondence, naturalness |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-carriage-direction.md, wiki/findings/results/RS-20260813g-register-quadrants.md, wiki/findings/results/RS-20260813b-affect-yardstick.md, wiki/goodness-senses.md, workshop/translations/flipperne/R07-v1/translation.md, workshop/translations/flipperne/R08-v1/translation.md, workshop/translations/flipperne/R25-v1/translation.md, workshop/translations/flipperne/R06-v1/translation.md, workshop/regimes/R07-fluency.md, workshop/regimes/R08-resistancy.md, workshop/regimes/R25-ennoblement.md, config/models.md, config/budget.md |
E-20260813h — carriage or elevation: when a reader calls a translation source-following, is it tracking direction or distance?
ARM-carriage-direction step 1, T2 (Poetics). Translation limb: T-flipperne-R07-v1, frozen
at 28677c9 before this design existed. Study limb: this experiment.
1. The question, in one sentence
Venuti's central claim is that a foreignizing translation shows its reader that a foreign text is
there; RS-20260813g has just shown that on this tale the formal difference between a source-ward
rendering and an elevated one is about as large as a formal difference gets — 16 of 16 frozen Danish
compounds carried against 5 — and this experiment asks whether that difference reaches a reader as a
direction at all, or only as distance from plain English.
The rival is not a strawman. RS-20260813b §7 found that an elaborate, latinate, perfectly
grammatical English document was matched to a foreignizing arm at 7 of 12 and to a plain arm at 0,
and wrote the conclusion this design is built to test: the style question "appears to measure
distance from plain English rather than direction of markedness." RS-20260813g §10 then recorded
that its own census "shows only that there is a formal difference available to be told apart … and
says nothing about whether a reader uses it."
2. Why this is the charter's question and not the apparatus's
What this unit teaches about translating literature: whether the source-ward moves a translator can actually make — calquing a compound, keeping a clause order — reach a reader as source-driven, or whether any marked English reads the same way, in which case a handbook cannot justify a source-ward recommendation by anything the reader experiences.
framework/v0.2 §8 Q-e is blocked on exactly this. RS-20260813g §6 P4 established that
published practice supplies no precedent for a source-ward recommendation, so if the reader-side
warrant also fails, Q-e has neither leg and the framework must say so.
3. Materials — four renderings of one tale, three of them frozen weeks or hours before this design
Andersen, «Flipperne» (1848), 756 Danish words, 26 paragraphs. Copy-text as
E-20260813g/corpus/DA.txt, byte-identical.
| arm | regime | hand | frozen | words | S1 lead audit |
inversions |
|---|---|---|---|---|---|---|
R06 |
none (unruled single pass) | lead | S174 | 873 | 7 / 16 | 1 / 6 (blind, RS-20260813g) |
R07 |
fluency (Venuti's F1–F10) | lead | this session, 28677c9 |
911 | 5 / 16 | 0 / 6 (translator-reported) |
R08 |
resistancy (Venuti's R1–R10) | lead | S178 | 829 | 16 / 16 | 4 / 6 (blind) |
R25 |
ennoblement (E1–E8) | lead | S174 | 1,070 | 5 / 16 | 2 / 6 (blind) |
The S1 column is the never-primary lead audit column of RS-20260813g §8, recomputed here by a
script that copies that run's rule verbatim; it reproduces that run's published figures for R06,
R08 and R25 exactly (0.438, 1.000, 0.312). That reproduction is the only reason the new R07
figure is admitted beside them. It is not the blind seat coding and is never reported as such.
S2 is not recomputed. A first draft of the stage-1 script coded the six inversion sites by
regex and disagreed with RS-20260813g's blind column on three of four checkable arms. It was
dropped rather than repaired, and the inversion figures above are quoted from the blind coding
for the three old arms and from the translator's own log for R07, labelled accordingly.
Five matched segments, by paragraph, frozen before any prompt was built:
G1 ¶1–4 · G2 ¶5–11 · G3 ¶12–17 · G4 ¶18–23 · G5 ¶24–26. Per-arm lengths in
analysis/stage1.json; R25 runs 20–47% longer than the others in every segment, which §11 treats
as a live confound and not a footnote.
4. The dependence table, computed before the design was frozen
Whole texts, all six pairs among the four lead arms (analysis/stage1.json):
| pair | 7-grams | 12-grams | longest run | verdict |
|---|---|---|---|---|
R08~R25 |
8 | 0 | 11 | clean |
R06~R25 |
7 | 0 | 11 | clean |
R07~R08 |
19 | 0 | 10 | clean |
R06~R08 |
42 | 6 | 15 | DEPENDENT? |
R07~R25 |
43 | 4 | 13 | DEPENDENT? |
R06~R07 |
253 | 155 | 41 | DEPENDENT? |
Both jury pairs are clean, which is the fact the design needed and did not control.
R06~R07 at a 41-token contiguous run is the largest lead-self-overlap ever measured in this
project — above note (bhb)'s previous record of 37, and above any lead-versus-published figure
ever recorded. It is not evidence that Venuti's fluency rules describe what an unruled pass does
anyway, because R07's contamination declaration records that R06 was read whole, in this
session, hours before R07 was written. The measurement is reported; the inference is refused.
Neither R06 nor R07 is in a jury cell with the other.
5. Design
Three conditions, five segments, three seats — 45 cells, one call each. Every call presents two renderings of one segment as Version A and Version B and asks one question with a closed answer.
| id | pair | source shown | role |
|---|---|---|---|
AB |
R08 vs R25 |
no | PRIMARY |
AS |
R08 vs R25 |
yes, the Danish | GATE G1 |
CB |
R06 vs R25 |
no | matched-carriage probe |
Seats: P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5.
P5 is excluded on note (bne); P4 is the pre-run critic and takes no cell.
Presentation order — which arm is Version A — is fixed per cell by
sha256("E-20260813h|" + condition + "|" + segment + "|" + seat) mod 2, so it is reproducible and
neither arm is systematically first.
The question, identical in all three conditions (the source block is prepended in AS only):
Two English translations of the same passage from a nineteenth-century short story are printed below. They translate the same original, which is in another European language. Neither is the original.
One of these two translators set out to follow the original's own way of putting things as closely as English allows. The other set out to do something else.
(Amended in place before dispatch on critic BLOCKING 1, §12. The wording frozen at 4bdd86b
glossed the phrase as "— the way it builds its words, and the order in which it arranges its
clauses —", which named the two measured features and is struck. What is quoted above is what was
dispatched, and run/cells.json is the record.)
Which of the two followed the original's own way of putting things?
Answer on ONE line, in exactly this form, and print nothing else:
AorB, then a semicolon, then at most fifteen words naming the one feature you used.
The language is withheld in AB and CB on purpose. Danish is Germanic, and a seat told so
could answer "pick the more Anglo-Saxon one" without reading for source-carriage at all. AS must
reveal it, which is one reason AS is the gate and not the primary.
Output form is closed by construction — one letter, a semicolon, ≤ 15 words — which is note (bng)'s prescribed remedy, and it is not a length that varies with the input. Caps are still sized from a pilot (§8).
6. The two hypotheses, and they make opposite predictions on the primary
The translation limb exists to make the rival quantitative. R07 is the declared plain-English
yardstick — a rendering made under a frozen fluency rule set — and distance from it is computed
per segment as Jaccard distance on token types, with mean word length as a robustness column
(analysis/stage1.json, computed before this design was frozen):
| arm | mean Jaccard distance from R07 |
mean word length gap |
|---|---|---|
R06 |
0.3331 | +0.062 |
R08 |
0.5588 | +0.325 |
R25 |
0.6383 | +0.439 |
Both metrics order the arms identically, and in every one of the five segments taken separately.
- H-DIRECTION — readers track which way a translation is marked. Predicts the arm with the
higher
S1:R08inABandAS;R06inCB. - H-DISTANCE — readers name whichever version is further from plain English, per
RS-20260813b§7. PredictsR25in all three conditions.
On AB, the primary, the two hypotheses predict opposite arms. R08 carries 16 of 16 compounds
and sits 0.0795 nearer the plain yardstick than R25 does. That opposition is the whole reason
this pair is the primary and is why R07 had to be written first: without a declared plain pole,
"distance from plain English" is an objection, not a prediction.
On CB the opposition is the same in sign and weaker in size — R06 carries 7 of 16 against
R25's 5, a two-site margin — so CB is registered as a probe, not a second primary.
7. Registered predictions, gates and failure criteria
G1 (gate, condition AS). With the Danish on the page, seats identify R08 in ≥ 12 of 15
cells. If G1 fails, P1 and P2 are WITHHELD and the run is reported as an instrument
failure: a jury that cannot see source-carriage with the source in front of it licenses no reading
of what it does without one. No override is available.
P1 (PRIMARY, condition AB, registered before dispatch). Seats identify R08 in ≥ 12 of
15 cells (exact binomial against p = 0.5: P = 0.0176). Fires → H-DIRECTION survives on this
pair and H-DISTANCE is refuted.
P2 (registered, condition AB, the opposite tail). Seats choose R25 in ≥ 12 of 15
cells. Fires → H-DISTANCE survives and H-DIRECTION is refuted: the arm carrying 5 of 16
compounds was called the source-follower over the arm carrying 16 of 16.
P1 and P2 cannot both fire. Neither firing is a real and informative outcome — it means the
question is answered at chance by seats that can answer it when the source is present, which would
make perceived source direction unavailable to a source-blind reader rather than mis-assigned.
That outcome is registered here so it cannot be reported afterwards as a null of the boring kind.
P3 (probe, condition CB). Seats choose R25 in ≥ 12 of 15 cells. Fires → elevation is read
as source-following even where the elevated arm carries less of the source's word-formation than
its partner.
P4 (exploratory, all 45 cells). The per-cell agreement rate between the seats' choices and
H-DISTANCE's prediction is reported for each condition, with the same for H-DIRECTION. No bar.
Failure criteria, fixed now:
F1—G1fails →P1,P2,P3all withheld, reported as numbers licensed for nothing, followingRS-20260813c§6 andRS-20260813g§7 precedent.F2— any seat whose reply cannot be parsed toA/Bin > 20% of its 15 cells has all its cells struck and the primary recomputed on the remaining seats, with the strike reported.F3— a cell returningfinish_reason: lengthis re-dispatched once at double the cap; the cost of both attempts is accumulated (theRS-20260813b§8 defect, repaired at S178). A cell dead after two attempts is struck.F4— more than 10% of cells struck in any condition → that condition is withheld entirely.F5(position lock) — a seat choosing the first-presented version in ≥ 14 of its 15 cells is reported as position-locked and excluded fromP1,P2andP3, which are recomputed.F6— if the verifier finds any prompt containing a banned provenance string (§9), the affected condition is void.
8. Procedure
- Stage 1 (
analysis/stage1.py, already run, committed with this design): dependence,S1lead audit, segments, the distance predictor. - Pre-run critic,
P4moonshotai/kimi-k3, given this design whole. Findings applied before dispatch; every finding recorded with accepted/overruled and a reason. - Pilot: three calls — one per seat, condition
AB, segmentG1— dispatched first, and the observed completion-token counts are read before the remaining 42 are sized and sent. This is note (bng)'s remedy applied as a measurement rather than a guess. Pilot cells are kept as real cells; they are not discarded and re-run. - Dispatch the remainder. Raw bodies stored per cell under
run/. analysis/analyse.pycomputes every reported number;analysis/verify.pyrecomputes them from the stored bodies, importing nothing fromanalyse.py.
9. Blinding and banned strings
Seats see only the two rendered segments (and the Danish in AS). The verifier asserts that no
dispatched prompt contains any of: R06, R07, R08, R25, resistancy, fluency,
ennoblement, Venuti, foreignizing, domesticating, Andersen, Flipperne, Danish,
compound, calque, lit-trans, or the name of any regime rule. Danish is banned in AB and
CB prompts and necessarily present as text in AS, where the assertion is instead that the
word Danish does not appear in the instruction block.
10. Budget
Declared ceiling: $0.70. UTC day 2026-08-13 stands at $4.268792634 of $5.00 across seven sessions, leaving $0.731207366. This run is sized to fit that headroom and is dispatched on the day it is designed; if any part of it is dispatched after 00:00 UTC it is ledgered against the day of dispatch and the split is stated.
Worst case built from max_tokens, per note (abc), at the pilot-confirmed caps:
| item | calls | worst case |
|---|---|---|
pre-run critic P4, cap 4,000 |
1 | $0.072 |
P1 cells, cap 1,200 |
15 | $0.119 |
P2 cells, cap 2,500 |
15 | $0.152 |
P3 cells, cap 1,200 |
15 | $0.131 |
| total worst case | 46 | $0.474 |
P2's cap is set higher because config/models.md's selection probe records ~1.7k tokens of hidden
reasoning on that seat; sizing it at 1,200 would be exactly note (bng)'s defect committed again.
Effort is pinned low on all cells. Lead translation of R07 is $0 and is not ledgered.
11. Known limits, written before the numbers exist
- One tale, one author, one language pair, one execution of each rule set by one translator, who also chose the segments and wrote the question. Nothing here generalises past that without replication, and the arm page names replication as its step 2.
R25is 20–47% longer than its partner in every segment. A seat could be answering "which is longer". This is not controllable without rewriting a frozen arm, so it is registered as a confound, andP4's reported per-segment agreement is the only handle on it: length ordering is constant across all five segments, so length cannot explain variation between segments, only a constant bias.- The distance yardstick is not independent of
R06.R07shares a 41-token run with it, soR06's low distance is partly a measure of how muchR07reproduced it. This inflates theCBgap in H-DISTANCE's favour — conservative for H-DIRECTION, which is the hypothesis this project would like to be true, and stated for that reason. - A forced binary choice with a false premise. In
CB, and arguably inAB, neither arm may be what the question describes. The design accepts this: the question is what a reader does with a forced attribution, which is the form every real reading takes. - Nothing here is about quality. Tier D is NOT PASSED; every evaluative sentence is
internal-judgment-onlyandprovisional.perceived-source-carriagemay not be cited as evidence that anything was carried (wiki/goodness-senses.md, four measured false-positive figures), and this design does not cite it so. - The seats are four models sharing a training distribution. Convergence among them is not evidence about human readers, and the sense entry's standing warning applies unchanged.
12. Pre-run critic pass — PROCEED-WITH-AMENDMENT, 5 findings, 4 BLOCKING, all five accepted
P4 moonshotai/kimi-k3, cap 4,000, effort low, stop, $0.071944200. Raw body:
run/critic.json. Amendments applied before any cell was dispatched; the design above is
amended in place only where §12 says so, and this section is the record of what changed and why.
Finding 1, BLOCKING — the prompt taught the answer. ACCEPTED IN FULL. The instruction glossed
"the original's own way of putting things" as "the way it builds its words, and the order in which
it arranges its clauses" — which is a verbatim description of the two features RS-20260813g
measured and R08 was written to maximise. A firing primary would then have shown only that seats
can apply a handed-over rubric. The gloss is struck from all three conditions, in identical
wording, leaving: "One of these two translators set out to follow the original's own way of putting
things as closely as English allows. The other set out to do something else." The critic's fallback
— keep the gloss in the gate only — is refused: a gate that asks a different question from the
primary is not a gate on it. The cost is accepted and named: with no gloss the task is harder, G1
is likelier to fail, and a G1 failure withholds the whole run. That is the right risk to take.
Finding 2, BLOCKING — the binomial arithmetic treated clustered cells as independent. ACCEPTED. Seats are fixed, not sampled, and five cells from one seat are not five draws. The primary unit is re-registered as the seat. The bars in §7 are replaced by:
G1(gate,AS) — each of the three seats identifiesR08in ≥ 4 of its 5 segments.P1(primary,AB) — each of the three seats identifiesR08in ≥ 4 of its 5.P2(AB, opposite tail) — each of the three seats choosesR25in ≥ 4 of its 5.P3(probe,CB) — each of the three seats choosesR25in ≥ 4 of its 5.
Per-seat exact binomial against p = 0.5 is 6/32 = 0.1875; the three-seat conjunction is 0.0066 if the seats are treated as independent, which they are not entirely — they are three models sharing a training distribution, and the inference is to these three models and not to a population of readers. The 15-cell counts and the full seat × segment grid are reported as secondary and descriptive, with their exact binomials given and labelled as the clustered figures they are.
Finding 3, BLOCKING — post-strike recomputation was unspecified and therefore gameable. ACCEPTED. The full decision table is fixed here, before any body exists:
| event | consequence |
|---|---|
one seat struck (F2 or F5) |
primary recomputed as each of the remaining 2 seats ≥ 4 of 5 (joint P = 0.0352), reported as the weaker bar it is, and the strike named in §1 of the result |
| two or three seats struck | condition withheld entirely; its prediction is declared uninformative |
F4 fires (> 10% of a condition's cells struck) |
condition failed, not recomputed |
Finding 4, BLOCKING — the distance predictor is confounded with length. ACCEPTED, and it changes
a registered prediction. Jaccard on token types rises mechanically with type count, and R25 has
26–37% more types per segment than its partners. A length-corrected column is therefore computed
and registered: overlap coefficient distance, 1 − |X ∩ Y| / min(|X|, |Y|), which is insensitive
to one set being larger. Computed before dispatch:
| arm | corrected distance from R07, mean |
per segment G1…G5 |
|---|---|---|
R06 |
0.1860 | 0.116 · 0.205 · 0.215 · 0.200 · 0.194 |
R08 |
0.3705 | 0.318 · 0.364 · 0.432 · 0.385 · 0.354 |
R25 |
0.4007 | 0.458 · 0.346 · 0.441 · 0.373 · 0.385 |
The mean ordering survives, and the per-segment ordering does not: on G2 and G4 the corrected
metric puts R08 further from plain English than R25. So H-DISTANCE's registered prediction is
now stated on both metrics and they disagree on two of five segments:
And P2's and P3's firing interpretations are reworded as the critic requires. A firing P2
now licenses only: the longer and more elevated arm is read as source-following, with length and
elevation not separated by this design. The attribution to elevation alone is available only if
the corrected column also predicts R25 in the segments where the seats chose it — which is a
three-of-five subset, reported segment by segment.
Finding 5, NON-BLOCKING, all three parts accepted.
(a) F5's position-lock bar is tightened from ≥ 14 of 15 to ≥ 13 of 15 first-presented
choices (exact P = 0.0037 under chance), because 14 of 15 let a nearly position-locked seat through.
(b) The ≤ 15-word stated feature is tabulated, per condition and per seat, and printed in the
result — a seat that picks the longer arm and calls it "compounds" is then visible rather than
hidden inside a count. (c) The asymmetric evidence standard on R07's inversion figure is stated
wherever the figure appears: 0 of 6 is translator-reported, while every rival figure in the same
table is blind-coded, and R07 sits in no jury cell precisely because of standards like this one.
Nothing was overruled. No new seat, no new translation and no additional budget was needed; the declared ceiling of $0.70 stands.