Repository path: workshop/experiments/E-20260812f-affect-unprompted/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260812f-affect-unprompted |
| status | frozen |
| created | 2026-08-12 |
| updated | 2026-08-12 |
| senses | affect |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-affect-unprompted.md, wiki/goodness-senses.md, wiki/decisions/resolved/D-20260804-16-affect-two-halves.md, wiki/findings/results/RS-20260804h-affect-halves.md, workshop/translations/bai-ganyo-opera/R06-v1/translation.md, workshop/translations/bai-ganyo-opera/R07-v1/translation.md, workshop/translations/bai-ganyo-opera/R08-v1/translation.md, workshop/regimes/R07-fluency.md, workshop/regimes/R08-resistancy.md, config/models.md, config/budget.md |
E-20260812f — do the two halves of affect separate when nobody has pointed the translator at either?
Frozen 2026-08-12, before any call was dispatched. The three renderings this design reads
were written and committed at 2e382d9, before this file existed.
1. The question, and the standing obligation it discharges
wiki/goodness-senses.md usage rule 4 — ratified B on 2026-08-04 as D-20260804-16, over an
adversarial review that returned A — requires an evaluation invoking affect to say which
half of the sense it scores: the experience produced in the target reader, or the
comparability of that experience to the one the source produces. The rule was ratified under six
binding conditions, and condition 3 is still open sixty sessions later:
it reverts, and
D-20260804-16reopens, unless a later run establishes unprompted separation under purposes that pass a paraphrase-level check.
The evidence the rule stands on (RS-20260804h-affect-halves) separated the halves by pointing
a translator at two readers — two frozen purpose specifications, which the run's own
independent critic found to paraphrase a half of the definition apiece. The adversarial review's
sentence was: "the distinction was not demonstrated; it was commissioned." Nothing since has
asked whether a translator who was pointed at neither reader produces renderings the two halves
select differently.
One sentence on what this teaches about literature (subject rule, wiki/tracks.md and
continue-prompt.md §4.5): it measures whether making the joke land in English and making the
English report what the joke does in the original are two different targets a translator has to
choose between at ordinary sites, or one target under two descriptions.
2. Materials
Source. Алеко Константинов, «Бай Ганьо в операта», from Бай Ганьо. Невероятни разкази за
един съвременен българин (1895). Public domain (Konstantinov 1863–1897). Copy-text
workshop/translations/bai-ganyo-opera/src.txt, from Bulgarian Wikisource, fetched 2026-08-12.
9 paragraphs, 677 Bulgarian words. The project's first Bulgarian source and its
nineteenth source language. Comic prose: a Sofia student in Vienna tells a company of
listeners how his provincial travelling companion disgraced him twice in one day.
The three arms, all lead, all $0, all frozen at 2e382d9 before this design existed:
| arm | regime | what it is |
|---|---|---|
R06 |
lead single pass | source only, no rule set, no return pass |
R07 |
fluency | R06 + Venuti's ten domesticating rules F1–F10, frozen 2026-07-28 |
R08 |
resistancy | R06 + Venuti's ten foreignizing rules R1–R10, frozen 2026-07-28 |
Why these three are the right arms for the word unprompted. F1–F10 and R1–R10 were written
at S048 for a different question entirely (RS-20260728e-venuti-specifiability), fifty-nine
sessions before affect's halves were distinguished. Neither rule set contains any sentence about
the reader's experience, about what a passage does, or about comparability to the source's
effect; R06 contains no rule set at all. That claim is not left to my reading of them: G2
below puts it to an independent seat as a paraphrase-level check, which is the exact form
condition 3 requires.
Segmentation. The chapter is cut into 15 segments at sentence boundaries, identically in
all four texts, by code.py, which reconstructs each arm's body from its segments and asserts
byte equality. Segments are the unit of analysis.
3. Procedure
Stage 0 — the two gates, run before anything else
G2— paraphrase-level check. Two source-side seats (P2, P3), neither of which judges in stage 2, are shown (a) the two halves of theaffectdefinition verbatim, and (b) each regime's procedure and rule set verbatim, and asked whether any rule, read at the level of paraphrase, instructs the translator toward either half. The primary is withheld if either seat says yes for any of the three regimes (F1).G0— segment classification, from the Bulgarian alone. The same two source-side seats (P2, P3) see the 15 numbered Bulgarian segments and no English, and answer for each: does this segment's effect on a Bulgarian reader depend on something English does not have — a word borrowed from another language and used as a social marker, a mangled foreign phrase, a slur, a diminutive, an untranslatable idiom, a fragment of spoken telling? Yes/no, with the feature named.FORK= both seats yes.NO-FORK= both seats no. Disagreements are dropped and reported. The classification is therefore not mine and does not see the renderings.
Stage 1 — the yardstick
For each segment, seat P2 sees the Bulgarian alone and no English at any point, and writes
an open description of what this passage does to a reader of the Bulgarian, and what in the words
does it. This is the open effect-description form that the affect entry's own standing
methodological caution prescribes and that agreed at 12 of 12 in E-20260804h; the yes/no coding
of a critic's category is the form that has failed twice and is not used. Seat P5 writes a
second description of every segment, used only for an agreement check and never shown to a
judge; P5 also judges in stage 2, in stateless calls, and never sees its own description. G3: a yardstick containing any Cyrillic character is void for that segment (F4).
Stage 2 — the two questions, in separate stateless calls
Three judging seats — P1, P4, P5 (config/models.md; all non-Anthropic). P2 and P3 are
the source-side seats and do not judge; P3 is additionally kept off the high-volume stages under
note (bmp), its reasoning not being bounded by max_tokens on its provider. For every segment, each seat answers both
questions, in separate calls with no shared state, with the three arms relabelled A/B/C in an
order drawn per (segment, question, seat) from a fixed seed:
H1— the first half. The three renderings, no source, no yardstick, no mention of translation: which does the most to you as a reader of English? A full ranking of the three, best first, and one sentence on the first.H2— the second half. The same three renderings plus the stage-1 yardstick'seffectfield: which comes closest to doing to its reader what the original does to its reader? A full ranking, closest first, and one sentence on the first.
Why a ranking and not a pick — amendment 6, made after G2 returned and before any judge call
was dispatched. G2 did not clear all three arms (§9a). A ranking makes the primary computable
on whichever subset of arms the gate cleared, because a pairwise preference between any two
arms is derivable from every cell, while a single top pick is not. It also removes three-seat ties
from the two-arm comparison. No judge data existed when this was written.
Stage 3 — parity control
G1. One seat (P1) sees the three renderings of each segment and answers whether
they state the same facts — same events, same actors, same objects — naming any fact present
in one and absent from another. Register/lexis/syntax differences are not parity failures; the
arms are built to differ there. The primary is withheld if more than one FORK segment fails
parity (F2).
4. Predictions, registered before dispatch
Let disagree(class) be the fraction of segments in that class at which H1's three-seat majority
pick and H2's three-seat majority pick are different arms. Three seats, three arms, so a
majority may fail to form; a segment with no majority under either question is excluded from the
primary and reported.
P1(primary), as amended on the pre-run critic's MAJOR 5 before dispatch.disagree(FORK) − disagree(NO-FORK) ≥ 0.50and the exact permutation P ≤ 0.05, where the permutation enumerates every assignment of the observed class sizes over the surviving segments.disagree(FORK) ≥ 0.75anddisagree(NO-FORK) ≤ 0.25were registered as conjunctive gates in the first freeze and are demoted to reported figures: with at most fifteen segments they would have made a single noisy cell decide the primary, and the permutation P already controls for the class sizes. This relaxation makesP1easier to satisfy and was made before any call was dispatched.P2(direction). AtFORKsegments,H1's majority isR07(the domesticating arm) andH2's majority isR08(the foreignizing arm), each at ≥ 2/3 of FORK segments.P3(seat-level). The localization margin ofP1is positive within each seat taken alone, at ≥ 2 of 3 seats.P4(yardstick agreement). P2's and P5's descriptions of the same segment are judged compatible — this is reported, not gating, and is not a prediction the primary rests on.
Why the primary is a difference and not a rate. With three arms, two independent random picks
differ two times in three; "the halves picked different arms" is therefore not evidence on its
own. The claim is localization: the halves diverge where the translator faced a real fork and
converge where he did not. That is the shape RS-20260804h measured (six of six against six of
six) and the shape a confound of mere size cannot manufacture — the arms differ most at FORK
segments, but a bigger difference moves both questions the same way unless the questions are
asking different things.
5. The registered result → option map
Required verbatim by usage rule 4's sixth binding condition (every future motion on this sense registers its result→option map before dispatch). Written here before any call.
| outcome | what follows |
|---|---|
P1 holds and F1–F5 all clear |
Unprompted separation is established on this material. Condition 3 of D-20260804-16 is discharged; usage rule 4 stands without the reversion clause hanging over it. ARM-affect-unprompted step 2 writes it into wiki/goodness-senses.md §affect with the figures and the limits. |
P1 fails with margin ≤ 0.25, or with margin ≥ 0.50 but permutation P > 0.05 |
The honest null. Unprompted separation is not established on this material. Condition 3 is not discharged and the entry records a first failed attempt against it. Reversion is not automatic and is not this session's to take: condition 3 has no deadline, one work is one work, and reverting a ratified rule is a motion. Step 2 opens that motion for a later session to ratify. |
| margin strictly between 0.25 and 0.50 | Withheld. Nothing is written into the entry and no motion is opened; the result page reports the numbers and says the design did not resolve. |
any of F1, F2, F5 fires |
Primary withheld regardless of the numbers; the run reports what the gate found. |
F3 fires (fewer than 4 segments in either class) |
The design cannot run; report the classification and stop. |
P2 and P3 never change the option taken. They qualify what may be written under it.
6. Failure criteria
F1—G2finds an arm paraphrasing or aiming at either half of the definition. Amended afterG2and before any judge call:F1no longer voids the run wholesale. An armG2does not clear is excluded from the primary and reported; the primary is computed on the cleared arms, and is withheld entirely if fewer than two arms clear. The wholesale reading was written when I expected the gate to pass; it would have thrown away the comparison the gate itself licenses.F2—G1finds more than oneFORKsegment failing propositional parity.F3— fewer than 4FORKor fewer than 4NO-FORKsegments surviveG0.F4— a yardstick contains Cyrillic; that segment'sH2is void and the segment is dropped.F5— more than 5% of dispatched bodies fail to parse, or any body isfinish_reason: lengthafter the one licensed re-dispatch.
7. Cost
Worst case built from max_tokens, per note (abc), and from the exact prompt lengths, by
run.py --dry-run. Seats and prices from config/models.md.
Declared ceiling: $1.40. Today's UTC ledger stands at $2.116807244 of $5.00 before this run; the ceiling fits the remaining $2.883192756 with margin. If the dry run exceeds the ceiling the run is scaled down, not re-priced.
Note (bmp) is implemented here, not merely cited. run.py writes runs/_config.json on its
first call — seats, slugs, caps, extras, seed, git HEAD — and refuses to start if an existing
_config.json differs from the configuration it is about to use. A second runner with a changed
seat configuration cannot silently share this output directory.
8. Declared exposures
- The three arms are one hand in one session, made in the order R06 → R07 → R08. Each is downstream of its predecessors; the ladder is matched on translator, source and reading depth and unmatched on order. Every existing ladder in this project has the same shape.
- Contamination is
suspectedand unmeasurable — the only English Бай Ганьо is in copyright and unreachable, so no overlap figure exists in either direction. No claim here turns on independence from it, because the contrast is within-lead. FORKis expected to coincide with the Turkish-and-oral-register stratum, which is also where the two rule sets differ most. The design cannot separate the halves diverge here from the arms differ most here; what it can do is show that the divergence is a change of which arm wins and not merely of margin, andP1's form is chosen for that reason.- Model seats are not readers,
affectisuntested, and Tier D is NOT PASSED (config/models.md). Everything here isinternal-judgment-onlyandprovisional. No jury verdict carries evidential weight, and nothing in this run may be called calibrated. - A positive result has a rival reading the pre-run critic named and the design cannot exclude:
that what separates the halves here is the domestication–foreignization axis rather than the
absence of a brief, the two regimes being that axis executed.
G2is the only thing standing between those readings, andG2is three seats' judgment of two rule sets. - The
R07interpretive rulings were written after the source was read, which the regime's S073 binding condition forbids. Declared onT-bai-ganyo-opera-R07-v1with the two things that limit the freedom it gave.
9. The pre-run critic pass, and what it changed
qwen/qwen3.7-max, non-panel, no role in this run. Verdict NEEDS-REDESIGN: 5 findings, 2
BLOCKING, 3 MAJOR. All five accepted; none overruled. Raw at runs/critic.json and
runs/critic.txt. Every amendment below was made before any study call was dispatched.
- BLOCKING — the yardstick turned
H2into feature-matching. The stage-1 prompt asked for the effect and what in the words does it, so a describer forbidden to quote the Bulgarian would name the mechanism ("uses a Turkish loanword"), and a judge could then pick the arm containing that mechanism without weighing any experience at all. Fixed by splitting the yardstick into two fields —effect, which may not mention any property of the language (an explicit list: loanwords, borrowings, dialect, slang, slurs, register, diminutives, idioms, syntax, fragments, tense, spelling, translation, other languages), andcause, which may. Onlyeffectis ever shown to a judge.causeis preserved and reported. - BLOCKING — unprompted may be false, because domestication and foreignization are arguably
the two halves under other names. The critic is right that a rule set aimed at the target
reader's ease is a near-neighbour of the experience produced in the target reader, and one
aimed at retaining the source's forms a near-neighbour of comparability to the source's
effect. The design already contains the instrument that decides this —
G2— butG2's prompt had been written so that it could not fire: it told the seats that a rule about words, grammar, register, foreignness or readability does not count merely because effects follow from words, which exempts exactly the equivalence at issue. That sentence is struck, and a second question is added at the level of the whole strategy — taken together, is what this procedure is FOR the same thing as either half? — withF1firing on any of the four answers. IfG2fires, the primary is withheld and the finding is that Venuti's two poles areaffect's two halves in disguise, which is a result and not a failure. - MAJOR —
G0's checklist made the classification circular. The prompt listed the features that count asFORK, which are the featuresR08retains, soFORKwas guaranteed to be where the arms differ most. The checklist is struck; the seats now judge holistically and must write a reason in their own words before answering. - MAJOR —
H1's does the most invites is best written. Kept, because narrowing it to a named emotion would be the commissioning error the run exists to avoid, but disambiguated: this is not a question about which is written best, which is most correct, or which is easiest to read. - MAJOR — the primary was underpowered against its own thresholds. Accepted;
P1is restated above with the two absolute rates demoted to reported figures.
One finding is recorded and not fixed, because it cannot be: the critic's point 2 also names an alternative reading of a positive result — that this run measures whether fluency versus resistancy maps onto the halves rather than whether an unpointed translator forks them. That reading is attached to §8 as exposure 6 and travels with any citation of this run.
9b. Amendment 8 — F3 fired on the classification, and the fallback primary registered in its place
Written after G0 returned and before a single judge call was dispatched. No H1 or H2 body
existed when this section was written; git log fixes the order.
G0 came back almost uniformly NO-FORK. Two competent readers of the Bulgarian, asked
holistically and with the critic's checklist removed, judge that an ordinary English sentence could
carry the effect of nearly every segment in this chapter — including the segment in which the
character says «санким» and «гут моргин», where the translator's own frozen log had recorded that
English has no equivalent. F3 therefore fires and the localization primary P1 is dead: with
fewer than four FORK segments there is nothing to localize against.
The localization figures are still computed and reported as measured, and the classification itself is a finding of this run rather than a failed control.
P1b — the fallback primary, registered here. Over every segment with a usable yardstick, take
the three-seat majority preference between the two arms G2 cleared, once under H1 and once
under H2. A segment is discordant when the two halves prefer different arms. P1b holds
when the discordant segments are one-sided at exact two-sided binomial P ≤ 0.05 — i.e. when the
half that asks what does this do to you and the half that asks how close is that to what the
original does systematically prefer different renderings of the same passage.
This is a weaker claim than P1: it establishes that the halves separate, not that they separate
where the translator faced a fork. The result→option map (§5) is read against P1b in P1's
place, with the weakening carried into whatever is written. P2, P3 and the reported
localization figures are unchanged.