Repository path: workshop/experiments/E-20260804h-affect-halves/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260804h-affect-halves |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| senses | affect |
| provisional | true |
| links | wiki/arms/ARM-affect-reception.md, wiki/goodness-senses.md, wiki/findings/results/RS-20260804b-unsmiling-gravity.md, workshop/experiments/E-20260804h-affect-halves/purposes.md, workshop/translations/hameleon/R11-stage-v1/translation.md, workshop/translations/hameleon/R11-seminar-v1/translation.md, workshop/translations/hameleon/R14-v1/translation.md, config/models.md, config/budget.md |
E-20260804h — do affect's two halves select the same translation?
Frozen 2026-08-04 (S107) before any API call. ARM-affect-reception step 2, T2.
1. The question, and why it is not a question about the apparatus
wiki/goodness-senses.md defines affect as: "The translation produces in its reader an
experience comparable to what the source produces in its reader: the joke lands, the dread
accumulates, the ending stings."
That sentence has two halves. (H1) the translation produces an experience in its reader.
(H2) that experience is comparable to the one the source produces in its. The sense is written
as though these name one property. RS-20260804b §12 found, on a matched pair differing by one
inserted clause, that they selected different arms — 33 of 42 for the funnier text, 40 of 42 the
other way for the source-match — and left the finding where step 1 could not decide it, because the
difference between those arms was an inserted narratorial joke, not a translator's choice.
What this unit teaches about translating literature (the subject rule, wiki/tracks.md): whether
a translator who aims at the effect and a translator who aims at the relation to the source are
aiming at the same thing, or at two things that come apart — and if they come apart, what the choice
between them costs, site by site, in real prose.
2. Materials
Source. Anton Chekhov, «Хамелеон» (Осколки, 1884), whole, 1,019 words, public domain.
Copy-text: ПСС 1974–1982 via ru.wikisource. Three редакции were fetched and diffed
(materials/variant-diff.json): the 1884 and 1886 printings carry a clause Chekhov later cut
(«Хрюкин виноват, что тронул её, а она бродячая»). The copy-text is followed.
Arms, all lead-authored, none published, none seen by any grader before:
| arm | id | what it is |
|---|---|---|
STAGE |
T-hameleon-R11-stage-v1 |
R11, purpose STAGE (heard aloud once, nothing in front of the listener, no gloss). Written first |
SEMINAR |
T-hameleon-R11-seminar-v1 |
R11, purpose SEMINAR (read on the page, must be able to say what the Russian is doing, no notes permitted). Written second |
FLAT |
T-hameleon-R14-v1 |
R14 matched-content flattening of STAGE, at the eight sites only. The positive control |
The purposes were frozen before rendering (purposes.md, commit d839d98) and are written as a
reader and a use, containing no sense id and no definitional wording of affect — R11's binding
constraint, checked by materials/check_specs.py, which fired once before the freeze on the word
comic and forced an amendment. Both translator's logs were frozen and committed (c943cae)
before this design existed.
The contamination gate ran between the freeze and this design, before any site was selected
(materials/gate.py → materials/dependence.json). Name-excluded 12-grams: garnett~stage 0,
garnett~seminar 0, seminar~stage 2. Raw: garnett~stage 60 shared 7-grams / longest run
16, garnett~seminar 45 / 16, seminar~stage 132 / 22. Both numbers are reported
because run length alone has failed as a proxy three times (CLAUDE.md).
Two things that matrix says, and they point opposite ways. The longest garnett~stage run —
"not a soul in the square the open doors of the shops and taverns look out" — is in paragraph 1,
the one paragraph on which an exposure event is declared: forty-four words of Garnett's opening
were displayed while the Gutenberg volume was being identified. Paragraph 1 is excluded from the
site list. But the longest garnett~seminar run is the same length, 16 tokens, sits in ¶8, and
no exposure to it occurred — a formulaic police sentence on which two translators converge at 16
tokens independently. The exclusion of ¶1 is therefore a precaution, not a demonstration, and the
per-site runs of every selected site are printed in §3 rather than used as a filter.
3. Sites
Selection rule, applied by materials/build_items.py and stated before the sites are named:
¶1 excluded (above); FORK = a paragraph at which both frozen logs record a decision and the
two decisions diverge in kind (substitute against retain-and-scaffold); NO-FORK = a paragraph
at which neither log records a decision taken to satisfy a purpose clause.
| site | ¶ | kind | the device the effect rides on | STAGE / SEMINAR |
|---|---|---|---|---|
N1-chase |
3 | no-fork | — plain narration | 93 / 109 |
F1-recognition |
5 | fork | the speaking name; вышеписанный; борзой щенок |
134 / 136 |
F2-khryukin-plea |
7 | fork | я человек, который работающий — a broken relative clause |
92 / 103 |
F3-tirade |
8 | fork | Я ему покажу Кузькину мать! — an untranslatable idiom |
108 / 113 |
F4-coat-on |
20 | fork | каждый свинья — a gender disagreement |
92 / 99 |
F5-brother |
25 | fork | братец ихний приехали — plural verb for one man |
33 / 37 |
F6-diminutives |
27 | fork | собачка · собачонка · цуцык — suffix morphology |
54 / 59 |
N2-crowd-laughs |
28 | no-fork | — plain narration | 20 / 18 |
Per-site longest Garnett runs (STAGE / SEMINAR): 7/10 · 11/9 · 5/5 · 14/16 · 8/9 · 4/2 · 5/8 · 6/4.
Key assignment is per site and per ordering, from seed 20260804, and o1's map is required to
differ from o0's at every site. A seat that answers globally has to track which key is which eight
times.
4. Procedure
Stage 0 — the yardstick, written by seats that see no English. Two independent non-Anthropic
seats are shown the eight Russian passages and nothing else — no translation, no crib, no gloss
— and asked, per site, what the passage does to a reader and what in it does that. This is
RS-20260804g's repair applied here (§7 of that page): the lead does not write the description its
own translations are marked against. A1 = P5 deepseek/deepseek-v4-pro, A2 = P4
moonshotai/kimi-k3.
Stage 1 — H1, "produces an experience". Three seats (P1, P2, P3) × two orderings. Shown: the eight sites, each with three keyed English versions. No Russian, no yardstick. Asked to rank the three keys 1/2/3 by which does more to them as a reader of English, with a verbatim quotation from the key ranked first.
Stage 2 — H2, "comparable to what the source produces". The same three seats × two orderings, in
separate stateless calls. Shown: the same eight sites, each with A1's effect statement and the
three keyed versions. No Russian. Asked to rank the three keys by which gives an English reader
an experience closest to the one described, with a verbatim quotation from the key ranked first.
A2's statements are never shown to a grader; they are FC6's check on A1.
Judgment is not parallelized. No seat judges its own output; the lead judges nothing.
5. Predictions, registered
- P1 — the split (primary). At FORK sites:
STAGEranks aboveSEMINARon H1 at ≥ 4 of 6, andSEMINARranks aboveSTAGEon H2 at ≥ 4 of 6. Both conjuncts required. Per site the value is the seat-majority over the six (seat × ordering) cells. - P2 — localization (the control that decides whether P1 means anything). The H1↔H2 disagreement rate — the proportion of cells where the two questions put a different arm first — is higher at FORK sites than at NO-FORK sites. If the two rates are equal or the no-fork rate is higher, P1 is reported as not localized and no motion is opened, because the split is then as consistent with the translator's hand as with the fork.
- P0 — the null, and it is a real outcome. If the same arm is ranked first on both questions at ≥ 5 of 6 fork sites, the clause names one thing on this material and the entry records that.
6. Failure criteria, registered
- FC1 — the instrument.
FLATmust rank last on H1 at ≥ 6 of 8 sites. Otherwise the graders are not registering an effect difference and H1 is uninterpretable; the run reports that and stops there. The bar is 6 and not 8 becauseN2-crowd-laughsis twenty words with one flattenable verb phrase and is very nearly unflattenable (T-hameleon-R14-v1§The operator sites). - FC2 — order. Any prediction whose sign reverses between
o0ando1is descriptive only. - FC3 — length.
SEMINARis longer thanSTAGEat 7 of 8 sites (674 words against 626). If the arm ranked first is the longer at 8 of 8 sites on a question, that question is reported length-confounded. Note the direction: a pure length preference predicts the same arm on both questions, so it works against P1, not for it. - FC4 — quotation. Every ranking must quote text present in the key it names. A seat below 0.90 verbatim is excluded from the primary and reported separately.
- FC5 — global preference. A seat giving an identical ranking at all eight sites in both
orderings on both questions is reported separately.
RS-20260804b§11.3 found exactly this behaviour in two of three seats. It too works against P1: a seat with one global preference answers both questions the same way. - FC6 — yardstick disagreement. A site where
A1andA2name incompatible effects is excluded from H2's primary.RS-20260804b§5 lost two predictions to this and the loss was the finding; it is registered here in advance.
7. What would change
If P1 and P2 both hold: affect's definition names two properties that a translator's choice
can separate, and an evaluation invoking affect must say which half it scores. A motion is
opened on the entry's wording — opened, not ratified: this session may not ratify a decision it
opened (charter §8).
If P0 holds: the clause survives its hardest available test and the entry records the null with the material that produced it. No motion.
If P1 holds and P2 fails: the split is real and unattributable, and the entry records it as unattributable. No motion.
8. Limits, written before the run
- L1 — the arms have one author, who knew the hypothesis. The deepest threat here is that the
split was written in rather than found: the same agent wrote both arms and could have made
SEMINARworthy and dull by construction. No control removes this.P2is the closest thing to a test of it — a hand that writes dull prose under one purpose writes it at the no-fork sites too — and the run is designed so that failingP2withholds the conclusion. - L2 — arm order.
SEMINARwas written second, by a translator who had already solved the story once (R11§6). Conservative on P1's first conjunct, anti-conservative on the second. - L3 — the purposes were written after the source was read. Declared in
purposes.md; the specs may have been chosen for a source that visibly affords this divergence. - L4 — no quality claim is made or implied. Tier D is NOT PASSED (
config/models.md). Every seat here is an uncalibrated reader reporting a preference, not a juror. Nothing in this design ranks these translations as translations, and nothing licenses a claim that either is better. - L5 — one story, one pair, one translator.
RS-20260804bwas ES→EN; this is RU→EN. Two is not a language-pair breadth claim. - L6 —
FLAT's propositional equivalence is unchecked by an independent seat and is the lead's own (T-hameleon-R14-v1§Equivalence). The failure direction is away from the run's conclusions.
9. Critic brief
The pre-run adversarial critic is asked for everything, and by name for two things:
(i) whether either purpose specification restates a half of affect in paraphrase — the
mechanical check catches literal strings only, and the S084 precedent is that a paraphrase got
through and would have manufactured the result; (ii) whether P2 is a real control on L1 or a
control that cannot fail.
9b. Amendments after the pre-run critic (applied before any grading call)
Critic: z-ai/glm-5.2, one pass, NEEDS-AMENDMENT, 3 findings (2 BLOCKING), $0.029674002. All
three accepted. Full text: critic.md.
A1 — from BLOCKING finding 1, accepted, and the specs are NOT rewritten. The critic read both
frozen specs and found each one paraphrasing a half of the definition: STAGE clause 5 ("a piece
that holds a room that has been drinking") names an effect on a reader; SEMINAR clause 3 ("say what
the Russian text is doing … what its words are up to") names the relation to the source's own
behaviour. The mechanical check passed them and the S084 precedent has repeated exactly as R11
warns it would. The specs are frozen and both arms are already rendered against them: rewriting
now would be a change to a freeze, which is worse than the defect. What changes instead is what
this run is allowed to conclude, re-registered here before any grading call:
- P1 is downgraded. If it holds, it does not show that a translator's choice separates the halves in general. It shows that when a translator is pointed at these two readers — readers the critic has shown were described in terms that lean towards the two halves — the resulting texts are selected differently by the two halves.
- P0 becomes the strong outcome. A null here, with specs that lean this way, is much harder to explain away than a split is. If the same arm is ranked first on both questions, the clause has survived a test loaded against it.
- The motion bar is raised. A motion on
affect's wording opens only if P1 and P2 both hold andFC7does not fire, and the motion page must carry this finding in its own text.
A2 — from BLOCKING finding 2, accepted. The critic showed P2 could not fail: with two no-fork
sites, one of them (N2) twenty words long and nearly unflattenable, the no-fork disagreement rate
was capped near 0.50 while the fork rate could reach 1.00, so P2 would have passed with the
artefact present. Four no-fork sites are added — ¶2, ¶6, ¶14, ¶18 — giving six fork and six
no-fork sites, and R14 is applied to the four on the same terms. Balance is now in count, not
in length: the no-fork sites are short exchanges (19–109 words) and the fork sites are mostly
paragraphs, which the critic asked for and this material cannot supply.
P2's bar becomes a margin: fork disagreement rate − no-fork disagreement rate ≥ 0.25.
A3 — new failure criterion FC7, from the second half of BLOCKING finding 2. The critic went
past the design and read the prose, and reported that SEMINAR looks written to lose at a no-fork
site — naming "The yelp of a dog is heard", "After her chases a man", "Sleepy physiognomies
protrude" against STAGE's "A dog yelps", "After it runs a man", "Sleepy faces poke out". That
is L1 with an address. It is now falsifiable:
FC7 — the artefact check. If
STAGEranks aboveSEMINARon H1 at ≥ 5 of 6 no-fork sites, the global-liveliness artefact is present, and P1's first conjunct is reported as not attributable to the fork whatever P2 returns.
A4 — from ADVISORY finding 3, accepted. FC1 is restated as what it is: a sanity check on the
graders, not a falsifiable test of anything. FLAT is a flattening of STAGE, so FLAT ranking
last is expected; its failure would invalidate H1, its success certifies nothing. Bar rescaled to
≥ 9 of 12 sites.
A5 — mechanical. materials/build_items.py regenerated at twelve sites; key maps re-drawn from
the same seed with the same per-site, per-ordering constraint.
9c. A deviation from the frozen design, found while the run was in flight
Stage 2 says "No Russian." The stage-2 prompt contains Russian. Not the passages — those are
withheld as designed — but the yardstick statements A1 wrote quote the Russian they are about,
often several times per site (не пущай, который работающий, Кузькину мать, собачонка,
цуцык). Those quotations travel into the stage-2 prompt inside the description.
This is recorded as a deviation, not a limit, because the frozen text says otherwise and the run
did not stop to fix it. What it risks: a stage-2 seat that reads Russian can partially reconstruct
the source from the fragments and answer an accuracy question instead of the registered one. The
direction is not neutral — the SEMINAR arm is the one that keeps source-specific items visible in
English, so a seat drifting towards accuracy would favour SEMINAR, which is the direction of P1's
second conjunct. Any P1 that holds is therefore reported with this deviation attached, and the
verifier records exactly how much Russian reached the stage-2 seats.
10. Budget
Pre-flight in config/budget.md, built from max_tokens per note (abc) with a routing margin
per the S022 caution and note (bhq). Fifteen dispatches: 1 critic, 2 stage 0, 6 stage 1, 6 stage 2.