Repository path: workshop/experiments/E-20260810-source-asymmetry/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260810-source-asymmetry |
| status | frozen |
| created | 2026-08-10 |
| updated | 2026-08-10 |
| senses | voice, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-forced-choice.md, wiki/findings/results/RS-20260809d-forced-choice.md, wiki/base/anchors/A-brown-calaveras-address/A-brown-calaveras-address.md, workshop/translations/brown-calaveras/R06-v1/translation.md, config/models.md, config/budget.md |
E-20260810-source-asymmetry — does the English carry the relation the translators had to write?
Frozen before dispatch. Session S148 · ARM-forced-choice step 2 of 2 · track T4.
1. Question
RS-20260809d established that seven published hands in three languages over 56 years all wrote a
symmetric address grid for the two men of Poe's "The Cask of Amontillado", and that whether the
English carries the asymmetry they declined to write is open: its P1 was withheld when the
registered speaker-halo gate G2 fired at 0.7917 against a bar of 0.75.
The withheld question, restated on new material:
On the English alone, do the subordinate speaker's turns position him below his interlocutor, relative to the interlocutor's turns — when each turn is read in isolation, with no speaker attribution and no surrounding dialogue?
If yes, a symmetric grid flattens something that is in the source. If no, a symmetric grid loses nothing and the obligatory target category is filled by convention with no source warrant.
2. Why this text and not the previous one
ARM-forced-choice step 1 named two defects a successor must fix, and both are fixed by design
rather than by control:
(a) The item-matching defect. RS-20260809d §5a: the four N items of the halo gate differed
in length and diction formality, so the gate could not distinguish a speaker halo from its own item
selection. This design removes the halo channel instead of controlling it: turns are rated one
per call, in isolation, unattributed, with no dialogue context, so there is no speaker for a halo to
attach to. The residual confound — length, since the subordinate speaker talks more — is not
controlled by hand-picked items either; it is handled by a pre-registered length-matched
re-computation (P2, §6) and a reported rank correlation.
(b) The recognition defect. RS-20260809d §5b: 11 of 12 bodies named "The Cask of Amontillado"
through the anonymisation. The material is changed and the recognition test is promoted from a
reported quantity to a pre-dispatch admission gate that can stop the run (G2, §6).
The text is Bret Harte, "Brown of Calaveras" (Overland Monthly, March 1870): a story by an
author who was translated into at least four European languages in his lifetime and is read in none
of them now. It carries, inside one story, the contrast the previous material could not supply —
a relation English leaves grammatically and lexically unmarked (Hamlin ↔ Brown) and a relation
English marks lexically (Hamlin ↔ the hostler: "his fiery patron", "Mr. Hamlin", "Stand aside!").
RS-20260809d §7 names the absence of that second kind as one of its limits.
3. Materials
- Source.
workshop/translations/brown-calaveras/source-en.txt, 71 paragraphs, 3,932 words, Project Gutenberg #6373, public domain (author d. 1902). - Translation limb, frozen at commit
53be89bbefore this design existed.T-brown-calaveras-R06-v1— the whole story into French underR06, 4,461 French words, with a translator's log carrying the registered address grid: mutualtufor Hamlin ↔ Brown,tudown /vousup for Hamlin ↔ the hostler,vousin the drawing room,tuin Mrs Brown's note. The lead never judges its own translation (charter §5) and no arm of this experiment contains any French. - Census hand. Wilhelmina Zyndram-Kościałkowska's Polish translation (Nowelle, Warszawa: S.
Lewental, 1885), stored whole at
wiki/base/anchors/A-brown-calaveras-address/brown-pl-koscialkowska-1885.txt. Public domain (translator d. 1926). The census rests on no API call and is not part of this run.
Anonymisation
Applied by build_items.py from a declared map, to the recognition-gate text and to every rated
item alike: Jack → Ned, Hamlin → Carver, Brown → Slade, Kate → Nell, Sue → May,
Magnolia → Nugget, Scott's Ferry → Vernon Crossing, Wingdam → Hartley, Calaveras →
Tulare. Nothing else is changed; the em-dashed profanity elisions, the dialect spellings and the
punctuation are Harte's.
Items
Admission rule, declared before the items were cut: every stretch of direct speech that is (i) spoken by Hamlin or by Brown, (ii) addressed to the other of them, and (iii) two words or longer. This admits 28 items — 8 Hamlin, 20 Brown — and excludes, by the rule and not by taste: ¶22 (addressed to Mrs Brown), ¶46 "Two out of three" (spoken aloud but to nobody), and the one-word turns ¶48 "Nothing", ¶54 "Yes", ¶60 "Smoke?", ¶62 "Light?". The 8 : 20 imbalance is the phenomenon, not a sampling fault: Hamlin is laconic and Brown is not, and no item is dropped to make the arms even.
Six constructed control items (G1), written by the lead in the same camp diction, three
spoken from clearly above and three from clearly below, matched on length (58 against 63 words in
total). They are interleaved with the real items in the shuffled order and are indistinguishable to
the seats from the rest.
4. Panel and stage structure
Five seats, panel v1, all non-Anthropic (config/models.md): S1 P1, S2 P2, S3 P3, S4
P4, S5 P5. Five is the minimum at which a seat-level sign test can reach 0.05 (1/32 = 0.031);
this is why the seat count is five and not four.
| stage | what | calls |
|---|---|---|
R |
recognition gate, dispatched first — the whole anonymised Hamlin/Brown dialogue, "name the work and the author, or say you do not know" | 5 |
A |
the rating stage — 34 items (28 real + 6 control), one item per call, no dialogue context, no speaker attribution, no other item visible | 170 |
Stage A is dispatched only if stage R passes (§6, G2). One item per call is the note-(blf)
shape: every answer is a small JSON object, so the token cap is never approached and no answer can
be truncated into a lost body.
Scale-direction counterbalance, registered: seats S1, S3, S5 receive the scale as written
below; seats S2, S4 receive it reversed, and their ratings are re-flipped (8 − x) in analysis
before any statistic is computed.
The scale
Considering only this line, and nothing you may guess about who is speaking: how does the speaker place himself in relation to the person he is speaking to?
- 1 — clearly above: commanding, granting, dismissing, laying down what will happen.
- 4 — level: neither above nor below.
- 7 — clearly below: appealing, deferring, submitting, asking to be granted something.
The anchors name the dimension and do not name superiority or inferiority as good or bad
(RS-20260809d's critic amendment A4, carried forward).
5. What is predicted, and by whom
Registered before dispatch. The direction is taken from the frozen translator's log (§1 of
T-brown-calaveras-R06-v1), which set out the warrant for an asymmetric grid — Brown's begging,
his "Fact, sir", his taking of orders — and then wrote a symmetric one anyway.
Registered direction: Brown's turns score HIGHER (more "below") than Hamlin's.
6. Quantities, bars and gates — all registered before dispatch
| id | quantity | registered bar | if it fails |
|---|---|---|---|
G2 |
recognition: seats naming the work ("Brown of Calaveras") | 0 of 5 (tightened by amendment A1) | stage A is not dispatched and P1 is not measured |
G2a |
seats naming the author but not the work | reported as a limit; does not fire | — |
G1 |
mean(constructed-below) − mean(constructed-above), per seat | ≥ 2.50, and in the registered direction on 5 of 5 seats | P1 is WITHHELD — the instrument is blind |
F1 |
answers quoting 3–8 words verbatim from the rated line | ≥ 0.90 | reported as a limit |
P1 |
mean(Brown items) − mean(Hamlin items), averaged over seats | ≥ 1.50, 5 of 5 seats in the registered direction (sign test P = 0.031), and one-sided Mann–Whitney U over the 28 items on seat-mean ratings at P ≤ 0.05 | the null is reported as measured |
P1b |
the same difference, secondary bar (added by amendment A4) | ≥ 0.75, 5 of 5 seats, U-test P ≤ 0.05 | licenses only the weaker statement in §7 |
P2 |
the same difference on the length-matched subset | ≥ 1.00 and in the registered direction | P1 is reported as not separable from length, and no claim about the source is made |
P2b |
the same difference on items of ≤ 10 words only (added by amendment A3) | reported, direction registered | second look at length |
P3 |
Spearman ρ between item word count and seat-mean rating | reported, no bar | context for P2 |
The length-matching algorithm, declared here so it cannot be chosen afterwards. Each of the 8
Hamlin items is paired with the not-yet-used Brown item nearest to it in word count (ties broken by
the lower item index); this yields 8 pairs, and P2 is the mean within-pair difference. The pairing
is computed by analyse.py from word counts alone and does not look at any rating.
No gate is weakened after it fires, and no seat is dropped. RS-20260809d recorded that
dropping one seat would have passed its gate and did not drop it; the same rule holds here.
7. Failure criteria, stated as such
G2fires → the run stops at 5 calls and the session reports that the recognition problem is not soluble by choosing an obscure text either. That is a publishable outcome and not a wasted run.G1fails →P1withheld; the instrument cannot read the dimension on isolated turns.P1misses its bar → the null is the result: on this text, read this way, the English does not carry an asymmetry for a symmetric grid to flatten, andframework/v0.1§7'sEN→row says so.P1passes andP2fails → the difference is not separable from how much each man talks, and the row says that.
8. Cost
Worst case built from max_tokens and not from an expected answer length (note (abc)): stage A
is 170 calls at a 400-token cap and ~450 prompt tokens; stage R is 5 calls at ~4,000 prompt
tokens; one critic pass. Declared ceiling: $1.00. Today's UTC ledger is empty at the time of
freezing.
9. Verification
verify.py recomputes every reported number from the raw bodies in runs/, independently of
analyse.py, including the scale re-flip, the length pairing, the exact sign-test probability and
a closed-form check of the Mann–Whitney P; and runs mutation tests that must be caught.
10. Amendments accepted from the pre-run critic
Pre-run critic mistralai/mistral-medium-3-5, one pass over this design and the built items,
VERDICT: NEEDS-REDESIGN, 8 findings, 4 BLOCKING. Raw at runs/critic.txt, prompt at
runs/critic-prompt.txt, cost $0.0139305. Four findings accepted, four overruled in writing.
Everything below was done before any rating call was dispatched.
A1 — accepted (BLOCKING 1). G2's bar is tightened from ≤ 1 of 5 to 0 of 5 naming the
work, and split: naming the author without the work is reported as a limit and does not fire,
because the halo the gate exists to exclude requires knowing who says what in this story, which
recognising Harte's dialect does not supply. The critic's stated reason (false positives from the
substituted names) is not the reason; the reason is that with a 5-of-5 sign test one seat that
knows the story can carry the primary.
A2 — accepted (BLOCKING 2). The G1 controls are doubled, 6 above and 6 below, written
before dispatch and interleaved by the same declared hash order. Items: 40 (28 real + 12 control);
stage A becomes 200 calls.
A3 — accepted (SERIOUS 3), as an addition rather than a replacement. P2b is added: the
difference restricted to items of ≤ 10 words, the band where both speakers have several items,
reported alongside the pairwise-matched P2. The critic's proposal — match all 28 items — is not
adopted because 20 Brown items cannot be matched to 8 Hamlin items without discarding twelve of
them, which is a larger distortion than the one it repairs.
A4 — accepted (BLOCKING 7). The 1.50 bar is kept as the registered primary, because it is
RS-20260809d's bar and comparability with the withheld predecessor is the point; but a
secondary bar P1b at ≥ 0.75 is registered now, before dispatch, so that a real-but-smaller
asymmetry is not thrown away and the goalpost cannot be moved afterwards. Which bar was met is
reported explicitly.
Overruled, with reasons
SERIOUS 4 — "the scale anchors are asymmetric". They mirror pairwise and the mirror is why they were written that way: commanding ↔ submitting, granting ↔ asking to be granted, dismissing ↔ appealing, laying down what will happen ↔ deferring. Overruled unchanged.
MINOR 5 — "include the one-word turns". The two-word floor was declared before the items were cut. Relaxing it now would add four items — "Nothing", "Yes", "Smoke?", "Light?" — all four of them Hamlin's, and all four peremptory, i.e. a change that can only move the result toward the registered direction. Overruled for that reason, which is recorded so it can be checked.
MINOR 6 — "counterbalance only 2 of 5". The design already splits 3 forward / 2 reversed (§4). Misreading; overruled.
SERIOUS 8 — "F1 is a red herring". F1 is not a comprehension check. It is a compliance
check that the answer came back against the item that was sent, which is the failure this
one-item-per-call shape is otherwise blind to. Overruled.
Consequence for §8
The ceiling is raised from $1.00 to $1.20 to cover the 30 extra calls A2 adds. Today's UTC ledger is empty; the raise is written here before dispatch, not after the outturn.