Repository path: workshop/experiments/E-20260809d-forced-choice/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260809d-forced-choice |
| status | frozen |
| created | 2026-08-09 |
| updated | 2026-08-09 |
| senses | style-correspondence, accuracy, voice |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-forced-choice.md, wiki/base/anchors/A-cask-forced-choice/A-cask-forced-choice.md, workshop/translations/cask-amontillado/R04-v1/translation.md, framework/v0.1/README.md, framework/v0.2/README.md, config/models.md, config/budget.md |
E-20260809d — is the asymmetry the seven hands flattened there in the English?
Frozen before dispatch. Nothing below is amended after any body returns except by a numbered amendment recorded in this file with its reason and its timestamp relative to dispatch.
1. Where this comes from
framework/v0.1 §2 declares R1 untested with English as source — the project's standing gap
since S015, and framework/v0.2 §2 states S1: across five language pairs no population of
sites was found at which a competent English rendering loses a relation the source marks
grammatically. Every one of the five had English as the target.
Turning the direction round does not turn R1 round, because the asymmetry between English and its usual partners is not symmetrical: English has fewer obligatory relational categories, not more. What the reversed direction produces is the mirror problem — the target's grammar demands a relation the source never states, and the translator must supply it.
The census is already done and is reported in the anchor (A-cask-forced-choice §4): on Poe's
"The Cask of Amontillado" (1846), seven published hands in three target languages over 56 years
each had to choose a second-person address grid, and 7 of 7 chose a symmetric one — six mutual
V (Baudelaire 1857; four Russian hands 1881–1906), one mutual T (Leśmian 1913). Not one hand
distinguished the two speakers. The lead's own rendering, frozen before any of them was opened,
chose an asymmetric grid.
So the question this run exists to answer is the one the census cannot: is there an asymmetry of standing in Poe's English for the hands to have flattened, or is the lead's reading an invention?
2. Question
Do independent readers of the English alone read Montresor's turns as more ceremonious toward his addressee than Fortunato's turns are toward his?
3. Materials
- Source:
wiki/base/anchors/A-cask-forced-choice/— Poe 1846, stored whole, PD. - Context block: the story's dialogue in story order, turns labelled
AandBand numbered, narration reduced to one-line stage directions. Built mechanically bybuild_items.py; the output is committed asmaterials/items.jsonbefore dispatch. - Anonymisation. Every name a reader could use to identify the story is replaced, by a fixed
table applied mechanically: Fortunato→Guarnieri, Montresor(s)→Belfiore,
Luchesi→Aldani, Amontillado→Madeira, "Lady Fortunato"→"Signora Guarnieri". Reason: the
canonical reading of this story is that Montresor is ingratiating, so a seat that recognises it
is biased toward the prediction. Recognition is probed anyway (§6
R) and reported whatever it says. This modifies the text and the modification is part of the design, not a defect discovered later. - Speaker mapping is fixed and is not disclosed to the seats:
A= Montresor (the narrator),B= Fortunato.
4. Items
| block | n | what |
|---|---|---|
M |
10 | Montresor's turns to Fortunato: source ¶5, 7, 13, 19, 21, 31, 35, 37, 52, 72 |
F |
10 | Fortunato's turns to Montresor: source ¶6, 14, 16+18, 20, 22, 36, 46, 53, 56+58, 66 |
N |
4 | neutral turns from the same dialogue with no address content: ¶44 (B), ¶45 (A), ¶47 (A), ¶48 (B) |
C |
4 | constructed control lines, written by the lead, presented as being from another work: two unambiguously from a superior to an inferior, two unambiguously from an inferior to a superior |
28 rated items per call.
5. Instrument
Each seat receives the whole anonymised dialogue as context — the trajectory is shown, not
withheld, because framework/v0.2 §6 records that a per-utterance instrument cannot see a
relation that lives in a text's arc, and address level is exactly such a relation.
Rating question, block 1 verbatim:
For each numbered line, how ceremonious is the speaker toward the person he is speaking to? 1 = blunt or familiar, speaks as a clear superior or an intimate. 4 = speaks as an equal. 7 = highly ceremonious, deferential, elaborately polite. Rate only the manner of address, not whether the speaker is sincere and not whether you like him.
Required with every rating: a quotation of at most eight words from that line that decided it. (Note (bky), third firing: a probe whose raters cannot quote the thing being measured is not measuring it. Quotes are analysed and reported whatever they show.)
Block 2 is the same context with the scale reversed — 1 = highly ceremonious … 7 = blunt or
familiar — and its ratings are inverted (8 − x) before analysis. This is the run's
counterbalance: the dialogue cannot be reordered without destroying the trajectory the design
insists on showing, so the scale is what is counterbalanced, which is also the bias most likely
to be present.
Seats. Panel v1 in full plus the declared reserve: P1 openai/gpt-5.6-terra, P2
google/gemini-3.6-flash, P3 x-ai/grok-4.5, P4 moonshotai/kimi-k3, P5
deepseek/deepseek-v4-pro, P6 qwen/qwen3.7-max. No Anthropic model (charter §2.2,
config/models.md). 6 seats × 2 blocks = 12 calls, 12 ratings per item.
No seat is told anything about translation, about pronouns, about address forms, about French, Russian or Polish, or about the hypothesis.
6. Registered quantities, bars and failure criteria
All of these are fixed before dispatch. x̄M and x̄F are means over the M and F blocks after
block-2 inversion.
| id | quantity | bar | if it fails |
|---|---|---|---|
P1 (primary) |
x̄M − x̄F |
≥ 1.50 scale points and the sign is the same for 6 of 6 seats (one-sided binomial P = 0.0156) | primary not licensed; report the figure and say so |
G1 (positive control, gate) |
mean(C inferior→superior) − mean(C superior→inferior) |
≥ 2.50 and 6 of 6 seats | P1 is WITHHELD. An instrument that cannot read ceremony off constructed English cannot be believed about Poe's |
G2 (speaker-halo control, gate) |
|mean(N A-turns) − mean(N B-turns)| |
≤ 0.75 | P1 is WITHHELD. A gap on turns with no address content means the seats are rating the speaker, not the words |
F1 (failure criterion) |
share of M and F items whose quote comes from the rated line | ≥ 0.90 | P1 demoted to descriptive |
R (recognition probe, reported not gating) |
seats naming the source work when asked at the end of block 2 | — | reported as a limit whatever it is |
P2 (secondary, no API) |
published hands whose grid is symmetric | descriptive — measured at 7 of 7 before this run was designed | — |
G1 and G2 are gates and are declared before dispatch. They will not be weakened after they
fire (RS-20260808e §"the gate was not weakened after it fired").
What each outcome means, written now so it cannot be chosen later:
P1holds → Poe's English carries an asymmetry of standing that all seven published hands deleted when their grammar made them choose. That is a loss with English as source, and it is not R1's loss: nothing English marks grammatically was lost. It is the mirror — an obligatory target category filled by convention, with the source's non-grammatical signals of the same relation not reaching it.P1fails → the two men's manner is not differentially ceremonious to readers of the English, the seven hands' symmetric grids are faithful, and the lead's asymmetric grid is an invention. The translator's log's §2 warrant would then be a reading the text does not support, which is a result about the lead and is reported as one.
7. Procedure
build_items.pywritesmaterials/items.json(context block, item list, speaker map, anonymisation table). Committed before dispatch.- Independent pre-run critic, one non-panel seat, given this file and the materials. Findings applied as numbered amendments in this file with reasons, or overruled in writing with a demonstration.
- Dispatch block 1 to all six seats, then block 2. Raw bodies preserved under
runs/. analyse.pycomputes every registered quantity.verify.pyrecomputes all of them from the raw bodies independently and checks the item counts, the inversion, and every number that reaches the result page.- Cost recorded from
usage.costper response and reconciled against the key-usage delta.
Operational notes applied from the start: (bkw) — "reasoning": {"enabled": false} on
every seat that accepts it, {"effort": "low"} where that form 400s, and nothing sent to
gemini, which 400s on both; (bhf) — zero content at cost means change the seat, not raise the
cap, after one reasoning-off retry.
8. Pre-flight cost estimate
Worst case is built from max_tokens, not from expected output (note (abc)).
max_tokens = 6,000. Prompt ≈ 2,000 tokens. 12 calls = 2 per seat.
| seat | in $/M | out $/M | 2 calls worst case |
|---|---|---|---|
| P1 gpt-5.6-terra | 1.00 | 6.00 | 0.0760 |
| P2 gemini-3.6-flash | 1.50 | 7.50 | 0.0960 |
| P3 grok-4.5 | 2.00 | 6.00 | 0.0800 |
| P4 kimi-k3 | 3.00 | 15.00 | 0.1920 |
| P5 deepseek-v4-pro | 0.435 | 0.87 | 0.0122 |
| P6 qwen3.7-max | 1.475 | 4.425 | 0.0590 |
Run worst case $0.5152. Pre-run critic, max_tokens 16,000, one non-panel seat: $0.12.
Re-dispatch contingency: $0.20.
Declared ceiling for this experiment: $0.90. UTC day 2026-08-09 stands at $1.749152 of $5.00 before this session, so the declared ceiling fits inside the remaining $3.250848 with $2.35 to spare.
9. Amendments
Pre-run critic: mistralai/mistral-medium-3-5, one call, max_tokens 16,000, $0.0409395,
10,775 characters, VERDICT: NEEDS-REDESIGN, 2 BLOCKING / 4 SERIOUS / 3 MINOR. Body at
runs/critic.txt, prompt at runs/critic-prompt.txt. Seven amendments accepted, five findings
overruled in writing. All of it before dispatch; not one body had been requested.
Accepted
A1(from BLOCKING 2's third limb —8 − xtreats an ordinal scale as interval). The inversion stays, because the primary is a mean difference and already assumes that much. What is added is an ordinal check that does not:P1a(forward block alone) andP1b(reversed block alone, computed on raw scores so that the predicted sign flips) must agree in direction. Registered as a condition onP1.A2(from SERIOUS 3 — the 6-of-6 binomial assumes seats are independent, and they are not). Accepted, and it is a real hole: these six models share training data. The registered significance test becomes an item-level Mann–Whitney U over the 20 items (10 M against 10 F, using each item's mean across all 12 ratings), which is a statement about items and does not need the seats to be independent. Exact U, two-sided. The 6-of-6 seat sign is kept as a robustness requirement, not as the P-value. Bar: U-test P ≤ 0.05 and the effect-size bar below.A3(SERIOUS 4, second limb). The four constructed control lines are re-worded to nineteenth-century diction. The asymmetry is not softened: a control exists to show the scale discriminates ceremony at all.A4(SERIOUS 5). The scale anchors no longer say "speaks as a clear superior" or "deferential". They now read "blunt or familiar; stands on no ceremony at all" against "highly ceremonious; elaborately polite, careful of the other man's dignity". The dimension is still named — an instrument that does not name its dimension measures nothing — but the power reading is no longer supplied.A5(SERIOUS 6, first and third limbs). The prompt now says the quote must be copied verbatim, word for word, from that line.F1's matching rule is defined here rather than left toanalyse.py: a quote counts as grounded when its case-folded, whitespace-collapsed, punctuation-stripped form is a substring of the rated line's same-normalised text.A6(MINOR 7). Anonymisation extends topalazzo→house,Medoc→claret,Signora→Madame.A7(MINOR 8). No item spans two turns.F3becomes line 14 alone andF9line 42 alone, so every rated item is exactly one speaker turn. Item counts are unchanged at 10/10/4/4.
Overruled, with the reason
- BLOCKING 1 — "
G1's 2.50 bar is arbitrarily high and the gate is circular." The circularity claim is wrong: a positive control that fails means the instrument was not shown to be sensitive, and withholding is the conservative response, not a loop. The bar is not arbitrary either —C1andC3are constructed to sit near 1 andC2andC4near 7, so the expected difference is about 4.5 and 2.50 is comfortably below it. Lowering a positive control's bar because it might fail is the moveRS-20260808eis on record refusing. - BLOCKING 2, first and second limbs — "a seat will notice the reversal and mechanically invert; use between-subjects." Demonstrably not available to it: every call in this run is a separate stateless HTTP request with no conversation history, so a seat cannot know it has seen the other block. Within-seat counterbalancing here is between-subjects counterbalancing plus a second rating per item.
- SERIOUS 3's "lower the bar to 0.75" — refused. The effect-size bar was registered before any
body was requested and is not moved because a critic guessed the effect might be small.
A2supplies the significance test the finding was actually reaching for. If the difference is significant but under 1.50, that is reported as significant and below the registered bar, andP1is not licensed as registered. - SERIOUS 4, first limb — "
N2/N3are not addresses, replace them." They are not addresses by construction:G2asks whether the seats rate the speaker rather than the words, and turns with no address content are exactly what that question needs. - MINOR 9 — "normalise for per-model rating bias." Already removed: every quantity is a
within-seat difference (
x̄M − x̄Fcomputed inside each seat before pooling), and an additive per-model offset cancels in a difference.
The registered table as amended
| id | quantity | bar |
|---|---|---|
P1 |
x̄M − x̄F, within-seat then pooled |
≥ 1.50 and item-level Mann–Whitney U P ≤ 0.05 and 6 of 6 seats same sign and P1a/P1b agree in direction |
G1 |
mean(C inf) − mean(C sup) |
≥ 2.50 and 6 of 6 seats — gate |
G2 |
|mean(N A) − mean(N B)| |
≤ 0.75 — gate |
F1 |
share of M/F quotes grounded in the rated line by A5's rule |
≥ 0.90, else P1 descriptive |
R |
recognition | reported, not gating |