Repository path: workshop/experiments/E-20260804b-quixote-affect/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260804b-quixote-affect |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| senses | affect, style-correspondence, voice, naturalness |
| internal-judgment-only | true |
| provisional | true |
| track | T2 |
| links | wiki/arms/ARM-affect-reception.md, wiki/base/sources/S-quixote-humour-reception.md, workshop/translations/quijote-I3/R04-v1/translation.md, wiki/goodness-senses.md, config/models.md, config/budget.md, workshop/experiments/E-20260804b-quixote-affect/critic.md |
E-20260804b — does the operation the reception record blames actually do what it is blamed for?
FROZEN, v2, after the pre-run critic pass and before any grading call. Materials
materials/sites.json, SHA-256
717f005b4d691286492c4e31752e1c2b5a51e4d9da8e3863b59a3d4220452369, built by materials/extract.py
and materials/sites.py, both committed. v1 (hash da2d476e…) is the version the critic saw; all
ten of its findings were accepted and are folded in below as A1–A10, recorded in critic.md.
No number below was computed after seeing any grader output.
1. Question
Ormsby's 1885 introduction (S-quixote-humour-reception) makes a claim with two separable halves:
- descriptive — Motteux's English adds narratorial signalling that Cervantes's Spanish does not have ("a flippant, would-be facetious style"; the essence of the humour is "the grave matter-of-factness of the narrative, and the apparent unconsciousness of the author that he is saying anything ludicrous");
- evaluative — that addition is "an absolute falsification of the spirit of the book".
Lockhart, editing Motteux in 1822, grants the first half and rejects the second: Shelton is
"perhaps even more successful in imitating the 'serious air' of Cervantes", and Motteux is
preferred anyway because "in spite of many defects and inaccuracies, it is by far the most
spirited." affect is defined on wiki/goodness-senses.md as "produces in its reader an experience
comparable to what the source produces in its reader." This design asks whether the two halves of
that definition — produces an experience and comparable to the source's — come apart on exactly
the operation the record argues about.
A2, stated as a limit on the inference before the run: a divergence between which arm is picked
funniest and which is picked closest to the Spanish shows that two questions are answered
differently. It does not by itself show that narratorial signalling causes the divergence. What
bears on causation is the matched lead / nudge / nudge-flat triple, where nothing varies but the
operator.
2. What this is not
No jury quality verdict is taken, and none would be licensed. Tier D is NOT PASSED
(config/models.md). SIGNAL is descriptive coding of a textual property against the Spanish,
which S015's instrument note licenses; MATCH, FUNNIEST-SP and FUNNIEST are uncalibrated
reader reports, reported as such and never as evidence that one translation is better than another.
Every claim on the result page carries internal-judgment-only and provisional.
3. Materials
Seven loci from Don Quijote I.3, seven English arms each (materials/sites.json):
| arm | provenance |
|---|---|
shelton |
Thomas Shelton, 1612 |
motteux |
Peter Motteux, 1712 — the arm the record accuses |
smollett |
Tobias Smollett, 1755 |
ormsby |
John Ormsby, 1885 — the arm the record's author wrote |
lead |
the lead agent, 2026 (T-quijote-I3-R04-v1), a labeled subject |
nudge |
POSITIVE CONTROL — lead plus an inserted narratorial signal that is also comic |
nudge-flat |
A6, the discriminating control — lead plus an inserted narratorial signal of matched length that is evaluative and intrusive but not comic |
nudge and nudge-flat differ from lead in the insertion and in nothing else, and from each other
only in whether the insertion is funny. That triple is what separates adds a signal from adds a
joke — the confound finding 6 named.
The seven loci — S1-close, S2-vigil, S3-innkeeper-cv, S4-arithmetic, S5-no-heed,
S6-ledger, S7-held-laughter — cover the record's three named mechanisms (unsmiling gravity,
sententious terseness, prosaic incongruity). That selection is lead-authored and the lead had read
the record first; it is a declared limit, not a controlled variable (§7 L1).
The graders are given the Spanish and no gloss. A lead-authored gloss would reinstate exactly the
defect RS-20260804-yardstick found one session ago.
4. Arms, stages, seats
Seats from config/models.md, all non-Anthropic. No seat grades its own output; no seat wrote any
arm; no Stage-0 seat takes part in Stages 1–2, and the critic seat (P4) takes part in neither.
- Stage 0 — the independent mechanism statement. TWO SEATS (A5).
P5(deepseek/deepseek-v4-pro) and theconfig/models.mdfirst reserveqwen/qwen3.7-max, each shown the Spanish at all seven loci and no English at all. Per locus: is it comic; in one sentence what produces the comic effect; and does the Spanish narrator himself signal that something funny is being said (evaluative epithet, aside to the reader, exclamation) — yes/no with the Spanish words quoted. 2 calls. - Stage 1 —
SIGNAL,MATCH,FUNNIEST-SP; source present.P1openai/gpt-5.6-terra,P2google/gemini-3.6-flash,P3x-ai/grok-4.5. Per locus: the Spanish, then the seven arms unlabelled under stable letter keys in a randomised order. Then, in this order: 1. per arm,SIGNALyes/no — does this English contain an evaluative or comic signal from the narrator (an aside to the reader, an evaluative epithet, a judgment on the characters, an exclamation) that the Spanish does not contain at this point? — with the exact English words quoted when yes; 2. oneMATCHpick — which single arm comes closest to producing the effect the Spanish produces; 3.FUNNIEST-SP(A2) — one pick, which single arm is funniest, asked in the same call on the same key order, so that the crux contrast varies task and nothing else. - Stage 2 —
FUNNIEST, source absent. Same three seats, new call, no Spanish, arms re-randomised under new keys. OneFUNNIESTpick per locus. Then, and only at the very end (A3), a recognition probe: for each key, name the translator if recognised, elseUNKNOWN. It is placed here, and nowhere in Stage 1, so that no authorship question is anywhere in the prompt whileSIGNALandMATCHare being answered.
Two orderings (o0, o1) per seat per stage, key maps drawn from sha256(stage|ordering|site)
and stored in runs/keymaps.json. 2 + 6 + 6 = 14 grading calls, plus the critic call already
spent.
5. Pre-registered predictions
| # | prediction | bar |
|---|---|---|
| PR1 | positive control: SIGNAL(nudge) − SIGNAL(lead) |
≥ 0.40 |
| PR2 | the record's descriptive half: SIGNAL(motteux) − mean SIGNAL(shelton, smollett, ormsby) |
≥ 0.15, reported with and without S3 (A10) |
| PR3 | SIGNAL(ormsby) is the lowest of the four published arms |
rank 4 of 4 |
| PR4 | nudge takes more FUNNIEST picks than lead |
strict > |
| PR4a | A6: SIGNAL(nudge-flat) − SIGNAL(lead) — the operator is detected when it is not funny |
≥ 0.40 |
| PR4b | A6: nudge takes more FUNNIEST picks than nudge-flat — it is the joke, not the signal, that buys the laugh |
strict > |
| PR5 | A1, the crux, directional: modal FUNNIEST-SP ∈ {motteux, nudge} and modal MATCH ∈ {shelton, smollett, ormsby, lead} |
both hold |
| PR6 | A4, the motivating reception fact: motteux FUNNIEST picks > ormsby FUNNIEST picks |
strict > |
SIGNAL(arm) is the proportion of yes over 7 sites × 3 seats × 2 orderings = 42 cells per arm
(fewer where F5 excludes a site). The uniform-null probability of a bare modal split between two
independent 42-pick draws over 7 arms is computed and reported beside PR5, so that the directional
bar is read against the noise level rather than instead of it (A1).
6. Failure criteria — pre-registered, each with its consequence
- F1 — the instrument. PR1 and PR4a both miss ⇒ the coding does not detect the accused
operation on pairs that differ in nothing else, and PR2 and PR3 are withheld; all
SIGNALfigures become descriptive only. - F2 — quote verification. Every
SIGNAL=yesmust quote a string occurring verbatim in the arm it is about (whitespace- and case-normalised). If < 0.80 of yes-quotes verify,SIGNALcoding is unreliable ⇒ descriptive only. Computed mechanically byanalysis/verify.py. - F3 — order (A9). Fires only where the order effect reverses the sign of the pick difference
a prediction is about (
nudge−leadfor PR4/PR4b,motteux−ormsbyfor PR6). Otherwise both orderings are reported and pooled. - F4 — seat agreement on the forced choices. If fewer than 5 of 7 sites take a 2-of-3 seat
majority on
MATCH,MATCHis descriptive only. Same rule applied separately toFUNNIEST-SPandFUNNIEST. - F5 — ill-posed sites (A5). A locus is excluded from PR2 and PR3 if either Stage-0 seat
answers that the Spanish itself signals the joke — excluded by default on disagreement. Exclusions
are reported with both seats' quotes.
S5-no-heedis the site expected to go, since the Spanish there carries(y fuera mejor que se curara, porque fuera curarse en salud). - F6 — length confound (A7). Pair-level: per site, the
nudge−leadpick difference against thenudge−leadword difference. If the sign of the pick difference tracks the word difference at 6 of 7 sites or more, PR4 is reported as confounded.nudge-flatis length-matched tonudgeby construction, so PR4b is structurally protected. - F8 — seat agreement on
SIGNAL(A8). Majority attainment (2 of 3) computed over every (site, arm, ordering) cell. Below 0.80,SIGNALis descriptive only.
7. Declared limits, written before the run
- L1 — site selection is lead-authored and record-primed. The lead read Ormsby's introduction before choosing the seven loci, so the loci may favour the mechanisms Ormsby names. Stage 0 fixes the mechanism statement, not the site list.
- L2 —
lead,nudgeandnudge-flatare all lead-authored. The PR1/PR4a/PR4b triple is a clean matched manipulation but sits inside one translator's prose. It generalises to the operation, not to Motteux. - L3 — the arms are not equal-length and were never made so.
motteuxatS3-innkeeper-cvis 21 words against 42–63 because he omits the innkeeper's list of crimes. Omission is a different operation from addition, is never folded intoSIGNAL, and is why PR2 is reported twice (A10). - L4 — contamination.
lead~ormsbyis the highest raw pair in the dependence matrix (91 shared 7-grams, longest run 15 tokens); name-excluded it is 6 shared 7-grams and zero 12-grams, as is every other pair.leadis therefore not used as an independent comparator anywhere; it is a labeled subject and the base of a matched control. - L5 — one reception record is one record. Nothing here establishes that Ormsby's account of Cervantes's humour is right about Cervantes. It tests whether the account describes the texts.
- L6 — recognition (A3). All four published arms are canonical public-domain texts, and the
criticism of them — including Ormsby's verdict on Motteux — is equally public. Letter keys do not
prevent a grader from recognising the texts and reciting the received view of them. No amendment
removes this. Stage 2's probe measures it; whatever it returns, no
SIGNAL,MATCHorFUNNIESTfigure on a published arm can be cleanly separated from the seats' prior knowledge of that arm's reputation. Thelead/nudge/nudge-flattriple is the part of this design that is immune, being unpublished and unseen, which is a further reason the causal weight sits there.
8. Analysis
analysis/score.py computes every figure; analysis/verify.py recomputes each one by an independent
path from the stored raw bodies, runs F2's quote check, asserts the key maps invert, and re-derives
the uniform-null modal-split probability by exhaustive enumeration. Mutation tests: three deliberate
corruptions of the stored scores, each of which the verifier must catch.
9. Cost pre-flight
Worst case built from max_tokens, per note (abc) — never from an assumed output length.
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic (P4) | 1 | 12,000 | $0.07368 actual |
| 0 (P5, qwen reserve) | 2 | 3,000 | $0.06 |
| 1 (P1,P2,P3 × 2) | 6 | 10,000 | $0.44 |
| 2 (P1,P2,P3 × 2) | 6 | 8,000 | $0.35 |
| re-dispatch reserve | — | — | $0.20 |
| declared worst case, total | 15 | $1.13 |
Declared as $1.25 with rounding. Today's UTC headroom before Stage 0: $4.674142391. Note
(b): reasoning: {effort: low} on the first dispatch to every seat. Note (bhf): P2 gets
its full token allowance from the outset.
10. Pre-run critic
Done. One pass, NEEDS-AMENDMENT, ten findings, all accepted, applied above as A1–A10 before any
grading call. Record: critic.md.