Repository path: workshop/experiments/E-20260729c-neutral-summary/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260729c-neutral-summary |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | accuracy, cultural-mediation, literary-quality, naturalness, style-correspondence, voice |
| internal-judgment-only | true |
| links | wiki/arms/ARM-tierD-repair.md, wiki/decisions/resolved/D-20260725-07-athenaeum-1906-condition-ii.md, wiki/decisions/votes/2026-07-25/D-20260725-07-ratification-record.md, wiki/decisions/resolved/D-20260725-06-heldout-arm-operationalisation.md, wiki/backlog.md, config/models.md, workshop/translations/dva-generala/R04-v1/translation.md |
E-20260729c — condition (iii): re-derive the primary, and re-run the vote on a summary the lead did not write
Frozen before any call is dispatched and before the translation limb's prose is written.
ARM-tierD-repair step 4. This design has two parts and they interrogate the same
1904 document from two directions.
Part A — the study limb
Question
D-20260725-07 ratified option C, narrowed: condition (ii) of D-20260725-06 is
satisfied for Garnett/Hapgood on the Memoirs of a Sportsman cycle only and for the
senses accuracy and cultural-mediation only. Its ratification record names one
un-taken check in terms:
"a future session that wants to re-test this decision should re-derive the primary itself and re-run the vote on a neutral summary. That is a real, cheap, and un-taken check."
wiki/backlog.md's merged owed entry makes that check condition (iii) of the three
that any future Tier D run must meet. This experiment discharges it, or reports why it
cannot be discharged.
The question, stated so it can fail: does the ratified outcome survive (a) an independent re-derivation of the primary from the scan, and (b) a re-run of the ratification protocol in which the evidence the voices read was not written by the lead agent?
Why this is not merely a re-run
The confound named on the ratification record is specific: S025 wrote the evidence summary the voices read, and both voices found its framing tilted toward the answer S025 preferred. A re-run that reused S025's excerpt files would inherit the confound at one remove, because the selection of which sentences to excerpt is also S025's. This design therefore does not use the excerpt files at all. It regenerates the whole scan page from the item id and gives a non-lead model the complete review, unedited, OCR errors intact.
Materials — frozen, and what makes each one admissible
| file | what it is | who made it |
|---|---|---|
materials/nation-1904-p93-rederived.txt |
raw output of tools/deinterleave_djvu.py on scan page 16 |
the tool, from the item id |
materials/nation-1904-p94-rederived.txt |
same, scan page 17 | the tool, from the item id |
materials/athenaeum-1906-body.txt |
the 1906 OCR body, lead header stripped | OCR; no lead prose |
materials/shared-question-and-options.md |
condition (ii) verbatim, the question, options A–E verbatim, contingent-artifact note | the lead (2026-07-25), identical in both arms |
materials/arm-L-evidence-lead-written.md |
the lead's own evidence sections, verbatim | the lead (2026-07-25) |
materials/arm-N-evidence-nonlead.md |
to be produced by stage 1 | a non-lead model |
The options are lead-written and stay lead-written. They are the decision's options; changing them would change the decision rather than test it. They are byte-identical across arms, so nothing in the contrast is attributable to them.
Procedure
Stage 0 — re-derivation of the primary (no API call). Fetch _djvu.xml for
sim_nation_1904-02-04_78_2014, run tools/deinterleave_djvu.py on scan pages 16 and 17,
and check every sentence D-20260725-07 and its excerpt file quote as load-bearing against
the regenerated text. Report each as REPRODUCED / VARIANT / NOT FOUND. (Executed before
this design was frozen; the outputs are the stage-1 materials above. Recording it here is
the record of what was done, not a prediction about it.)
Stage 1 — the non-lead summary. One call. Model P3 x-ai/grok-4.5 — chosen because
it took no part in D-20260725-07 and is neither of the two voices being re-run. Input: the
three primary files and nothing else, plus the text of condition (ii). Task: write a
neutral factual summary of what these two documents say bearing on condition (ii). The
model is not shown the options, is not told a decision exists, and is not asked for
a recommendation. Its output becomes materials/arm-N-evidence-nonlead.md verbatim, with
no lead editing of any kind — including if it is wrong.
Stage 2 — the protocol, run twice. continue-prompt.md §3's protocol is an independent
adversarial review followed by a routed ratifying vote, and the vote governs. Both roles run
in both arms:
| call | role | model | arm | evidence block |
|---|---|---|---|---|
| 2a | adversarial review | P1 openai/gpt-5.6-terra |
N | arm-N-evidence-nonlead.md |
| 2b | ratifying vote | P2 google/gemini-3.6-flash |
N | arm-N-evidence-nonlead.md |
| 2c | adversarial review | P1 | L | arm-L-evidence-lead-written.md |
| 2d | ratifying vote | P2 | L | arm-L-evidence-lead-written.md |
Same two models as 2026-07-25, deliberately: this is a reproduction check, and changing the instrument would confound the answer with a model change. Each call receives the shared block plus exactly one evidence block. No call is told the outcome of the 2026-07-25 ratification, that a ratification has ever happened, or that another arm exists. Calls are independent; the review's output is not shown to the vote, which departs from the 2026-07-25 protocol and is declared as a departure — showing it would leak arm identity through the review's wording.
Stage 3 — analysis. Option letter, sense scope and work scope extracted from each of the
four verdicts by the lead, with the verbatim sentence each is extracted from quoted on the
result page. analysis/verify.py recomputes every reported count from the raw bodies.
Pre-registered predictions
- A1. Arm N's vote (2b) returns C.
- A2. No call in either arm returns B, D or E.
- A3. Arm N's vote scopes to the Memoirs cycle and to
accuracy+cultural-mediation. - A4. Arms N and L return the same option letter in the vote role.
- A5. At least one Arm N voice notes that every paired extract the review prints is cited to Hapgood vol. iv / Garnett vol. ii, i.e. to A Nobleman's Nest, and that the review prints no style exhibit from the Memoirs of a Sportsman. This is visible only to a reader of the whole page; the lead's excerpt file records the citations in an inventory section and never draws the consequence. A5 is the prediction this design most expects to fail, and it is registered because it is the one thing the re-derivation bought that a re-run of the excerpt could not.
Failure criteria, declared in advance
- If A1 fails — if Arm N's vote returns anything but C — then condition (iii) is
discharged with the finding that the ratified outcome does not reproduce on a summary
the lead did not write.
D-20260725-07's binding clauses are then in doubt, and the consequence is a new decision page, which this session opens and may not ratify (charter §8). It is not this session's job to decide it. - If A1 holds and A4 fails, the difference is reported as a difference and is not read as a measured framing effect: see the power statement below.
- A null is a result. If all four verdicts agree with 2026-07-25, condition (iii) is discharged as met and the arm may close.
Power, stated before the numbers exist
Four verdicts across two arms cannot measure a framing effect. With one call per
(role × arm) cell there is no within-cell variance estimate, and this project established
at S053 that a temperature: 0 repeat measures backend determinism rather than judgment
variance, so the cheap repair is unavailable. The N-vs-L contrast is therefore
descriptive and is reported as a description. The load-bearing claim of this experiment is
A1 alone: the ratified outcome does or does not reproduce when the evidence is not
lead-written. That is a one-cell question and one cell answers it.
And Arm L is not a replica of the 2026-07-25 prompt. It omits the lead's recorded preference and the "convergence" section, so that the two arms differ only in the authorship of the evidence block. The comparison N-vs-2026-07-25 therefore differs in more than one thing and is not used as a control.
Part B — the translation limb, and the wire
The wire, in one sentence
The 1904 review is the sole external ground for D-20260725-07's sense-narrowing, and it
makes exactly one falsifiable claim about translation — that a Russian second-person
address-slur "could be reproduced in English only by a stage-direction"; Part A tests whether
that review's verdict survives being summarised by someone other than the lead, and Part B
tests the review's own claim by translating a tale whose entire satire is carried by a
three-way Russian address system.
The claim, verbatim from the re-derived primary
"In Russian the singular pronoun helps to express an old lady's contempt for the town gossip. The slur could be reproduced in English only by a stage-direction."
The reviewer's own illustration shows the two options he can see: Hapgood's "Thou sayest that, my good sir, because thou hast never been married thyself" against Garnett's "You say that, my good sir, because you have never been married yourself" — an archaic pronoun, or silence. The registered question is whether the choice is that binary.
Material, and why this one
М. Е. Салтыков-Щедрин, «Повесть о том, как один мужик двух генералов прокормил» (1869), 2,059 words, public domain, freely reachable whole on ru.wikisource; a published English (Thomas Seltzer, Best Russian Short Stories, 1917) is freely reachable whole on Project Gutenberg, so the contamination gate can actually run.
The tale carries a three-way address system that English has no pronoun for:
- the two generals address each other with
вы+ «ваше превосходительство»; - they address the muzhik with
ты+ «лежебок», «дружок», «каналья»; - the muzhik addresses them with
вы+ «господа генералы».
That is a strictly harder case than the reviewer's, which has one asymmetry; and it is the same phenomenon, in the same language, criticised by the same critic in the same year the project's evidence base is built on.
Procedure, and the order is the point
- Freeze the census before translating. Every site in the source where a second-person
address form or a status vocative occurs is enumerated in
census.json, with its direction (G→G,G→M,M→G), before any English is written. Committed first. - Translate unit A (the opening, through the generals' first exchange) and run the
contamination gate against Seltzer before the rest is drafted — the order
CLAUDE.md's standing rule prescribes and the order S049/S050 established. - Translate the whole tale from the Russian alone. The Seltzer file is on disk for the gate and is not read; the one exception already incurred is declared in the log.
- Classify each census site by the device actually used —
vocative,register,syntax,lexis,stage-direction,lost— after the translation is frozen.
Pre-registered predictions
- B1. The finished translation uses a stage-direction — an authorial gloss naming the pronoun or the address form, of the kind the reviewer says is the only route — at zero of the census sites.
- B2. At least three distinct non-stage-direction devices are used across the
G→Msites. - B3. At every exchange in which both directions occur, the asymmetry — that one party addresses the other more familiarly than he is addressed in return — is recoverable in the English by a reader with no access to the Russian.
- B4. The contamination gate returns a longest common run with Seltzer at or below 12 tokens on the whole work. (Registered as a gate, not as a hoped-for result: above 12 the translation is declared contaminated and Part B's device counts are reported with that declaration attached.)
What Part B does and does not license
It licenses a descriptive sentence: a translator working from the Russian alone rendered
N address sites using D distinct devices and M stage-directions. B3's "recoverable" is the
only evaluative element and is the lead's own reading of its own prose, so it carries
internal-judgment-only and is not offered as evidence about readers. The lead does not
judge the quality of its own translation (charter §5); no quality claim is made here and
none is licensed.
What would refute the reviewer, and what would not. A single translation showing devices between "thou" and silence refutes "only by a stage-direction" as a claim about possibility. It does not show that those devices are good, that a jury would prefer them, or that Hapgood should have used them. The narrower claim is the one this design makes.
Cost, pre-flight
Worst case built from max_tokens (note (abc)), at list out-price plus the prompt at list
in-price:
| call | model | max_tokens | worst case |
|---|---|---|---|
| pre-run critic | P4 moonshotai/kimi-k3 |
16,000 | $0.29 |
| stage 1 summary | P3 x-ai/grok-4.5 |
8,000 | $0.07 |
| 2a, 2c review | P1 openai/gpt-5.6-terra |
10,000 ×2 | $0.36 |
| 2b, 2d vote | P2 google/gemini-3.6-flash |
8,000 ×2 | $0.16 |
| total | ≈ $0.88 |
Against $4.360490 headroom on UTC day 2026-07-29. Fits. Part B costs nothing — lead translation is free and is never ledgered (charter §3, A4).
Declared fall-through. P4 has returned an empty body once in this project at
max_tokens 6,000 (note (b)); note (bdl) says the remedy that worked was raising the
cap, not the declared fall-through, so the cap is raised to 16,000 here. If P4 still returns
no body, fall through to P5 deepseek/deepseek-v4-pro at 32,000, then to P3. P3 as
critic would collide with P3 as summariser; if that fall-through is reached the collision is
declared on the result page and the summariser moves to P5.
AMENDMENTS (v2) — made 2026-07-29 after the pre-run critic pass, before any other call
critic/dispositions.md carries the findings verbatim and the reasoning. Five findings,
all accepted. What changes in this design:
V2-1 (F1). Failure criterion 1 is replaced. A non-C verdict from Arm N's vote is reported as "the ratified outcome failed to reproduce in a single instance under a changed protocol; the cause is not identifiable between summary authorship, vote run-to-run variance, and the declared protocol departure." It still opens a decision page. It no longer asserts a cause. And the design's sentence "That is a one-cell question and one cell answers it" is withdrawn — it contradicted the S053 result the same paragraph cited.
V2-2 (F1). A second non-lead summary is added. Stage 1 runs twice: N1 from P3
x-ai/grok-4.5 and N2 from P5 deepseek/deepseek-v4-pro (max_tokens 32,000, note
(bdl)), from byte-identical prompts. Arm N's vote runs on both. This bounds
summary-authorship variance and nothing else; it does not measure vote variance and is not
reported as if it did. If P5 returns no body, the arm runs on N1 alone and V2-1 carries the
result.
V2-3 (F2). "Neutral" is withdrawn. The artifact is a non-lead summary throughout. The obligation's own words are "a summary the lead did not write"; the ratification record's looser "a neutral summary" is a stronger thing this design does not deliver and does not claim.
V2-4 (F2). Two gates on each summary, both frozen before stage 1 runs
(materials/loadbearing-checklist.json): a coverage check over 17 load-bearing points,
mechanical by regex plus a labelled lead reading; and a tilt gate counting directive,
verdict-shaped, recommendation, hedge and intensifier language — reported for each summary
and for the lead's own Arm L block, so the comparison is same-corpus rather than against an
absolute standard. Whatever the gates say is reported.
V2-5 (F3). Part B's device classification moves to a non-lead model. P1
openai/gpt-5.6-terra, shown the frozen census, the frozen English and the device
vocabulary, and nothing else — not the design, not the predictions, not the 1904 claim. The
lead's own classification is kept beside it as a labelled second reading,
internal-judgment-only, with disagreements printed rather than resolved. P1 also runs
Part A's two review calls; the calls are stateless and share no material, so this is a
weights-correlation concern and not a leak, and it is declared as such.
V2-6 (F3). Part B's licensed sentence narrows. B1 remains partly a compliance measure — an author who has registered "zero stage-directions" will not write one. What Part B may say is: these devices exist and an independent reader finds them in this text. It may not say that a translator ignorant of the prediction would have found them.
V2-7 (F4). A3 and A5 get declared consequences. A3 failure (C returned but not narrowed) counts as A1 failure — the narrowing is the binding part of the outcome. A5 failure means condition (iii) is still discharged with an explicit note that whole-page evidence did not surface the scope problem, which weakens the rationale for the re-derivation repair itself.
V2-8 (F5). "Part B costs nothing" → "Part B incurs no API cost." And the two priming exceptions are named rather than gestured at:
- Andreyev's «Баргамот и Гараська» was considered as the translation limb and rejected: triaging it printed the tail of the stored Lowe comparator and primed the lead at the exact site that made the story attractive. Method note (bdn).
- Of Seltzer's English of the tale actually chosen, one fragment has been seen — the final clause, "five kopeks. Now, Muzhik, rejoice.", printed by the slice-boundary check that extracted the comparator from the Gutenberg anthology. Nothing else. It falls on census site S31 and the log declares it there.
Revised cost, worst case: the two added calls (N2 summary, P1 device classification) take the pre-flight worst case from ≈$0.88 to ≈$1.10, against $4.293 headroom after the critic call. Fits.