Repository path: workshop/experiments/E-20260806-carrier-or-knowledge/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260806-carrier-or-knowledge |
| status | frozen |
| created | 2026-08-06 |
| updated | 2026-08-06 |
| senses | voice, style-correspondence, naturalness |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-voice-persona.md, wiki/goodness-senses.md, workshop/regimes/R19-isolated-carrier.md, workshop/experiments/E-20260806-carrier-or-knowledge/materials/specifications.md, workshop/translations/zapiski-oct3/R19-v1/translation.md, workshop/translations/zapiski-oct3/R19-v2/translation.md, workshop/translations/zapiski-oct3/R19-v3/translation.md, wiki/findings/results/RS-20260805d-two-persons.md, config/models.md, config/budget.md |
E-20260806 — is a narrator's person carried by the five properties voice names, or by the one it omits?
Frozen 2026-08-06, S117, ARM-voice-persona step 2. No call has been dispatched. The three
renderings and their logs were frozen first; the specifications were committed before any of them
existed, at 96249bb.
AMENDED after stage 0, before any other dispatch —
critic.md, amendmentsA1–A6. The independent pre-run critic returnedNEEDS-AMENDMENT, six findings, four BLOCKING, all six accepted (one in part, with the departure recorded). The amendments add a fourth rendering (A′,T-zapiski-oct3-R19-v4), rebuild the content-parity screen that was blind to the arm it most needed to screen, raise the manipulation check from two seats to three with numerical isolation thresholds, raise the primary from three seats to four, add failure criterionF7, and narrow what the primary is licensed to say. Every amended passage below is marked[A<n>].
Every sentence this design licenses is provisional. Tier D is NOT PASSED. No translation is
judged in this run. Every seat is asked who is speaking or what does this text say; none is
asked whether anything is good, and nothing here may support a quality claim about any member.
1. The question, and why the arm is spending its second session on a run
wiki/goodness-senses.md §voice asserts that a narratorial perspective's "observable carriers
are register, rhythm, diction temperature, idiosyncrasy and distance." RS-20260805d §4 measured
those five against three properties the entry does not name and found the unnamed three moving more
(1.667) than the five named (1.067), with idiosyncrasy saturated at 7 in all eight profiles —
a scale that agreed Multatuli's prose is mannered and discriminated nothing.
ARM-voice-persona step 2 is "carry the verdict into the entry." The verdict that would change the
entry most is the carrier finding, and it cannot be carried on RS-20260805d alone: one passage, one
author, one language, one translator, and — decisively — that run never manipulated the carriers.
It manipulated a person and measured the carriers as a side effect. A side-effect measurement
cannot say which properties do the work, because nothing in it was held.
So this run manipulates the properties directly and disjointly, in a second language and a second author, and asks the only question that discriminates:
Hold a narrator's self-knowledge fixed and move all five named carriers as far as the propositions allow; then hold all five fixed and move his self-knowledge. Which change makes a source-blind reader say he has met a different person?
What this teaches about translating literature (the subject-rule sentence, wiki/tracks.md): it
tests whether the features a translator is told to hold on to when carrying a narrator across a
language are the features that decide whether the reader meets the same man.
2. Materials
| id | what | words |
|---|---|---|
SRC |
Gogol, «Записки сумасшедшего» (1834), Октября 3 ¶1. Public domain. materials/source.txt |
306 RU |
C |
T-zapiski-oct3-R19-v1 — baseline, no MOVE set |
433 EN |
A |
T-zapiski-oct3-R19-v2 — MOVE K1–K5, HOLD K6 |
540 EN |
B |
T-zapiski-oct3-R19-v3 — MOVE K6, HOLD K1–K5 |
465 EN |
A′ |
[A5] T-zapiski-oct3-R19-v4 — MOVE K1–K5, HOLD K6, under a hard budget of ten edited sites. Added by critic finding 6; drafted after stage 0 and before any further dispatch |
455 EN |
N |
an unbriefed paraphrase of C, generated at stage 1 by a non-panel seat told only to say the same thing a different way |
— |
CMP |
Field 1916 English, public domain, extracted mechanically by heading index, never displayed, for the post-freeze contamination measurement only | 1,189 EN |
The six declared carriers K1–K6 and both variant specifications are at
materials/specifications.md, frozen before drafting.
3. Why N is in the design, and what it is a control for
A rewrites nearly every word; B changes eight sites (its log enumerates them exhaustively).
A raw comparison of C–A against C–B is therefore confounded with the extent of textual
change, and that confound is named here rather than discovered later.
N is the extent control. It rewrites nearly every word of C while aiming at nothing, so it
measures how much "different person" a reader will report for wholesale rewording alone. It is
the floor A must clear to have bought anything with its five moves, and — because B rewrites far
less than N does — any separation B achieves over N is achieved against the extent confound
rather than with it.
This is a deliberately weaker use of the paraphrase arm than RS-20260805d's F4, and the
difference is registered here before the data. There, PN reaching parity with the deliberate
rewrite withheld an absolute claim that a persona reaches readers, and rightly. Here the primary
is a contrast between two manipulations, and rewording noise common to both does not undermine a
contrast the way it undermines an absolute claim. N therefore gates the primary only under F3
below — when it beats both manipulations, i.e. when neither is distinguishable from noise.
A second, mechanical extent measure is registered as a secondary: token-level change fraction
from C, computed by analysis/extent.py (difflib.SequenceMatcher over lowercased word tokens,
1 − matching ratio), reported for A, B and N, with each arm's mean rating divided by it.
Computed before dispatch and it is the sharpest fact in this design: A changes 0.7012 of C's
tokens and B changes 0.0396 — a factor of 17.7. So P1p and P2p are not symmetric bets. P2p
(the entry's claim) can be satisfied by extent alone and would be weak evidence even if it fires;
P1p requires the arm that rewrote 4% of the text to outrun the arm that rewrote 70% of it,
and cannot be explained by extent at all. A P2p firing therefore does not vindicate the entry's
list, and this is registered now, before the data, so that it cannot be decided afterwards.
[A1] The demotion of P2p is now formal: a P2p firing may be reported only as not
distinguishable from extent, never as support for the entry's carrier list. And F7 is added at
§6: if A does not beat the aimless floor N, P1p is unreadable, because B would then be
outrunning noise rather than a manipulation.
[A5] The extent confound is no longer a single comparison but a ladder of four arms, the
critic's fix (b) to finding 6:
| arm | what moved | change fraction from C |
|---|---|---|
B |
K6 alone, eight sites |
0.0396 |
A′ |
K1–K5, ten sites |
0.1767 |
A |
K1–K5, unbudgeted |
0.7012 |
N |
nothing, aimlessly | measured at stage 1 |
A′ is 4.5× B's extent where A is 17.7×, so the B-vs-A′ contrast is the fairer one and
is registered as P1q beside the primary rather than instead of it.
4. Procedure
Five stages. Every raw body is written to runs/ before any parse (note (bdt)); finish_reason:
length is a seat failure, never a partial answer; a reserve is declared per stage before dispatch
(note (bfc)); key-usage snapshots are write-once (note (bgq)).
- Stage 0 — independent pre-run critic. One seat, non-panel, shown this design, the three
specifications, all three renderings and all three logs. Reserve:
qwen/qwen3.7-max. Findings are applied before any other call is dispatched and the amendments are recorded incritic.md. - Stage 1 —
N. One call tomistralai/mistral-medium-3-5(non-panel, asE-20260805d), prompt: the text ofCand "Say the same thing a different way." Nothing about persona, voice, narrators, register or self-knowledge. Reserve:z-ai/glm-5.2. - Stage 2 — the manipulation check (transmission, not a finding).
[A3][A5]Five texts (C,A,A′,B,N) × three seats:z-ai/glm-5.2,qwen/qwen3.7-maxandmistralai/mistral-medium-3-5. Each seat rates one text, with no comparison and no source, on the six declared carriersK1–K6, 1–7.mistral-medium-3-5writesNat stage 1 and therefore does not rateN— 14 calls, andN's profile rests on two seats where the other four rest on three, declared here rather than discovered later.max_tokens6,000. Reserve: P4. - Stage 3 — the content-parity screen.
[A2][A5]Three pairs (C–A,C–A′,C–B) × two seats:z-ai/glm-5.2andqwen/qwen3.7-max. Each is shown two passages and asked to enumerate every difference in what the two texts assert, excluding only differences of wording, tone, formality and sentence length, and to label each differenceEVENT(something that happened, or something stated about a person other than the narrator) orSELF(something stated about the narrator's own thoughts, habits or handling of his material). No exclusion category is named to the seat and neither label is flagged as the one the design cares about.max_tokens6,000. Reserve: P4. - Stage 4 — the primary.
[A4][A5]Four pairs (C–A,C–A′,C–B,C–N) × four seats (P1openai/gpt-5.6-terra, P2google/gemini-3.6-flash, P3x-ai/grok-4.5, P5deepseek/deepseek-v4-pro) × two presentation orders = 32 calls, each independent and stateless. None of these four seats holds any other role in this run — the critic, the paraphraser, the three manipulation-check seats and the two parity seats are all drawn from outside them. Each call shows two passages labelled Passage 1 and Passage 2, no source, no authorship, no mention that either is a translation, and asks:Q1. On a scale of 0 to 6, are these two passages narrated by the same person? (0 = certainly the same person; 6 = certainly two different people.) Q2. In one or two sentences, what in the words makes you say so? Q3. If you judged them different, name the single property that differs most.
max_tokens 4,500. Reserve: z-ai/glm-5.2. Judgment is not parallelised across a seat's own
cells: each cell is one call and no seat sees another seat's answer.
5. Predictions and the primary statistic — registered before dispatch
Let r(X, s, o) be Q1's rating for pair C–X, seat s, order o; six cells per arm.
[A4] d (s,o) = r(B,s,o) - r(A ,s,o) 8 paired cells (4 seats x 2 orders)
[A5] d'(s,o) = r(B,s,o) - r(A',s,o) 8 paired cells
T = sum d(s,o) exact sign-flip permutation over 2^8 = 256
P1p— THE PRIMARY.mean d ≥ 1.00andd > 0in ≥ 6 of 8 cells. Holding: moving the one property the entry omits changes who the reader meets more than moving all five it names.P1q— the extent-nearer contrast[A5]. The same bars ond′.A′is 4.5×B's extent whereAis 17.7×, soP1qis the fairer test and is reported besideP1p, never instead of it. A design that swapped its primary after seeing the critic would be choosing its statistic; this one adds and keeps both.P2p— the registered alternative, which is the entry's own claim.mean(-d) ≥ 1.00andd < 0in ≥ 6 of 8 cells: the five named carriers are where the person is.[A1]Reportable only as not distinguishable from extent — see §3.- Neither firing is a null on the contrast, and it is a real outcome: the two manipulations are not distinguishable at this power.
P3d(descriptive). Q2/Q3 free text forC–Bnames knowledge, irony, self-awareness or self-relation more often than forC–A;C–A's names diction, formality, register or sentence length more often thanC–B's. Coded by two independent string-match passes inanalysis/score.pyagainst keyword sets fixed in this file (below), never by hand.P4d(descriptive).mean r(N) < mean r(A)and< mean r(B).P5d(transmission check, never reported as a finding). Stage 2 profiles move as specified:AmovesK1–K5more thanK6;BmovesK6more thanK1–K5.
Keyword sets, fixed here (case-insensitive substring, on Q2+Q3 concatenated).
KNOW = {self-aware, self-awareness, self-knowledge, aware of, irony, ironic, ironical, wry,
self-deprecat, knows himself, sees himself, self-critical, candid, honest with himself, insight,
detach} · SURF = {register, formal, formality, diction, vocabulary, syntax, sentence length,
sentence structure, latinate, elevated, style, tone, ornate, verbose, colloquial}
The lead's own expectation, registered before dispatch as RS-20260805d §6 did. I expect
P1p to fail and P2p to come nearer: point predictions mean r(A) = 3.5, mean r(B) = 3.0,
mean r(N) = 2.0, i.e. mean d = −0.5. The reasoning is that A's change is enormously louder and
readers asked same person? will hear the loudness. I am genuinely unsure of this, and it is
written down so the run can be wrong about it in public.
6. Failure criteria — pre-registered, each with what it does
F1— the manipulation did not isolate.[A3]WithΔ_X(K) = profile_X(K) − profile_C(K)averaged over the three stage-2 seats, all three of the following must hold, andF1fires and the primary is WITHHELD if any fails:mean|Δ_A (K1..K5)| - |Δ_A (K6)| >= 0.50 mean|Δ_A'(K1..K5)| - |Δ_A'(K6)| >= 0.50 |Δ_B(K6)| - mean|Δ_B(K1..K5)| >= 0.50What a firing would establish is not nothing: these six properties are not independently movable in translation, which is itself a claim about the entry's list. Stage 4 is dispatched either way — the departure from the critic's requested fix, with its reason, is incritic.md§F3.F2— content parity.[A2]If the union over the two screen seats ofEVENT-labelled differences exceeds 2 for any ofC–A,C–A′,C–B, that arm is WITHHELD. The screen's own counts and the screen's own labels govern; the lead does not adjudicate them (RS-20260805d§3).SELF-labelled differences are reported and do not gate — and they are a check on the lead's own logs, which declare three forBand none forAandA′. If the screen finds more than the logs declare, the logs are wrong and the result says so. A declared, unscreened difference, stated here in advance:B's log enumerates three added propositions, all of them first-person statements about how the narrator handles his own material, and the screen is instructed not to count them. They are the manipulation. Self-knowledge has no purely substitutional exponent in English —B's log argues this at length — so a design that demanded zero addition would be a design in whichK6cannot be moved at all. What follows is a hard limit on the primary and is written into any result:Bdiffers fromCby three self-descriptive assertions as well as by their marking effect, and this run cannot separate the two.F3— the floor beats every manipulation. Ifmean r(N)is≥the mean of all three ofA,A′,B, the primary is WITHHELD: no manipulation is distinguishable from wholesale rewording.F4— seat degeneracy.[A4]A seat returning an identical Q1 rating across all eight of its cells is not discriminating. If ≥ 2 of 4 seats do, the primary is WITHHELD.F7— the loud arm must beat the floor.[A1]Ifmean r(A) ≤ mean r(N),P1pmay not be read as support for anything:Awould not then be a manipulation that reached readers at all, andBoutrunning it would beBoutrunning noise.P1qis subject to the same condition onA′.F5— order. If the mean absolute within-pair order difference exceeds 2.0 scale points, every magnitude claim is withheld and only signs are reported.F6— dispatch. Fewer than 5 of 6 admissible cells on any arm withholds that arm.
7. What this run cannot establish, whatever it returns
[A4] This run supports a pilot-level signal on one passage, one author, one pair, one translator,
and licenses no generalisable claim about voice. Eight paired cells over four seats is enough to
see a large effect and not enough to size one.
[A6] What the primary is licensed to say, fixed here before the data. The critic's finding 6 is
that this design cannot isolate a latent property; it can only compare operations a translator can
actually perform. So the licensed sentence is about operations:
A translator who changed N% of a passage at the site of the narrator's self-relation moved source-blind readers' judgment of who was speaking further, or less far, than one who changed M% of it along all five properties the entry names.
Not the five carriers do not carry voice. A′'s log already establishes, from the translator's
side and before any reader was asked, why the two operations cannot be made parallel: K4 and K5
are located properties that ten edits move, K1 moves partially, and K2 and K3 are distributed
over every sentence boundary and every concrete noun in the paragraph.
- That the lead's Poprishchin is Gogol's. No source-side yardstick is in this design, deliberately:
the question is which properties carry a person, not whether the person is apt.
voiceis a source–target relation sense and this run measures only the target-side half of it. - That any member is good. Nothing is judged; Tier D is NOT PASSED.
- Anything about
voicebeyond this passage, this pair, this translator. - Anything that separates
B's three added propositions from their marking effect (F2).
8. Contamination
Declared suspected on all three members and measured after the freeze, not before, because
this design's validity does not turn on independence from any published rendering: all three
members are the lead's, compared against one another, and an overlap with Field 1916 that is shared
by all three cannot separate them. The measurement is reported for the record, per the standing rule
that a contamination: declaration without a measurement is a placeholder. tools/dependence_check.py
unmodified, C/A/B against CMP; no line of CMP was displayed to the translator.
9. Pre-flight cost estimate
Built from max_tokens, not from expected output (note (abc)), with a 2× routing margin on every
stage (the S022 provider caution).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| 0 critic | 1 | 16,000 | $0.0477954 actual |
| 1 paraphrase | 1 | 4,000 | $0.05 |
2 manipulation check [A3] |
14 | 6,000 | $0.58 |
3 parity screen [A2] |
6 | 6,000 | $0.30 |
4 primary [A4] [A5] |
32 | 4,500 | $1.76 |
| total | 54 | $2.74 |
Declared worst case: $2.90, revised upward from $2.40 by the critic's amendments; the fourth rendering itself cost $0. UTC day 2026-08-06 opens with the full $5.00 and no session has spent today, so the revised worst case is 58% of the cap. A stage that does not fit remaining headroom is dropped, not shrunk silently.