Repository path: workshop/experiments/E-20260805d-persona-two-ways/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260805d-persona-two-ways |
| status | frozen |
| created | 2026-08-05 |
| updated | 2026-08-05 |
| senses | voice, style-correspondence, naturalness |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-voice-persona.md, wiki/goodness-senses.md, workshop/regimes/R18-declared-persona.md, workshop/canon/max-havelaar-i-b/manifest.md, wiki/findings/results/RS-20260802-voice-warrant.md, wiki/findings/results/RS-20260731d-sense-axes.md, wiki/method-notes.md, config/models.md, config/budget.md |
E-20260805d — the same narrator written as two different people
ARM-voice-persona step 1 (T2). Frozen before any seat is addressed and before either
rendering was drafted. Commit order is the guarantee.
This page carries senses: and judges nothing (note (bha)). No seat in this run is asked whether
any text is good, faithful, natural or well made; every seat is asked who is speaking. Nothing
here licenses a quality claim about either rendering, and Tier D is NOT PASSED, so nothing would
license one anyway.
1. What is being asked
wiki/goodness-senses.md §voice says the sense assesses "whether the translation realizes, for its
readers, the source work's characterized authorial or narratorial perspective", names five carriers —
register, rhythm, diction temperature, idiosyncrasy, distance — and says the sense exists so that a
translation cannot score well for "inventing an attractive but source-inapt persona". Its
reachability note (S097) then states, in one sentence, the design that would reach the positive claim
and does not exist:
a design in which the same source persona is realised two ways on purpose and a blind jury is asked which reader met which person, with the source withheld from them.
This is that design. It asks two questions and a third falls out:
- Q1 — determinacy. Do readers who see one English text and nothing else converge on the same
person? If they do not,
voiceis not scoreable by a source-blind jury at all, whatever its definition says. - Q2 — aptness. Of two renderings of the same passage built to present two different people, is the one built to present the source's narrator the one that a source-only reader's description of that narrator picks out?
- Q3 — carriers.
voicestipulates five carriers and has never tested them. Which of them move when the person is deliberately changed, and do three unstipulated properties move more?
2. Materials
- Source. Multatuli, Max Havelaar (1860), chapter I, body paragraphs 7–9. 521 Dutch words.
workshop/canon/max-havelaar-i-b/manifest.md;sha256of the frozen extent recorded there. A disjoint extent from the project's only previous Dutch unit (¶1–4, S077); nothing is retranslated. - Why this passage. Droogstoppel is the strongest case the project can reach for the sense: the whole rhetorical point of Multatuli's first chapter is that the reader meets a determinate person and understands him better than he understands himself. If a persona is transmissible at all, it is transmissible here — which makes a null in this run a strong null.
PA—T-max-havelaar-i-b-R18-v1, rendered by the lead under persona specification A.PB—T-max-havelaar-i-b-R18-v2, rendered by the lead under persona specification B.PN— an unbriefed paraphrase ofPA, written by a seat told nothing about the source, the study, the persona, or anything to preserve. This isE-20260805c'sPARAarm imported as a control (note (bix)), and it is the arm that decides whether the measure is about a person or about wording.- Comparator (Nahuÿs 1868, whole chapter I) extracted and not opened; used only for the post-freeze contamination measurement. It enters no seat prompt.
2.1 The two persona specifications — frozen here, before either rendering exists
They are character sketches, not descriptions of textual properties, per R18 constraint 2 and
note (bio). The mapping from a specification to a predicted movement on the rating instrument is a
prediction (§5), declared separately and below.
⚠ Both specifications were REWRITTEN after the pre-run critic pass and before any rendering existed
— amendment A1, critic.md. The critic found that the frozen wording reused the rating
instrument's own anchor phrases, which is note (bio)'s failure. Note (bio) forbids rewriting a
freeze once the material it governs exists; nothing existed, which is exactly why the critic runs
here. The originals are preserved verbatim at the end of this section and in commit b2fff5e. The
substantive overlap between a specification and a scale is not removed and cannot be — a design that
changes a property on purpose and then measures that property must name it twice — and the consequence
is registered in A2 below rather than concealed.
Specification A — the man who does not know. The man speaking is a middle-aged Amsterdam broker. It has never crossed his mind that he might be a figure of fun, or that in telling you about the theatre he is telling you about himself. He says a thing, and then the next thing, and stops when there is no more to say; he breaks off to correct a word or to name a sum, and the sum is exact to the penny. He has one test of whether something is true, which is whether it happened to him, and one test of whether something is good, which is whether it pays. He is talking to you across a counter, at a person whose agreement he takes for granted. His indignation is real indignation.
Specification B — the man who knows. The man speaking is a practised talker with an audience in mind. Everything he reports about the theatre he has chosen because it is preposterous, and he lays the ground before he lets it off; he keeps the best of each item back until the last possible moment. He knows perfectly well what impression he is making, and he is making it on purpose. He is no kinder than the other man and pretends to be no kinder; he speaks straight at you and over nobody's head; and he entertains no doubt at all that he is right. But he is giving a performance, and he assumes you can see that.
Declared and not concealed: specification B's penultimate sentence deliberately holds fixed
three of the properties the rating instrument measures. Those three scales are held by construction,
and §5's P2p is a manipulation check on transmission, not a discovery. The same is now registered
for the five shifted scales: a shift that shows up is evidence that the change of person reached a
source-blind reader, and is not evidence that the design discovered which properties constitute a
person.
The superseded specifications as first frozen, at commit b2fff5e
> **A (superseded).** The man speaking is a middle-aged Amsterdam broker. He has no notion that anyone
> could find him funny, and no notion that he is telling you anything about himself. He says things in
> the order they occur to him and stops when he is finished; he interrupts himself to correct a word
> or to name a sum, and the sum is always exact. He has one test of whether a thing is true, which is
> whether it has happened to him, and one test of whether a thing is good, which is whether it pays.
> He talks to you across a counter, at somebody he is sure of, and expects to be agreed with. When he
> is indignant he is indignant in earnest.
>
> **B (superseded).** The man speaking is a practised, urbane talker who finds all this funny and
> expects you to. Every absurdity he reports he has chosen because it is absurd, and he sets it up
> before he delivers it. His sentences are built: they begin somewhere and they arrive, and what is
> funny waits at the end. He is entirely aware of the figure he cuts and uses it. He is not a kindlier
> man than the other; he talks straight at you and never over your head; and he has no doubt whatever
> that he is right. But he is performing, and he knows you know.
2.2 Construction rule for the pair — content held fixed
R18 P5: every proposition asserted in one rendering is asserted in the other, in the same order.
Where the person makes a proposition awkward, the proposition is kept and the wording changes. Both
renderings declare British English and hold it (R18 P6), so orthography cannot separate them.
The renderings are written in the order A then B, B from the Dutch and not from A, and the log
records whether that held.
3. The rating instrument — eight bipolar scales, frozen
Every seat that describes a person, whether from the Dutch or from an English text, answers the same
eight items. Scales S1–S5 are voice's own five stipulated carriers, one scale each, in the order
the entry names them. S6–S8 are three properties the entry does not name, included so that Q3 has
something to compare the five against.
S1 REGISTER 1 = plainly colloquial, the way a man talks in a shop
7 = formal and official, the way a document is written
S2 RHYTHM 1 = abrupt, short, broken off, one thing after another
7 = flowing and built, long sentences that arrive somewhere
S3 TEMPERATURE 1 = cold and dry
7 = warm and effusive
S4 IDIOSYNCRASY 1 = anonymous; this could be anybody's prose
7 = strongly mannered; unmistakably one particular person's
S5 DISTANCE 1 = close; buttonholing you directly
7 = remote; addressing nobody
S6 SELF-AWARENESS 1 = has no idea how he sounds
7 = knows exactly how he sounds and is in control of it
S7 ASSERTIVENESS 1 = tentative, hedging, unsure
7 = dogmatic; states things as settled
S8 HUMOUR 1 = entirely in earnest; not joking
7 = deliberately funny; joking on purpose
S8's wording is load-bearing and is chosen, not inherited. It separates this narrator is comic from this narrator is joking — the distinction the whole passage turns on, since Droogstoppel is funny and is not making jokes. A scale worded "how funny is this" would have been answered the same way for both renderings and could not have failed.
Each seat returns eight integers and then, in two to four sentences, who the speaker is in its own words. The free description is for the result page's exhibits and for the leak screen; no primary statistic is computed from it.
4. Procedure — stages, seats, order
Seats are named as roles; slugs resolve from config/models.md and are logged as provenance.
| # | stage | who sees what | slugs | cap | attempts |
|---|---|---|---|---|---|
| 0 | seat probe (note (bit)) | a 90-word English passage + the instrument | P1, P3, P5, qwen/qwen3.7-max, z-ai/glm-5.2 |
5,000 | 1, no retry |
| 1 | pre-run critic | this design + both specifications + the Dutch source | google/gemini-3.6-flash (P2), fallback mistralai/mistral-medium-3-5 |
16,000 | 2 |
| — | lead translates PA and PB, freezes both logs, commits |
— | — | — | — |
| 2 | source-side yardstick Y |
the Dutch only + the instrument | qwen/qwen3.7-max, z-ai/glm-5.2 |
8,000 | 2 |
| 3 | paraphrase PN |
PA only, with S111's PARA wording and nothing else |
mistralai/mistral-medium-3-5, fallback P5 |
8,000 | 2 |
| 4 | ratings | one English text only, no source, no comparison | P1, P3, P5 × {PA,PB,PN} |
5,000 | 2 |
| 5 | content-parity screen | PA and PB side by side |
qwen/qwen3.7-max, z-ai/glm-5.2 |
8,000 | 2 |
Why the critic comes before the translations. Note (bio)(ii): where a specification governs production, the critic must see it before the material is produced. Both specs are frozen above and the critic reads them; if it finds a spec restating the measure, the claim is amended and the freeze is not.
Why the parity screen comes last. It is a gate on what may be concluded, not on what may be collected, and its seats see both renderings together. Dispatching it after stage 4 keeps the rating seats' naivety independent of it. Its verdict still withholds the primary if it fails.
Seat exclusions, declared. moonshotai/kimi-k3 (P4) is used nowhere in this run. Note
(bhf) has fired on it five times for zero-content bodies, twice as a critic; and note (bhf)(i)'s
own remedy — do not hang a registered control on a seat that has failed the shape — applies. P2 is
used once, on the critic, where a failure costs a re-dispatch and not a control.
No seat rates a text it wrote. The paraphraser (mistral-medium-3-5) is not among the rating
seats. The yardstick seats see the Dutch before they see any English, and the dispatch order enforces
it; every call is stateless and independent.
5. Statistics, predictions and the null of every branch
Let Y_s = mean(Y1_s, Y2_s) be the source-side profile on scale s, and r_{s,i}(X) seat i's
rating of text X on scale s. Ratings are integers 1–7, so every quantity below is a multiple of
0.25.
Primary statistic. For each of the 24 cells (s, i), s ∈ S1..S8, i ∈ {P1,P3,P5}:
d_{s,i} = |Y_s − r_{s,i}(PB)| − |Y_s − r_{s,i}(PA)| T = Σ d_{s,i}
d > 0 means the seat's reading of PA is closer to the source-only reading of the narrator than its
reading of PB is.
P1p(primary) fires ifT > 0and the exact permutationP ≤ 0.05and the same statistic computed on profile-centred ratings has the same sign.- What
P1pis licensed to say (amendmentA2, narrowed after the critic and before any datum): a deliberate change of narratorial person, with propositional content held fixed, moves source-blind readers' description of the speaker away from an independently authored source-side description of that narrator. Nothing stronger. It cannot separate the person transmitted from the lead's reading of the narrator happening to coincide with the yardstick's — and the interpretive rule for that is registered here, in advance: ifS3is large (the yardstick reads the Dutch narrator differently from the translator who wrotePA) andP1pstill fires, the result is not conditional on the lead's reading and is stronger; ifS3is near zero,P1pis close to a transmission check and the result page says so in those words. T_excl12(amendmentA4), registered as a subsidiary and reported always:Trecomputed with S1 and S2 dropped — the two scales a rater can answer from sentence length and subordination alone, both of which specification B moves by construction. 18 cells, its own exact sign-flip null. IfTfires andT_excl12does not, the result is attributed to sentence architecture and not to person, in those words.- The null. Under H0 the labels
PA/PBare exchangeable within a cell, so eachd_{s,i}is equally likely to be+dor−d.Pis the exact one-sided tail ofΣ ±d_{s,i}over all2^24sign assignments, enumerated by dynamic programming over the (integer × 4) sums. Not sampled. - Why centring is a conjunct and not an alternative.
Yis read off Dutch and the ratings off English, so a constant per-profile offset between the two rating conventions is possible and would bias a raw distance. Subtracting each profile's own mean across the eight scales removes it. Requiring both statistics to agree is stricter than either alone and is fixed here, before any datum exists, so that neither can be selected after the fact. - Power, stated rather than assumed. With 24 cells and within-cell noise of about 1.5 scale
points, a uniform per-cell effect of +0.5 gives
P ≈ 0.06and +0.75 givesP ≈ 0.01. Only five of the eight scales are shifted by construction, so the manipulation must move those five by roughly 1.2 scale points for the primary to fire. The three held scales dilute the primary and are included anyway, because choosing the scales after seeing which moved is the defect this arithmetic exists to prevent.
Declared movement, PB minus PA — the prediction, written before either rendering exists:
| scale | direction | shifted or held |
|---|---|---|
| S1 register | + more formal | shifted |
| S2 rhythm | + more flowing | shifted |
| S3 temperature | 0 | held by construction |
| S4 idiosyncrasy | − less mannered | shifted |
| S5 distance | 0 | held by construction |
| S6 self-awareness | + much more | shifted |
| S7 assertiveness | 0 | held by construction |
| S8 humour | + joking on purpose | shifted |
P2p(transmission of the held dimensions) fires if mean|r(PB) − r(PA)|over the three held scales is below the same quantity over the five shifted scales. It is a manipulation check, not a discovery: §2.1 says the held properties are held in the specification itself.P3p(carriers, descriptive, no threshold). Per-scalemean_i |r_{s,i}(PB) − r_{s,i}(PA)|, reported for all eight, with the five stipulated carriers (S1–S5) and the three additions (S6–S8) totalled separately. No hypothesis is registered on which group moves more, because the project has no prior that would make one; the number is the finding either way.
The lead's own reading, registered so it can be wrong (S3 below). Before dispatch the lead's
prediction of Y is: S1 3 · S2 2 · S3 2 · S4 6 · S5 2 · S6 1 · S7 7 · S8 2. S111's exhibit S10 —
where a divergent yardstick penalised the translator's reading — is the reason this is written down in
advance rather than discovered.
Secondaries, all descriptive:
S1source determinacy.mean_s |Y1_s − Y2_s|. Two independent readers of the same Dutch who do not agree about who is speaking would be a finding about the sense larger than anything else here.S2target determinacy (Q1). For each text,mean_s mean_{i<j} |r_{s,i} − r_{s,j}|. Chance is 16/7 = 2.2857 for two independent uniform draws on 1–7; that figure is the comparison, not a vague "they agreed".S3lead-vs-yardstick divergence.mean_s |Y_s − lead_s|against the registered prediction above. This is the quantityA2's interpretive rule is keyed on.S4is aptness cheap? (amendmentA7).mean_s |Y_s − r_s(PN)|againstmean_s |Y_s − r_s(PA)|.E-20260805c's finding asked of a person rather than of a marking: if an aimless rewording sits as close to the source-side description of the narrator as the rendering written to be apt, note (bix) applies tovoicetoo.A_magpersona fragility (amendmentA3).mean_{s,i} |r_{s,i}(PN) − r_{s,i}(PA)|— how far a persona drifts when a machine is asked, with no brief at all, to say the same thing differently. A first-class result, and no longer a gate.
6. Failure criteria — what withholds what
Each states the branch it is on and, where computable, the probability that it fires by accident.
F1source determinacy. Ifmean_s |Y1_s − Y2_s| ≥ 2.00, or the two profiles differ by ≥ 3 scale points on more than two of the eight scales, the source-side person is not determinate andP1pis reported descriptively only.Yremains the mean in either case. Null firing rate, computed by the critic and imported: ≈ 0.85 under two uninformative raters, because chance disagreement is 2.286 and the bar is 2.00.F2target determinacy. IfS2exceeds 1.50 forPAor forPB, source-blind readers do not converge on a person andP1pis withheld. 1.50 is 66% of the 2.2857 chance value. Null firing rate: ≈ 0.9999. PassingF2is licensed to mean seat agreement better than chance by the stated margin, never the readers agreed, and the result page prints both numbers side by side. (The critic called these two rates a defect — "shutting down the experiment by accident". The arithmetic is imported and the reading is rejected in writing atcritic.mdA6: a determinacy gate that rarely fired on noise would be the defect. A persona nobody can agree on is the null this design exists to be able to return.)F3content parity. Two screen seats each enumerate propositions asserted in one rendering and not the other. If the union of substantive divergences (a claim about the world, not a wording difference) exceeds 3, the pair is not content-matched andP1pis withheld.F4the(bix)control — REBUILT DIRECTIONALLY, amendmentA3. The frozen version compared undirected L1 magnitudes and the critic showed it would have withheld the primary on a success: a paraphrase that normalises a fragile persona away makesAlarge exactly whenPAhad a persona. Withu_sthe declared direction of the shift on each of the five shifted scales (§5 table), over those five scales and the three seats:
B_dir = mean u_s ( r(PB) − r(PA) ) movement along the persona axis, intended
A_dir = mean u_s ( r(PN) − r(PA) ) movement along the same axis, unintended
P1p is withheld if A_dir ≥ 0.75 × B_dir — an unbriefed rewording walks the persona axis as
far as a deliberate change of person, so the axis is a property of wording. A_mag, the undirected
drift, is no longer a gate and is reported as a result (§5). If B_dir < 0.50 the manipulation
did not transmit at all, the run's result is that null, and F4 is not evaluated. Both
outcomes of F4 remain attainable, which is what note (bip) requires of a control.
- F5 leak (note (biw)). Every returned body is screened for (i) Dutch tokens, (ii) any quoted
run of ≥ 4 words from PA, PB or PN, (iii) any mention of translation, of a comparison, or of a
study. A hit on any stage-2 body triggers re-request of both stage-2 bodies, mechanically and
uniformly, with the lead declaring that it had seen the first set. Quotation, not just vocabulary.
- F6 incompleteness. If fewer than three rating bodies return for any text, the run reports what
returned, says which cells are missing, and does not impute.
- F7 length confound — GIVEN TEETH, amendment A5. As frozen it only logged a note, which the
critic correctly called decoration. If |len(PA) − len(PB)| / mean(len) > 12%, T_excl12 becomes
the reported primary and raw T is demoted to secondary. Mechanical, computed before dispatch,
with a consequence.
- F8 anchor-phrase screen — NEW, amendment A1, computed before dispatch. No content word from
any scale anchor may stand in PA or PB in the anchor's own sense. The screen enumerates every
hit; a hit that is an anchor phrase is repaired before dispatch, and the whole hit list is
published. This is the checkable half of the critic's F2 finding: substantive overlap between a
specification and a scale is unavoidable, an anchor phrase inside the measured prose is not.
7. What this run cannot establish
- Not a quality claim. Nothing here says either rendering is better, and Tier D is NOT PASSED.
Every sentence of the result is
provisional. - Not a claim that the lead's reading of Droogstoppel is right.
Yis two seats' reading of the Dutch. If it diverges from the lead's registered prediction, that is a fact about two readings and not an adjudication between them. - Not a general claim about
voice. One passage, one author, one language pair, one translator, three rating seats. Whatever fires, fires on Droogstoppel. - Cross-language commensurability of the scales is an assumption, mitigated by the centring conjunct and by the three held scales, and not eliminated.
- The held scales are held in the specification, so
P2pcannot discover that they are held.
8. Pre-dispatch computations (note (bhr))
Filled in before stage 0 dispatches. Every criterion computable without returned data is computed here, not at analysis time.
F7length ratio: (recorded inanalysis/predispatch.json)F2chance value:E|X−Y| = 16/7 = 2.285714…for independent uniform draws on {1..7}; verified by enumeration inanalysis/predispatch.json.P1pnull enumeration: the DP is unit-tested against brute-force enumeration on a reduced cell set inanalysis/verify.pybefore it is used on the real one.
9. Pre-flight cost
Built from max_tokens × attempts × slugs with a ×2 routing margin (notes (abc), (bgk),
(bhq), and S079's correction).
| stage | calls | cap | worst |
|---|---|---|---|
| 0 probe (1 attempt, no retry) | 5 | 5,000 | $0.23 |
| 1 critic (+1 fallback slug) | 1 | 16,000 | $0.54 |
| 2 yardstick | 2 | 8,000 | $0.33 |
| 3 paraphrase (+1 fallback slug) | 1 | 8,000 | $0.08 |
| 4 ratings | 9 | 5,000 | $0.77 |
| 5 parity screen | 2 | 8,000 | $0.33 |
| input tokens, all stages | 20 | — | $0.20 |
| declared worst case | 20 | $2.50 |
Checked against the day's headroom at stage 0 and again before stage 4; a stage that does not fit is
deferred and the deferral goes to NEXT.md.