Repository path: workshop/experiments/E-20260806f-persona-crossing/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260806f-persona-crossing |
| status | frozen |
| created | 2026-08-06 |
| updated | 2026-08-06 |
| links | wiki/arms/ARM-voice-crossing.md, wiki/goodness-senses.md, wiki/base/sources/S-venuti-invisibility.md, workshop/regimes/R04-lead-close.md, workshop/regimes/R14-matched-flattening.md, wiki/findings/results/RS-20260806-same-man.md, wiki/findings/results/RS-20260805d-two-persons.md, config/models.md, config/budget.md |
| senses | voice, style-correspondence |
| provisional | true |
| internal-judgment-only | true |
E-20260806f — does the person the source presents arrive in the translation, and is he determinate on the source side at all?
Frozen 2026-08-06 (S123), before any API call. Nothing below was written after a datum existed.
Arm: ARM-voice-crossing, step 1. Track T2.
1. Question
wiki/goodness-senses.md §voice defines the sense as a source–target relation: "whether the
translation realizes, for its readers, the source work's characterized authorial or narratorial
perspective… the reader of the translation meets someone, and the question is whether it is the
someone the source presents."
Two runs have now tried to reach that positive claim by the route the sense's own reachability note
specified — two renderings of one source built to differ in persona, a blind jury asked which
reader met which person — and both failed, the second one structurally: RS-20260806-same-man
established that the propositional parity that makes a paired rendering a controlled contrast is
what readers use to answer same person?, so no number of seats or languages repairs that design.
Its closing sentence names the successor: what would reach the claim is "a jury shown one
rendering and asked to describe the person, against a source-side description written by a reader who
saw no English, with the two descriptions compared for fit rather than for difference" — and it
adds that the next obstacle is on the source side, because in RS-20260805d two competent
readers of the Dutch disagreed at 5 of 7 about the narrator they had both just read.
This run builds that instrument, and it puts the source-side obstacle first, as a gate.
The Tier 2 challenge this bears on. Venuti's simpatico argument (S-venuti-invisibility ch. 6,
recorded against voice since S048 and standing unabsorbed) is that the sensation of having
captured an author's voice is "evidence about the translator, not about the source" — cultural
narcissism, the reader recognising himself. That argument has an empirical consequence this design
can test in one direction: if the narrator is not a determinate property of the source, independent
readers of the source will not converge on him. A null on the source side is Venuti's result. A
positive is not a refutation of him — he can still say the convergence is a shared projection — and
the result page will say so.
One sentence on the wire between the limbs (continue-prompt.md §4): the four translations are
the target-side objects whose independently described narrators are matched against source-side
descriptions the same run measures for determinacy, so the translation limb is what makes the study
limb's crossing measurable at all.
2. Materials, and the contamination gate
Four works, four languages, four language families. Every locus was chosen on one source-side rule, fixed before any English existed: a span in which the narrator presents himself and almost nothing happens — so that what a reader has to describe is a manner, not a plot.
| id | source | locus | size | copy-text |
|---|---|---|---|---|
dost |
Достоевский, «Записки из подполья» (1864), RU | I.1 ¶1–2 | 411 words | ru.wikisource |
daudet |
Daudet, Lettres de mon moulin (1869), FR | «Installation» ¶1–4 | 380 words | fr.wikisource (Charpentier 1895) |
sandmann |
Hoffmann, «Der Sandmann» (1816), DE | the address to the reader, ¶1 | 337 words | de.wikisource |
wagahai |
夏目漱石『吾輩は猫である』(1905), JA | 一, opening | 599 chars | Aozora 000148/789_14547 |
SHA-256 of every span is in materials/sources.json.
Contamination — measured, and the deviation from the standing rule's letter is declared here rather
than buried. CLAUDE.md requires the measurement before anything is translated. That is not
satisfiable together with R04 §Procedure 1, which forbids the translator to read a published
rendering before its own log is frozen: the check needs the lead's English on one side. The order
actually taken was: loci fixed on source-side grounds → four R06 drafts frozen → four R04
revisions and logs frozen → comparators fetched and measured → this design written. The rule's
purpose — that the number is not a diagnostic produced inside a running experiment to explain a
result away — is met: every figure below existed before a single line of the study limb was
specified.
| work | comparator | 7g | 12g | 15g | longest run | verdict | declared |
|---|---|---|---|---|---|---|---|
dost |
Garnett 1918, PG #600 | 33 | 6 | 2 | 16 | DEPENDENT? |
high |
daudet |
Keith Adams, PG #30442 | 1 | 0 | 0 | 7 | clean |
none |
sandmann |
Bealby 1885, PG #31377 | 3 | 0 | 0 | 8 | clean |
none |
wagahai |
— | — | — | — | — | no comparator reachable | none, on an absence |
materials/contamination.json; command line
python3 tools/dependence_check.py materials/contamination-cells.json.
What the dost figure does to this design, decided before the run. It does not disqualify
the item. Nothing here claims the lead is an independent third translator of anything, no published
rendering is an arm, and the primary asks only whether the narrator a source-only reader finds is the
narrator a target-only reader finds. What it does is weaken one confound for that item — the
English of dost is partly Garnett's hand, not the lead's — and that cuts against the run's own
"one hand across four works" worry rather than for it. dost is kept, flagged, and the primary is
recomputed without it as a pre-registered robustness check (§7 R2).
3. The instrument
A persona profile: nine integer scores, 0–6, plus one free sentence. Every profile is written from one text and nothing else. The same nine axes are used on both sides.
| family | axis | 0 | 6 |
|---|---|---|---|
| carrier | register |
plain, colloquial, everyday | elevated, literary, formal |
| carrier | rhythm |
short, even, level sentences | long, uneven, surging periods |
| carrier | diction-temperature |
cool, dry, clinical words | hot, charged, exclamatory words |
| carrier | distance |
the narrator stands right beside what he describes | he observes from far off |
| persona | warmth |
cold or hostile toward what he describes | affectionate |
| persona | irony |
says what he means, straight | his words continually undercut what they report |
| persona | intrusion |
never appears as an "I", never addresses anyone | constantly speaks in his own person or turns to the reader |
| persona | judgment |
passes no verdict on anything | constantly evaluates |
| persona | certainty |
hesitant, self-correcting, unsure of his own account | wholly certain |
Provenance of the axes, stated because the order of work makes it matter. The four carrier
axes are voice's own named carriers, taken verbatim from the entry. idiosyncrasy, the entry's
fifth, is deliberately excluded, on the entry's own record: it has failed to separate anything in
two runs, two languages and two disjoint seat sets (RS-20260805d §4, pinned at 7 in 8 of 8
profiles; RS-20260806-same-man §4, a range of 0.33 across five renderings). The five persona axes
are ordinary narratological properties. None was chosen by looking at the renderings — but the
renderings did exist when this list was written, because charter §3 rule 8 requires the logs frozen
first, and that ordering risk is a declared limit (§9).
The free sentence never enters a number. All matching is arithmetic on the nine integers.
This is the design's answer to the confound that killed ARM-voice-persona: a number cannot leak
plot, so a match cannot be produced by shared content. It is not an answer to the recognition
confound, which §6 measures instead.
4. Arms
| arm | what it is | who made it | n texts |
|---|---|---|---|
SOURCE |
the four source passages, read in the original | — | 4 |
CLOSE |
the lead's R04 renderings |
lead, $0 | 4 |
FLAT |
R14 matched-content flattening applied to two CLOSE renderings |
lead, $0 | 2 |
PANEL |
an unbriefed plain translation from the source by a non-panel model | mistralai/mistral-medium-3-5 |
2 |
The two works carrying FLAT and PANEL are daudet and wagahai — the alphabetically first
and last of the four slugs. The rule is mechanical and is fixed here so that no property of any
datum could have selected them.
FLAT uses R14 v0.1 exactly as written: it renders from the live English, never from the
source, and nothing may change what the text says. Its operators F1 (cadence levelling),
F2 (figure de-specification), F3 (repetition flattening), F4 (verb deadening), F5 (register
levelling) are applied to every sentence they reach. R14's own propositional-equivalence
requirement is checked by the recognition stage's seats reading both texts (§6c), not asserted.
5. Procedure
Every raw body is written to runs/ before anything is computed from it. Temperature 0. Axis order
in the prompt is counterbalanced by a hash of (work, seat, side) — half the calls get the list
reversed — so a scale-order effect averages out rather than aligning with a work.
| stage | what | calls |
|---|---|---|
| 0 | independent adversarial pre-run critic over this frozen design | 1 |
| 1 | SOURCE profiles — 4 works × 4 seats, source language only |
16 |
| 2 | CLOSE profiles — 4 × 4, English only |
16 |
| 3 | FLAT profiles — 2 × 4 |
8 |
| 4 | PANEL translations — 2 works, one non-panel hand, unbriefed |
2 |
| 5 | PANEL profiles — 2 × 4 |
8 |
| 6 | recognition probe — 4 CLOSE texts × 2 seats, asked to name work and author |
8 |
| total | 59 |
Seats. Profiling seats are P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3
x-ai/grok-4.5, P5 deepseek/deepseek-v4-pro (routing pinned per the S121/S122 record). Critic is
P4 moonshotai/kimi-k3, which profiles nothing. The PANEL translator is
mistralai/mistral-medium-3-5, non-panel, so that no profiling seat ever scores its own prose
(charter §5). Roles are the design's; slugs are logged as provenance.
Nobody judges quality anywhere in this run. Every seat is a describer of a narrator. No
sense is scored, no translation is ranked, and nothing here is licensed to say any rendering is
better than any other. Tier D is NOT PASSED and every sentence of the result will carry
provisional.
6. Analysis, fixed here
Centring. For each (seat, side), subtract that seat's mean per axis across its four works. This
removes seat-level offsets, which would otherwise make a seat match itself trivially. FLAT and
PANEL profiles are centred with that seat's CLOSE constants, because a two-work mean is not
comparable.
Matching. For a target profile from seat s on work i: the source consensus for work j is
the mean of the centred SOURCE vectors of the seats other than s (leave-one-seat-out, so a
seat never matches against itself). Predicted work = argmin Euclidean distance over the four
consensus vectors. A tie counts as a miss.
G1— the source-determinacy gate. Each seat'sSOURCEvector for work i, matched leave-one-seat-out against the other three seats' source consensus. 16 trials, chance 4.P1— the primary.CLOSEtarget profiles matched against the source consensus. 16 trials, chance 4.P2— flattening. Paired per (seat, work) over the twoFLATworks: is the distance from theFLATprofile to its own source consensus greater than from theCLOSEprofile? 8 pairs.P3— the independent hand (positive control). The same paired comparison forPANELagainstCLOSE, andPANEL's own match count. 8 pairs.S1— which family carries it.P1recomputed on the fourcarrieraxes alone and on the fivepersonaaxes alone.
Null. Exact, by exhaustive enumeration: permute the work-labels within each seat independently, all 24⁴ = 331,776 assignments, and count how many reach the observed number of correct matches or more. That P is what gets reported; the binomial figures below are only how the bars were set.
Registered bars. G1 fires at ≥ 8 of 16 (binomial P(X≥8 | p=¼) = 0.0271). P1 fires at
≥ 8 of 16. P2 and P3 claim only at ≥ 7 of 8 in the predicted direction (sign test
P = 0.0352); below that they are descriptive.
7. Failure criteria — declared before dispatch
F1— the gate. IfG1< 8 of 16,P1is withheld and the run's finding is the source-side null:voice's source–target relation has no measurably determinate source end, which is the empirical half of Venuti's charge, and no crossing claim may be made in either direction.F2— degenerate seat. Any seat whose nine centred axis values have pooled SD < 0.40 across its four profiles on a side is giving every text the same profile; its trials are reported separately and the primary is recomputed without it. If two or more seats are degenerate on a side, the primary is withheld.F3— the length confound. A one-dimensional matcher on token count alone is run over the same trials. If it reaches the observed count or more, the primary is withheld: the axes would then be carrying nothing the word count does not.F4— recognition. Stage 6 asks two seats to name the work and author of eachCLOSEtext. This does not gateP1— it cannot, since a profile written by a seat that recognises the work is still a profile of that work's narrator — but the recognition rate is reported beside every figure, and if it is at ceiling (8 of 8), the result page states in its headline that this design cannot separate carriage from recognition.FLATandPANELare the partial answer: both carry the same recognisable content, so a match driven by recognition alone should survive flattening.F5— non-return. A cell that does not yield nine parseable integers after two attempts is empty. If more than 4 of the 48 profile cells are empty, the primary is withheld.F6— inert axis. Any axis constant across all 16SOURCEprofiles is reported as inert and the primary is also recomputed without it, alongside the registered figure, never instead of it.R2— robustness, not a failure criterion.P1is recomputed withdostdropped (3 works, chance ⅓, 12 trials), because that item iscontamination: high.
What a null means here. If P1 fails with G1 passing, the finding is that a determinate source
narrator does not arrive in a close English rendering by this instrument — which would be the
strongest thing this project has said against voice as a scoreable source–target relation, and it
is as publishable as a pass. The arm said at birth that a null is completion.
8. Pre-flight budget
UTC day 2026-08-06 stands at $2.587897236 of $5.00 before this session; headroom $2.412102764.
Worst case is built from max_tokens, not from expected output (note (abc)): 57 calls at
max_tokens 1,200 and 2 translation calls at 2,500, priced at the most expensive panel rate
($7.50/M out, $2.00/M in) with a 2× routing margin on top of that, per config/models.md's
standing caution that routing alone can move a bill 4×.
57 × (1200 × 7.5 + 1800 × 2.0)/10⁶ + 2 × (2500 × 7.5 + 1200 × 2.0)/10⁶ = $0.72, doubled:
declared worst case $1.50, 62% of headroom. A stage that will not fit is dropped in the order
6 → 5+4 → 3, and the drop is reported.
9. Declared limits, written before the run
- The axes were fixed after the renderings were made, because charter §3 rule 8 requires the
translator's logs frozen before an evaluation is designed. The mitigation is provenance, not
procedure: four axes are
voice's own, five are standard, none was picked off a rendering. - All four
CLOSEtexts are one hand. That works againstP1, not for it. - The seats are panel models. Nothing here licenses a claim about a human reader.
- Recognition is not excluded, only measured (
F4). - A pass does not refute Venuti, who can hold that convergent readers share a projection. It removes only the strongest version of the charge, the one that says there is nothing there to converge on.
- Four works is four works. Chance is ¼ and the bar is set accordingly; no claim about persona in general follows from four narrators.
10. Amendments after the pre-run critic — applied before any profile call
Critic: nvidia/nemotron-3-ultra-550b-a55b (non-panel; it profiles nothing here),
NEEDS-AMENDMENT, 15 findings, 6 BLOCKING, $0.01652940, finish_reason: stop, full text in
critic.md.
Seat note, and note (bhf)'s thirteenth firing, recorded because it was avoidable. The critic seat
was first moonshotai/kimi-k3 (P4). It returned finish_reason: length with zero content
characters at max_tokens 16,000 — the whole cap spent on hidden reasoning — billing
$0.2414856 for nothing. That slug had already failed this exact role at S106 and note (bhf)
rule (iii) says to change the seat rather than the ceiling. The design chose it without reading the
note. The seat was changed, and the replacement returned a complete critique for $0.0165, one
fifteenth of the failure. The dead body is kept as runs/critic-kimi-length-DISCARDED.json and is
ledgered.
Accepted — five BLOCKING and three advisory
- A1 (BLOCKING 3) —
F3is fully specified, in two parts. (a) Single-axis matchers. The match is recomputed nine times, each on one axis alone, same centring, same leave-one-seat-out consensus, same tie rule, same exact null. If the best single-axis count is ≥P1's count, the primary is not withheld, but the result page must say that the nine-axis profile adds nothing over that one axis and name it. (b) The length matcher, specified. Tokenisation is whitespacestr.split(). Target size = word count of theCLOSEEnglish. Source size = word count of the source fordost,daudet,sandmann; forwagahai, whose script has no spaces, non-space characters ÷ 1.702, the ratio that item's own materials give (599 ÷ 352). That divisor is fitted to the item, which makes the surrogate as strong as it can be, which is the conservative direction for a withholding rule. Matching rule: argmin |target size − source size|; the matcher is seat-independent, so its trial count is 4 × its per-work correct count. If it reachesP1's count or more, the primary is withheld, exactly as first registered. - A2 (BLOCKING 4) — the recognition probe is specified. Seats
S1= P1openai/gpt-5.6-terraandS3= P3x-ai/grok-4.5. Prompt verbatim: system — "You are shown an English passage. Name the work it comes from and its author if you recognise them. If you do not recognise it, say so. Do not guess wildly."; user — theCLOSEtext between---rules, then "Reply with JSON only: {\"recognised\": true|false, \"work\": \"...\", \"author\": \"...\"}. Use false and empty strings if you do not recognise it." Scoring rule: a row counts as recognised iffrecognisedis true and the returnedauthorstring contains the true author's surname, case-insensitively (Dostoevsky / Достоевский, Daudet, Hoffmann, Sōseki or Natsume). Titles are not scored: they vary between translations. Every response is logged verbatim. - A3 (BLOCKING 5) —
R2gets a bar. 12 trials at chance ⅓: ≥ 8 of 12 (P(X≥8 | 12, ⅓) = 0.0188), with the exact permutation P reported beside it. - A4 (BLOCKING 6) — the
PANELtranslation prompt is fixed. System: "You are a translator. Translate the passage into English. Output the English translation and nothing else — no preamble, no notes, no title." User: the source passage alone. Temperature 0,max_tokens2,500. It is unbriefed in the sense that matters: it is told nothing about narrators, personas, register or this experiment. - A5 (advisory 9) —
F2is defined exactly. For each seat and side, the standard deviation of the 36 centred axis scores (4 works × 9 axes). Degenerate if SD < 0.40. - A6 (advisory 7) — accepted as a reporting duty. The result page states that 4 of the 9 axes
are
voice's own named carriers and 5 were added, and repeats §9 limit 1. - A7 (advisory 8) — accepted.
P2andP3are descriptive. Eight pairs at a 7-of-8 bar has little power, and neither may be headlined in either direction. - A8 (advisory 10) — accepted. The 4-work and 3-work
P1figures are reported side by side.
Rejected, with the reason — BLOCKING 1 and 2
The critic requires a single centring constant per seat across all eight profiles in place of the registered per-side centring, on the ground that per-side centring makes a seat that scores high on the source and low on the target look artificially distant. The algebra says the opposite, and the finding is declined. Write σ for a seat's mean source vector and τ for its mean target vector.
- Registered (per-side): distance is ‖(T_w − τ) − mean_j(S_j − σ)‖ = ‖(T_w − S̄_j) − (τ − σ)‖. The seat's side offset τ − σ is removed.
- Critic's (common constant γ): distance is ‖(T_w − γ) − mean_j(S_j − γ)‖ = ‖T_w − S̄_j‖. γ cancels exactly, so the proposed change does not centre the comparison at all — it leaves the side offset τ − σ inside every distance, where it shifts the four candidate distances unequally and can move the argmin.
The registered scheme is the one that does what the finding asks for. BLOCKING 2 falls with 1.
What the finding does earn is a robustness figure, which is added: P1 is also reported under the
critic's common centring, so that the choice is visible rather than argued.
11. Amendment A9 — declared mid-run, before the stage it changes
S2 (google/gemini-3.6-flash) returned finish_reason: length on all four SOURCE calls at
max_tokens 1,200, with the JSON truncated mid-object after three or four of the nine scores. The
bodies billed $0.0417 and carry no usable profile. Note (b) forbids accepting a truncated
body; note (bhq) prescribes raising capacity and re-dispatching as a declared amendment, and says
to size the cap from the seat's measured reasoning appetite rather than from the answer.
Measured, off the stored bodies: completion_tokens_details.reasoning_tokens = 1,152 / 1,104 /
1,124 / 1,113 against a total completion allowance of 1,196. So the cap was consumed by hidden
reasoning and the answer never fitted.
A9: max_tokens is raised to 3,000 for S2 alone (1,152 measured + ~250 for the answer,
doubled), written into run.py as SEAT_MAX_TOKENS so a resume picks it up. The four truncated
bodies are moved to runs/discarded/ and re-dispatched with a byte-identical payload. No other
seat's cap changes: S1, S3 and S4 returned stop on 12 of 12.
The worst case is rebuilt from the raised cap (note (abc)). S2's 12 remaining profile calls at
3,000 and the other three seats' 24 at 1,200, priced at $7.50/M out with a 2× routing margin, plus
the two translation calls and eight recognition calls: $0.95 remaining against $0.3416 already
spent, so the session's declared worst case is revised from $1.50 to $1.70 — the increase is
entirely the two dead bodies (the kimi critic at $0.2415 and these four at $0.0417) and the raised
S2 ceiling.
Note (bhf) has now fired twice in this one session, on two different slugs, in two different roles, and both were foreseeable from the note itself. Recorded in the result page's limits rather than softened.