Repository path: workshop/experiments/E-20260809c-generic-voice/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260809c-generic-voice |
| status | frozen |
| created | 2026-08-09 |
| updated | 2026-08-09 |
| track | T2 |
| senses | voice, naturalness, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-voice-generic.md, wiki/goodness-senses.md, wiki/findings/results/RS-20260808d-carriage-decoupled.md, wiki/findings/results/RS-20260807f-carriage-or-strangeness.md, workshop/translations/ruska/R04-v1/translation.md, workshop/translations/ruska/R06-v1/translation.md, config/models.md |
E-20260809c — does genericising a narrator move voice and leave naturalness alone?
ARM-voice-generic step 1. Frozen before dispatch. Nothing below was written after a number
existed. The translation and its log were frozen first, at 51d2fcc, before this file was begun.
1. The question
wiki/goodness-senses.md, voice: "Distinguish from naturalness: a natural voice can be a
generic voice." Never measured. S135's reachability survey on the style-correspondence entry
names the operator: "an operator that genericises a narrator while holding register would move one
and not the other."
One direction is not enough. If voice were simply a noisier reading of naturalness, a
genericising operator that happened to cost a little fluency would produce a voice drop too, and a
single-direction design could not tell that from separation. So this run builds both operators
and tests the interaction:
GENERICgenericises the narrator's manner and should costvoiceand notnaturalness.AWKWARDroughens the English at sites that are not voice carriers and should costnaturalnessand notvoice.
A double dissociation is also the only form immune to the obvious objection about extent — that an arm which changes more text moves more scales. The primaries are differences between two senses measured on the same arm, so an account in which editing simply moves everything predicts both primaries at zero.
What this unit teaches about translating literature (subject rule, wiki/tracks.md): whether the
thing a translator is doing when they try to keep a narrator — as against keeping the propositions
and writing well — is visible to a reader as a separate property at all. The alternative outcome,
that it is not, is the same finding RS-20260807f reached about perceived-source-carriage, and it
would say that this project's controlled typology has two labels for one reader impression.
2. Materials
- Source. Jan Neruda, «O měkkém srdci paní Rusky», Povídky malostranské, 2nd edn (Praha:
Valečka, 1885), pp. 107–112; first published 1875. Public domain. 630 Czech words in seven
segments (
workshop/translations/ruska/source-cs-segments.txt), the story's opening five paragraphs, all of it narration. Single witness, uncollated — declared on the translation artifact. The project's first Czech, its first West Slavic source, its eighteenth source language. - Why this passage. The narrator is present in the first person twice (
Myslím,Měl jsem), disclaims his own gossip twice (prý,prý), announces that a thing is nobody's business and then tells it, keeps two sentences with no verb, and runs sixty words of an imagined scene in the present tense. Those are manner, not content, which is what makes the genericising operator constructible at all. CLOSE=T-ruska-R04-v1, frozen at51d2fccbefore this design existed, with its translator's log (D1–D23).- The three derived arms are built and asserted by
materials/arms.py, which runs 6 families of mechanical check and refuses to export an arm that fails one.
The four arms
| arm | what it is | words | token edit distance from CLOSE |
|---|---|---|---|
CLOSE |
T-ruska-R04-v1, the lead's close rendering |
794 | — |
GENERIC |
CLOSE with the narrator's manner genericised at 34 declared sites — rhythm, diction temperature, idiosyncrasy of phrasing, distance. Register held, propositions held, imagery held |
852 | 143 |
AWKWARD |
CLOSE with 21 declared wooden but grammatical edits — nominalisation, agentive passive, relative padding, periphrasis — at sites that are not voice carriers |
861 | 92 |
WRONG |
CLOSE with five planted propositional errors and no formal change |
794 | 5 |
What holds the two operators apart, mechanically. arms.py asserts that ten named CLOSE
strings — the inverted ledger name, the three-verb asyndeton, both verbless sentences, the
present-tense soup, the gnomic always say, the fronted object with its two they says, the
direct-speech punchline, the asyndetic tricolon — are absent from GENERIC and present verbatim in
AWKWARD. GENERIC exists to move exactly those; AWKWARD may not touch one.
What holds content, mechanically. Two checks, not one. (i) Class C: 41 declared strings —
every image, every evaluative word, every named entity, every number — must appear byte-identical
in CLOSE, GENERIC and AWKWARD. (ii) Content-word conservation: every content word of
CLOSE must survive into each derived arm on a 4-character stem match, or be listed in
DECLARED_LOSSES with the operator kind that removed it. There are 20 declared losses and every one
is a lexical move belonging to a named operator. This is the analogue of RS-20260808d's Class A /
Class B split, made stricter: that run held its image-bearing strings identical, this one holds those
and audits the whole content vocabulary.
Non-calque discipline. RS-20260807f's pre-run critic caught six of eighteen unlicensed edits
that were calques of the source at the site, and RS-20260808d banned fronted inversion outright
because Danish is V2. Czech is article-less, case-marked and orders constituents freely, so an
English article omission, a bare-noun subject or a fronted object would be a calque here. None of
the 21 AWKWARD edits is of those kinds — every one is a wooden habit of English officialese — and
each row of AWKWARD_SITES states what the Czech does at that site. Two rows say so explicitly:
a leather grocer's apron → an apron of leather for grocers moves away from the Czech, whose
shape is adjective + adjective + noun.
3. Panel and stages
Seats J1 = P1 (openai/gpt-5.6-terra), J2 = P2 (google/gemini-3.6-flash), J3 = P5
(deepseek/deepseek-v4-pro) — the S135/S141 seats, unchanged, so the naturalness figures sit on the
same panel as the runs they will be compared with. Critic and the parity control are non-panel.
The lead judges nothing (charter §5).
| stage | what | source shown | scales in the call |
|---|---|---|---|
| 0 | pre-run critic over this frozen design | — | — |
| 1 | blind naturalness, dispatched first, before any seat sees Czech (S134 amendment A10) |
no | naturalness alone |
| 2 | blind register placement — a separate call, after stage 1, so the fluency question is not primed by a formality question | no | formality alone |
| 3 | Czech competence screen, 5 items | yes | — |
| 4 | voice alone, Czech present, with two required quote fields |
yes | voice alone |
| C1 | independent propositional-parity control (non-panel, qwen/qwen3.7-max): are the paired texts saying the same things, and where do they differ? |
yes | — |
28 items (4 arms × 7 segments) per seat at each rating stage; 3 seats; 84 naturalness
cells + 84 register cells + 84 voice cells plus 15 competence cells. Item ids are opaque hashes,
order is seed-fixed and independently shuffled per seat, and arms are interleaved.
voice is scored alone. Usage rule 6 (S135) requires accuracy to be scored in its own call
where a formal comparison is load-bearing, because the same difference measured +0.19 alone and
+0.57 alongside three senses in one experiment. The same argument applies with more force here:
the whole question is whether voice and naturalness move together, and putting them in one call
would manufacture the correlation the run exists to measure. They are in different calls, in
different stages, and naturalness is answered before any seat has seen a word of Czech.
The two required quote fields at stage 4 are note (bky)'s remedy, made a first-class measure:
each voice rating must name (i) the English phrase most responsible for it and (ii) the Czech it
answers to. Gate G6 reads the first and gate G1b reads the second, and both are checked before
the primary.
4. Gates — declared here, none movable after a number exists
| id | statistic | bar | what it protects |
|---|---|---|---|
G0 |
token edit distance GENERIC : AWKWARD |
ratio ≤ 2.0 | extent parity. Met before dispatch: 143 : 92 = 1.55 |
G1a |
Czech gloss screen, per seat | ≥ 4 of 5 | that stage 4 means anything |
G1b |
stage-4 Czech quote is a substring of that item's Czech segment | ≥ 0.80 of items | that the seats read the Czech they were given, not their memory of it |
G2 |
|Δ register placement (CLOSE − GENERIC)|, blind, 7 seat-pooled segment means, 21 ratings per arm behind them — identical in both respects to the primary (note (bkv)) |
≤ 0.75 | the operator's register-hold clause. If genericising also shifted formality, a naturalness null could be two effects cancelling |
G3a |
independent parity: CLOSE/GENERIC pairs judged propositionally equivalent |
≥ 6 of 7 | content parity, outside the lead |
G3b |
independent parity: CLOSE/AWKWARD pairs judged equivalent |
≥ 6 of 7 | the same for the other operator |
G3c |
same control: planted WRONG errors named |
≥ 4 of 5 | that G3a/G3b is not a rubber stamp |
G4 |
Δ blind naturalness (CLOSE − AWKWARD) |
≥ 1.00 | AWKWARD is live. P2 is uninterpretable if the roughening did nothing |
G6 |
share of GENERIC stage-4 items whose quoted English phrase falls inside a GENERIC_SITES edit |
≥ 0.50, and above the 0.2887 chance baseline committed at §9 A3 |
note (bky): that the seats are rating the manipulation and not the events |
There is deliberately no gate requiring GENERIC to move voice. That is P1, and a run may not
gate its primary on itself.
5. Predictions, registered
Read on seat-pooled segment means, n = 7, with the 21-cell form reported alongside. Exact paired
permutation over sign flips, 2⁷ = 128 relabelings, minimum attainable P = 0.0078. Δnat is from
stage 1, Δvoice from stage 4; any constant offset between the two stages cancels inside each
difference, which is why the primaries are differences of differences.
P1— the forward dissociation.D_G= Δvoice(CLOSE−GENERIC) − Δnaturalness(CLOSE−GENERIC) ≥ +1.00, exact P ≤ 0.05. Holds ⇒ genericising the manner costs the person-sense and not the fluency-sense.P2— the reverse dissociation.D_A= Δnaturalness(CLOSE−AWKWARD) − Δvoice(CLOSE−AWKWARD) ≥ +1.00, exact P ≤ 0.05. Holds ⇒ roughening the English costs the fluency-sense and not the person-sense.P3— the parenthesis in its own words. "A natural voice can be a generic voice": Δnaturalness(CLOSE−GENERIC) ≤ +0.25 (amended at §9A4; as first written this was "≤ 0 or the interval contains 0", which was close to unfalsifiable). Registered as a directional claim and not as an equivalence test, becauseRS-20260808candRS-20260809bhave now both shown that an equivalence interval on seven segments cannot pass at thisn(note (bkm)); an equivalence claim here would be unfalsifiable-by-arithmetic and is not made.P4— the crossing.GENERICscores higher thanAWKWARDonnaturalnessin 7 of 7 segments and lower onvoicein 7 of 7 (amended at §9A5: 6 of 7 gives an exact sign-test P of 0.0625, which does not reach 0.05).P4requires both halves; the counts are reported whatever they are.P5— reported, not predicted. r(blindnaturalness,voice) over the 28 items. Ifvoicewerenaturalnessunder another name, this is where it shows:RS-20260807fmeasured −0.975 for the last sense that turned out to be one.
6. Failure criteria — declared, and not weakened after firing
F1.G1afails for ≥ 2 seats, orG1b< 0.80 ⇒ stage 4 is uninterpretable;P1,P2,P4WITHHELD.F2.G2fails ⇒ the genericising operator did not hold register;P1andP3WITHHELD,P2unaffected (it does not useGENERIC).F3.G3a< 6 of 7,G3b< 6 of 7, orG3c< 4 of 5 ⇒ the arms are not content-matched; everything in §5 is descriptive and no sentence enterswiki/goodness-senses.md.F4.G4fails ⇒P2andP4WITHHELD.F5.G6fails ⇒P1is reported as descriptive only and may not be written into the typology as a measured separation.F6. Any seat returning < 80% of its cells is dropped and the loss is declared; cell completeness is reported whatever it is.F7. IfP1holds andP2fails, the result is reported as a single dissociation, which is compatible withvoicebeing a noisiernaturalness, in those words, and the typology sentence says so.
7. Known limits, written before the run
- The lead wrote all four arms, chose the passage, and wrote the operator tables.
G3a/G3b/G3cput content parity outside the lead andG2puts the register-hold clause outside it; nothing puts the two operator tables themselves outside it. A reader who thinksGENERICalso removed something the lead did not classify as manner hasmaterials/arms.pyand the site tables to check, and no third party checked them before dispatch except the stage-0 critic. - One passage, one language pair, one narrator. Neruda's narrator is unusually overt. A result
here does not establish that
voiceseparates fromnaturalnesson a covert narrator, and the result page will not say it does. AWKWARDis a made object. Real unnatural translations are unnatural for reasons — calque, over-literalism, register error — and this arm is unnatural for no reason, exactly asRS-20260807f's andRS-20260808d'sSTRANGEarms were. Its nulls are about manufactured woodenness.- Note (bkz) does not bind and the reason is stated rather than assumed. That note requires a
non-lead arm whenever a finding is about what a regime does. This finding is about what two
operators applied to one text do to two scales; there is no regime in it, and all four arms are
derived from one rendering, so a lead-specific translating habit sits identically in every arm and
cannot produce a difference between them. What the note would bite on, and does: any sentence
claiming that translators who genericise lose
voice— no such sentence is registered here. contaminationissuspectedand unmeasured, with no reachable comparator: both English Povídky malostranské collections are in copyright. Declared on the translation artifact. The design does not turn on it (limit 4's reasoning).GENERICedits 18% ofCLOSE's tokens andAWKWARD12%.G0bounds the ratio and the interaction form makes extent non-explanatory for the primaries, but it remains true that the two operators are not the same size, and any figure compared across arms rather than across senses carries that.- Tier D is NOT PASSED. No jury verdict here carries evidential weight; every evaluative sentence
in the result is
provisionalandinternal-judgment-only.
8. Pre-flight cost estimate
Built from max_tokens, not from expected length (note (abc)).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
0 critic (nvidia/nemotron-3-ultra-550b-a55b, reasoning off per (bkw)) |
1 | 16,000 | $0.10 |
1 blind naturalness |
3 seats × 2 blocks | 3,000 | $0.11 |
| 2 blind register | 3 seats × 2 blocks | 3,000 | $0.11 |
| 3 Czech screen | 3 | 3,000 | $0.05 |
4 voice |
3 seats × 4 blocks | 6,000 | $0.40 |
C1 parity (qwen/qwen3.7-max) |
2 | 8,000 | $0.15 |
| retries / recovery headroom | — | — | $0.28 |
Declared worst case: $1.20. UTC day 2026-08-09 stands at $1.460842 of $5.00 before this session, so the headroom is $3.539158 and this run fits it three times over.
9. Amendments after the pre-run critic, before any rating call
Critic: nvidia/nemotron-3-ultra-550b-a55b, reasoning off per note (bkw), $0.0185328, one call,
runs/critic.txt. Verdict NEEDS-REDESIGN, on three findings marked BLOCKING. All three are
factually false, and they are overruled with the demonstration rather than with an argument.
Overruled — BLOCKING 1 and 2, "GENERIC drops the Class C string the shavings, because it
appears there as of the shavings." The check is a substring test and "the shavings" is a
substring of "of the shavings". python3 arms.py exits 0 on that string and did so before the
critic ran. The critic simulated the code instead of the semantics of in.
Overruled — BLOCKING 3, "AWKWARD edits the declared voice carrier a leather grocer's apron
(line 312)." That string is not in the carriers list, at line 312 or anywhere. It is in
AWKWARD_SITES and nowhere else, which is exactly where a non-carrier edit belongs.
Overruled — BLOCKING 4/verdict item 3, "carriers names a string GENERIC does not edit
(her face, her lips began to tremble, copious tears sprang out), and GENERIC has no S7 edit."
GENERIC_SITES has three S7 rows, one of them at precisely that string, and the assertion
c not in joined_gen passes for it.
Six findings are accepted, and one of the false ones pointed at a real weakness.
A1(from the false BLOCKING 1). The Class C check ran on the seven segments joined, so a string could in principle have been satisfied by a coincidental match in a different segment. It is now checked per segment, against the segment ofCLOSEthe string actually comes from. All 41 strings still pass.A2. A kind-disjointness assertion is added: theGENERICkind set{RH, DT, ID, DI}and theAWKWARDkind set{NOM, PAS, PAD, PER}must not intersect, and no kind outside those may appear.overlaps()additionally reports three spans both operators touch — S2every comer, S4appointed for four, S6and once they beat him, each aDTedit in one arm and aNOMedit in the other. Site overlap is declared, not forbidden: the design's claim is disjointness in kind, the arms are independent derivations fromCLOSE, and a span edited by two different kinds in two different arms is not a confound between them.A3(critic SERIOUS b, accepted).G6's bar of ≥ 0.33 was near chance. The chance baseline is now computed and committed before dispatch: 0.2887 — the share ofGENERIC's tokens lying inside a declared edit span,arms.py: edit_span_share().G6's bar is raised to ≥ 0.50 and must additionally exceed 0.2887.A4(critic SERIOUS c, accepted).P3as registered ("≤ 0, or its 90% interval contains 0") was close to unfalsifiable, which is the defect it was written to avoid.P3becomes a plain directional bar: Δnaturalness(CLOSE−GENERIC) ≤ +0.25. The interval is reported as description and decides nothing.A5(critic SERIOUS d, accepted — the arithmetic is right). At 6 of 7 an exact sign test gives P = 0.0625, which does not reach 0.05, soP4could have been called at a significance it does not have.P4now requires 7 of 7 on both halves. The observed counts are reported whatever they are.A6(critic SERIOUS a, overruled on substance, wording fixed). The critic readG2as having 21 observations against the primary's 7. Both are 7 segment means, each pooled over 3 seats, 21 ratings per arm behind them — the precisions are equal, which is what note (bkv) asks. The sentence in §4 is reworded to say both numbers so the ambiguity cannot recur.A7(critic MINOR, accepted). Two no-op rows removed fromDECLARED_LOSSESand eight dead single-character entries removed fromSTOP(the checker already skips tokens shorter than three characters).
The critic's power paragraph is accepted as a limit and not as an amendment. It is right that
P1's and P2's conjunction of ≥ 1.00 with exact P ≤ 0.05 is demanding at n = 7; it is wrong
that power is "near zero", because the sign-flip test's P depends on the consistency of the seven
differences and not on their size — seven same-signed differences give P = 0.0156 whatever their
magnitude. What is true, and now travels as limit 8: the ≥ 1.00 bar is the binding half of each
primary, and a real separation smaller than a scale point fails these predictions as registered.