Repository path: workshop/experiments/E-20260802-voice-warrant/design/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260802-voice-warrant |
| status | frozen |
| created | 2026-08-02 |
| updated | 2026-08-02 |
| senses | voice, style-correspondence, accuracy |
| internal-judgment-only | true |
| provisional | true |
| track | T2 |
| links | wiki/arms/ARM-typology-derivation.md, wiki/findings/sense-dossier.md, wiki/findings/results/RS-20260731d-sense-axes.md, wiki/goodness-senses.md, workshop/translations/harzreise/R06-v1/translation.md, config/models.md, config/budget.md, wiki/method-notes.md |
E-20260802-voice-warrant — is voice distinguished by the extent of the feature, or by the extent of the warrant?
Frozen before dispatch. The translation it is measured on was frozen first, at
cc6c394, before this design was written — the A4 freeze condition, satisfied by
commit order rather than by assertion.
1. The question
wiki/goodness-senses.md distinguishes voice from style-correspondence in one
parenthesis: "voice is global and cumulative; style-correspondence is local and formal."
Two independent measurements now sit against the first half of it.
RS-20260731d-sense-axes(S072). Three raters unanimously agreed which of the two definitions applied to eight control sites, and then declined to score thevoice-pole sites as global: the formal contrast separated the poles by +2.917 of 4 (p = 0.0001) and the scope contrast by 1.000 (p = 0.027), one fifth the size. The registered positive control failed on exactly that half, which voided the run's seam measurement.wiki/findings/sense-dossier.md§0 (S084). Across two independent corpora — 380 logged translation decisions from twenty-one translator's logs in nine source languages, and 48 from two more —voiceis the home of 0 of 14 decision classes and 3 of 48 decisions. A translator's log records decisions, a decision happens at a site, andvoiceis almost never there.
Read together these two facts look like a demotion of the sense, and the dossier says explicitly that they are not one. This experiment tests a single explanation for both.
The hypothesis.
voiceis not distinguished from its neighbours by how far the feature in the source extends. It is distinguished by how much text is needed to warrant the choice. A voice decision is executed at one word and licensed by the whole passage. If that is right, then the page's words name the wrong property: what is "global and cumulative" is not the phenomenon but the evidence a reader needs in order to judge the rendering of it — and a site-level log will look empty ofvoicefor the same reason a site-level rater declines to scorevoicesites as far-reaching.
Subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5), written before the
design. This unit teaches what the criterion voice actually picks out when a translation
is evaluated — whether it names an extent in the text or an extent of warrant — and it does so
by translating a passage whose whole difficulty is voice and putting the sites of that
difficulty to independent readers. The claim is about evaluating translations, not about the
project's instruments; the reliability figures the run will also produce are limits, not the
finding.
2. What is NOT asked
No quality judgement about any translation is made or elicited by anyone. Every question
put to a rater is how much text do you need to decide or how far does this feature reach —
never whether a rendering is good. Tier D is NOT PASSED (RS-20260802-tierD-verdict), and
nothing here needs it to have passed, because nothing here is a quality claim. The design's
own conclusions are internal-judgment-only and provisional regardless.
3. Materials
3.1 The translation
T-harzreise-R06-v1 — Heinrich Heine, «Die Harzreise» (1826) ¶124, 421 German words → 466
English, regime R06 (lead single pass, source-only, no revision), with a 30-decision
translator's log frozen at cc6c394 before this file existed.
The passage was chosen because its difficulty is voice and almost nothing else: no
dialect, no orthographic or typographic play, three items of realia, and few marked formal
features — but a stance that is built across a hundred words, named in one five-word sentence
(The Brocken is a German), and then dismantled. This directly answers RS-20260731d §7
limit 1, which records that its own material was skaz, where every formal property is
simultaneously a property of the person the text sounds like, so that a null there is weak.
3.2 The item set — 24 sites, three strata of eight
Each item gives the German phrase, the English rendering, and one or two alternatives that were genuinely live. The translator's stated reason is withheld from every item, because the reason is exactly what would give the answer away. The rater is given the whole German paragraph and the whole English rendering above the items, so that consulting more text is possible — a warrant question is meaningless if the wider text is not there to be consulted.
Strata, declared by the lead here and frozen:
| stratum | home sense | items (log ids) |
|---|---|---|
| V | voice |
D1, D5, D6, D9, D14, D19, D26, D27 |
| A | accuracy |
D3, D10, D12, D17, D18, D22, D23, D30 |
| F | style-correspondence |
D7, D11, D16, D20, D21, D25, D28, D29 |
Six logged decisions are not used and the reasons are stated so the selection is auditable:
D8 (paragraphing) has no site to point at; D13 is a decision to do nothing and has no
alternative to offer; D4 and D24 are cultural-mediation, a fourth stratum this design
does not carry; D2 and D15 are minor and would have unbalanced the strata.
A / F / V is the contrast set, and V–F is the replication contrast: voice against
style-correspondence is the pair RS-20260731d measured, and accuracy is added as the
uncontroversially site-local floor that run did not have.
3.3 The contamination gate, and the candidate it discarded
Run before the locus was selected and before ¶124 was translated, on a separate gate text
(¶118, T-harzreise-R06-gate-v1, 226 English words), against two published English
renderings of the whole «Harzreise» plus five null cells from other works by the same
two translators in the same two volumes (note (bgf): a floor built from text the lead did not
write cannot be gamed):
| shared 7-grams | shared 12-grams | longest run | |
|---|---|---|---|
| Storr 1887, whole Harz section | 1 | 0 | 7 |
| Leland 1869, whole Harz section | 0 | 0 | 6 |
| five nulls | 0 | 0 | 3 each |
Both figures are reported because run length alone is a poor proxy and has failed as one three
times (notes (bcd), (bez); RS-20260731g §4.2). Verdict none.
One candidate span was discarded by note (bdn)'s rule while this was being set up. ¶45 was the first choice; a grep over the comparator files during material triage printed three lines of Storr's and two of Leland's English of that paragraph into the session transcript. Nothing had been translated from it. It was dropped, ¶124 was chosen instead, and no comparator prose for ¶124 has been seen. This is (bdn) firing for the fourth recorded time — you check a comparator by looking at its edges, and the edge you look at is chosen by the thing that made the work interesting — and the cost of it firing was one candidate paragraph.
4. Procedure
4.1 Stage 0 — pre-run critic, and an independent second home-assignment
Two calls, both before any rater sees anything, both to seats that do not rate in this run:
- Critic — P5
deepseek/deepseek-v4-pro. Given this design in full and the 24 items, it returns a verdict (ACCEPT/NEEDS-AMENDMENT/NEEDS-REDESIGN) with numbered findings, and independently assigns each of the 24 items a home sense from the seven inwiki/goodness-senses.md, without being shown the lead's strata. - Second assigner — P4
moonshotai/kimi-k3. Given the 24 items and the seven sense definitions verbatim, nothing else. Assigns a home sense to each. Not shown the design, the hypothesis, the strata, or the critic's answers.
This is the control RS-20260731d §7 limit 3 says it did not have. There, "the lead wrote
the items, the axes and the control declarations", the critic endorsed all eight controls
sight-unseen, and nothing was excluded, so the control set was the lead's set with a second
opinion on it. Here the primary is computed on the agreed subset only — items where the
lead, P5 and P4 all name the same home — and the full 24 is reported as a secondary. Exclusion
is the point; an endorsement that cannot exclude is not a control (the S070 move that S072 did
not make).
4.2 Stage 1 — the two axes, counterbalanced
Three non-Anthropic rater seats, temperature 0, the same 24 items in two question shapes, one shape per call, item order reshuffled and item ids reassigned between shapes so that no rater can match an item to its own earlier answer:
- WARRANT — Suppose you had to decide whether the choice made here is the right one for
this translation. How much of the text do you need in order to decide?
0the site alone ·1the sentence it sits in ·2the surrounding few sentences ·3the whole passage ·4more than the passage — the work, or the translator's whole approach. - EXTENT — Consider the feature of the German that creates the problem here. How far past
this site does that feature itself extend in the text?
0confined to the site ·1the sentence ·2the surrounding few sentences ·3the whole passage ·4beyond the passage — a property of the work as a whole.
The distinction is registered in one line: WARRANT is about the judgement, EXTENT is about
the phenomenon. EXTENT is RS-20260731d's REACH re-asked in substance, so that this run
contains its own replication of the axis that failed there.
Order is counterbalanced across seats, which RS-20260731d §7 limit 2 records that it was
not — there, "CAT ran first for all three raters and may have anchored GRAD":
| seat | slug | first call | second call |
|---|---|---|---|
| P1 | openai/gpt-5.6-terra |
WARRANT | EXTENT |
| P2 | google/gemini-3.6-flash |
EXTENT | WARRANT |
| P3 | x-ai/grok-4.5 |
EXTENT | WARRANT |
Three seats cannot be balanced 2–2; the split is 1 / 2 and is stated rather than smoothed. The primary is a within-seat contrast between strata, so order enters as a between-seat covariate and per-seat figures are reported.
4.3 Nulls, registered before dispatch
- N1 (leak). A frozen mechanical rule over the lead's own item prose — the number of words in the item's alternatives line plus the item's total character length, ranked — correlated against the pooled WARRANT ranking. Its Spearman must not reach the rater↔rater mean Spearman. The margin is quoted whatever the verdict, per note (bfx): a null that nearly reaches is reported as nearly reaching.
- N2 (length). Spearman(item character length, WARRANT) and (…, EXTENT), pooled.
|ρ| must be < 0.5 on both, the threshold
RS-20260731dused. - N3 (permutation). The stratum labels are shuffled 10,000 times with a fixed seed; the p-value of the observed V−A WARRANT gap is read off that distribution.
4.4 Reporting rules, registered
- Per-rater realised vocabulary on both 0–4 scales is printed before any agreement or mean is quoted (notes (bfq), (bdq)): a seat that uses two of five levels is evidence about the seat, and a pooled check cannot see it.
- Every option offered is reported with its take-up rate (note (ber)).
- The exclusion count from §4.1 is reported before the primary, not after.
5. Registered predictions
| prediction | why it is falsifiable | |
|---|---|---|
| P1 | WARRANT(V) − WARRANT(A) ≥ +1.00 of 4, on the agreed subset | the core claim; if voice sites need no more text than accuracy sites, the hypothesis is simply wrong |
| P2 | the dissociation: (WARRANT(V) − WARRANT(A)) − (EXTENT(V) − EXTENT(A)) ≥ +0.50 | if both axes move together, "warrant" is "extent" renamed and nothing is added to RS-20260731d |
| P3 | WARRANT(V) − WARRANT(F) ≥ +0.75, against RS-20260731d's EXTENT-analogue gap of 1.000 on the same pair |
the replication contrast; a smaller gap than the axis it is meant to beat is a failure |
| P4 | |Spearman(WARRANT, EXTENT)| < 0.60 pooled | two questions that correlate above this are one question asked twice |
| P5 | N2 holds: |ρ(length, WARRANT)| < 0.50 and the same for EXTENT | the length confound RS-20260731d found at +0.35 on REACH |
| P6 | N1 does not reach rater↔rater agreement, with the margin quoted | note (bfx) |
| P7 | ≥ 16 of 24 items survive the three-way home agreement of §4.1 | if the strata cannot be reproduced by two independent seats they are the lead's opinion and the primary is withheld |
| P8 | per-rater realised vocabulary reported before any mean | note (bfq) |
Registered direction of the interesting failure. If P1 holds and P2 fails, the honest
report is that voice sites need more text and reach further, which supports
wiki/goodness-senses.md as written and refutes this design's reason for existing. That
outcome is recorded as a confirmation of the page, not as a partial success.
6. Failure criteria — what voids the run
- F1. P5 fails (a length confound at |ρ| ≥ 0.5) → the affected axis is void and no gap on it is reported as a finding.
- F2. P6 fails (the leak null reaches the raters) → the WARRANT measurement is void, and the run reports that the raters were reproduced by a word count over the lead's prose.
- F3. P7 fails (< 16 items agreed) → the primary is not reported at all. Only the secondary over the lead's full 24 stands, explicitly labelled as the lead's own strata.
- F4. A seat returns fewer than 24 answer lines, or
finish_reason == "length"→ seat failure, re-dispatched once, then the seat is dropped and the run reports 2 raters. - F5. Any per-seat realised vocabulary with fewer than 3 of 5 levels used on an axis → that seat's contribution to that axis is reported separately and excluded from the pooled mean, because a seat operating a five-point scale as a two-point one is not measuring the scale (note (bfq)'s case, applied as a rule rather than as a caution).
An honest null is a first-class result. If P1 fails the finding is that voice sites are
not warrant-hungrier than accuracy sites, the hypothesis is dead, and
wiki/goodness-senses.md's parenthesis keeps whatever standing RS-20260731d left it.
7. Pre-flight cost
Worst case built from max_tokens × attempts × slugs (note (abc), as sharpened at S079),
priced at the worst plausible provider for each seat (note: the S022 routing caution — a
list price can be exceeded ~4× by routing alone):
| call | seat | max_tokens |
attempts | worst case |
|---|---|---|---|---|
| critic | P5 deepseek/deepseek-v4-pro |
8,000 | 2 | $0.079 |
| home-assign | P4 moonshotai/kimi-k3 |
4,000 | 2 | $0.168 |
| WARRANT + EXTENT | P1 openai/gpt-5.6-terra |
4,000 | 2 each | $0.145 |
| WARRANT + EXTENT | P2 google/gemini-3.6-flash |
4,000 | 2 each | $0.150 |
| WARRANT + EXTENT | P3 x-ai/grok-4.5 |
4,000 | 2 each | $0.136 |
Declared worst case $0.70, against $4.486 of headroom on UTC day 2026-08-02 (one
session, S086, $0.513968989 of $5.00). Note (bgk) — a slug may bill past max_tokens — is not
priced in, because the headroom is six times the declared ceiling and the note's remedy
(pricing the reasoning-inflated figure) is for tight-headroom runs; the after-the-fact check
against the cap is still run.
8. Verification
analysis/verify.py recomputes every number that reaches the result page from the stored
bodies, with mutation tests each asserting that the mutation changes bytes on disk (note
(bgu)) and that every file the mutated run can write is snapshotted and restored (note (bhd)).
tools/dependence_check.py, tools/build_index.py and tools/check_balance.py are used
unmodified.