Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260802-voice-warrant/design/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260802-voice-warrant
statusfrozen
created2026-08-02
updated2026-08-02
sensesvoice, style-correspondence, accuracy
internal-judgment-onlytrue
provisionaltrue
trackT2
linkswiki/arms/ARM-typology-derivation.md, wiki/findings/sense-dossier.md, wiki/findings/results/RS-20260731d-sense-axes.md, wiki/goodness-senses.md, workshop/translations/harzreise/R06-v1/translation.md, config/models.md, config/budget.md, wiki/method-notes.md

E-20260802-voice-warrant — is voice distinguished by the extent of the feature, or by the extent of the warrant?

Frozen before dispatch. The translation it is measured on was frozen first, at cc6c394, before this design was written — the A4 freeze condition, satisfied by commit order rather than by assertion.

1. The question

wiki/goodness-senses.md distinguishes voice from style-correspondence in one parenthesis: "voice is global and cumulative; style-correspondence is local and formal."

Two independent measurements now sit against the first half of it.

Read together these two facts look like a demotion of the sense, and the dossier says explicitly that they are not one. This experiment tests a single explanation for both.

The hypothesis. voice is not distinguished from its neighbours by how far the feature in the source extends. It is distinguished by how much text is needed to warrant the choice. A voice decision is executed at one word and licensed by the whole passage. If that is right, then the page's words name the wrong property: what is "global and cumulative" is not the phenomenon but the evidence a reader needs in order to judge the rendering of it — and a site-level log will look empty of voice for the same reason a site-level rater declines to score voice sites as far-reaching.

Subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5), written before the design. This unit teaches what the criterion voice actually picks out when a translation is evaluated — whether it names an extent in the text or an extent of warrant — and it does so by translating a passage whose whole difficulty is voice and putting the sites of that difficulty to independent readers. The claim is about evaluating translations, not about the project's instruments; the reliability figures the run will also produce are limits, not the finding.

2. What is NOT asked

No quality judgement about any translation is made or elicited by anyone. Every question put to a rater is how much text do you need to decide or how far does this feature reach — never whether a rendering is good. Tier D is NOT PASSED (RS-20260802-tierD-verdict), and nothing here needs it to have passed, because nothing here is a quality claim. The design's own conclusions are internal-judgment-only and provisional regardless.

3. Materials

3.1 The translation

T-harzreise-R06-v1 — Heinrich Heine, «Die Harzreise» (1826) ¶124, 421 German words → 466 English, regime R06 (lead single pass, source-only, no revision), with a 30-decision translator's log frozen at cc6c394 before this file existed.

The passage was chosen because its difficulty is voice and almost nothing else: no dialect, no orthographic or typographic play, three items of realia, and few marked formal features — but a stance that is built across a hundred words, named in one five-word sentence (The Brocken is a German), and then dismantled. This directly answers RS-20260731d §7 limit 1, which records that its own material was skaz, where every formal property is simultaneously a property of the person the text sounds like, so that a null there is weak.

3.2 The item set — 24 sites, three strata of eight

Each item gives the German phrase, the English rendering, and one or two alternatives that were genuinely live. The translator's stated reason is withheld from every item, because the reason is exactly what would give the answer away. The rater is given the whole German paragraph and the whole English rendering above the items, so that consulting more text is possible — a warrant question is meaningless if the wider text is not there to be consulted.

Strata, declared by the lead here and frozen:

stratum home sense items (log ids)
V voice D1, D5, D6, D9, D14, D19, D26, D27
A accuracy D3, D10, D12, D17, D18, D22, D23, D30
F style-correspondence D7, D11, D16, D20, D21, D25, D28, D29

Six logged decisions are not used and the reasons are stated so the selection is auditable: D8 (paragraphing) has no site to point at; D13 is a decision to do nothing and has no alternative to offer; D4 and D24 are cultural-mediation, a fourth stratum this design does not carry; D2 and D15 are minor and would have unbalanced the strata.

A / F / V is the contrast set, and V–F is the replication contrast: voice against style-correspondence is the pair RS-20260731d measured, and accuracy is added as the uncontroversially site-local floor that run did not have.

3.3 The contamination gate, and the candidate it discarded

Run before the locus was selected and before ¶124 was translated, on a separate gate text (¶118, T-harzreise-R06-gate-v1, 226 English words), against two published English renderings of the whole «Harzreise» plus five null cells from other works by the same two translators in the same two volumes (note (bgf): a floor built from text the lead did not write cannot be gamed):

shared 7-grams shared 12-grams longest run
Storr 1887, whole Harz section 1 0 7
Leland 1869, whole Harz section 0 0 6
five nulls 0 0 3 each

Both figures are reported because run length alone is a poor proxy and has failed as one three times (notes (bcd), (bez); RS-20260731g §4.2). Verdict none.

One candidate span was discarded by note (bdn)'s rule while this was being set up. ¶45 was the first choice; a grep over the comparator files during material triage printed three lines of Storr's and two of Leland's English of that paragraph into the session transcript. Nothing had been translated from it. It was dropped, ¶124 was chosen instead, and no comparator prose for ¶124 has been seen. This is (bdn) firing for the fourth recorded time — you check a comparator by looking at its edges, and the edge you look at is chosen by the thing that made the work interesting — and the cost of it firing was one candidate paragraph.

4. Procedure

4.1 Stage 0 — pre-run critic, and an independent second home-assignment

Two calls, both before any rater sees anything, both to seats that do not rate in this run:

This is the control RS-20260731d §7 limit 3 says it did not have. There, "the lead wrote the items, the axes and the control declarations", the critic endorsed all eight controls sight-unseen, and nothing was excluded, so the control set was the lead's set with a second opinion on it. Here the primary is computed on the agreed subset only — items where the lead, P5 and P4 all name the same home — and the full 24 is reported as a secondary. Exclusion is the point; an endorsement that cannot exclude is not a control (the S070 move that S072 did not make).

4.2 Stage 1 — the two axes, counterbalanced

Three non-Anthropic rater seats, temperature 0, the same 24 items in two question shapes, one shape per call, item order reshuffled and item ids reassigned between shapes so that no rater can match an item to its own earlier answer:

The distinction is registered in one line: WARRANT is about the judgement, EXTENT is about the phenomenon. EXTENT is RS-20260731d's REACH re-asked in substance, so that this run contains its own replication of the axis that failed there.

Order is counterbalanced across seats, which RS-20260731d §7 limit 2 records that it was not — there, "CAT ran first for all three raters and may have anchored GRAD":

seat slug first call second call
P1 openai/gpt-5.6-terra WARRANT EXTENT
P2 google/gemini-3.6-flash EXTENT WARRANT
P3 x-ai/grok-4.5 EXTENT WARRANT

Three seats cannot be balanced 2–2; the split is 1 / 2 and is stated rather than smoothed. The primary is a within-seat contrast between strata, so order enters as a between-seat covariate and per-seat figures are reported.

4.3 Nulls, registered before dispatch

4.4 Reporting rules, registered

5. Registered predictions

prediction why it is falsifiable
P1 WARRANT(V) − WARRANT(A) ≥ +1.00 of 4, on the agreed subset the core claim; if voice sites need no more text than accuracy sites, the hypothesis is simply wrong
P2 the dissociation: (WARRANT(V) − WARRANT(A)) − (EXTENT(V) − EXTENT(A)) ≥ +0.50 if both axes move together, "warrant" is "extent" renamed and nothing is added to RS-20260731d
P3 WARRANT(V) − WARRANT(F) ≥ +0.75, against RS-20260731d's EXTENT-analogue gap of 1.000 on the same pair the replication contrast; a smaller gap than the axis it is meant to beat is a failure
P4 |Spearman(WARRANT, EXTENT)| < 0.60 pooled two questions that correlate above this are one question asked twice
P5 N2 holds: |ρ(length, WARRANT)| < 0.50 and the same for EXTENT the length confound RS-20260731d found at +0.35 on REACH
P6 N1 does not reach rater↔rater agreement, with the margin quoted note (bfx)
P7 ≥ 16 of 24 items survive the three-way home agreement of §4.1 if the strata cannot be reproduced by two independent seats they are the lead's opinion and the primary is withheld
P8 per-rater realised vocabulary reported before any mean note (bfq)

Registered direction of the interesting failure. If P1 holds and P2 fails, the honest report is that voice sites need more text and reach further, which supports wiki/goodness-senses.md as written and refutes this design's reason for existing. That outcome is recorded as a confirmation of the page, not as a partial success.

6. Failure criteria — what voids the run

An honest null is a first-class result. If P1 fails the finding is that voice sites are not warrant-hungrier than accuracy sites, the hypothesis is dead, and wiki/goodness-senses.md's parenthesis keeps whatever standing RS-20260731d left it.

7. Pre-flight cost

Worst case built from max_tokens × attempts × slugs (note (abc), as sharpened at S079), priced at the worst plausible provider for each seat (note: the S022 routing caution — a list price can be exceeded ~4× by routing alone):

call seat max_tokens attempts worst case
critic P5 deepseek/deepseek-v4-pro 8,000 2 $0.079
home-assign P4 moonshotai/kimi-k3 4,000 2 $0.168
WARRANT + EXTENT P1 openai/gpt-5.6-terra 4,000 2 each $0.145
WARRANT + EXTENT P2 google/gemini-3.6-flash 4,000 2 each $0.150
WARRANT + EXTENT P3 x-ai/grok-4.5 4,000 2 each $0.136

Declared worst case $0.70, against $4.486 of headroom on UTC day 2026-08-02 (one session, S086, $0.513968989 of $5.00). Note (bgk) — a slug may bill past max_tokens — is not priced in, because the headroom is six times the declared ceiling and the note's remedy (pricing the reasoning-inflated figure) is for tight-headroom runs; the after-the-fact check against the cap is still run.

8. Verification

analysis/verify.py recomputes every number that reaches the result page from the stored bodies, with mutation tests each asserting that the mutation changes bytes on disk (note (bgu)) and that every file the mutated run can write is snapshotted and restored (note (bhd)). tools/dependence_check.py, tools/build_index.py and tools/check_balance.py are used unmodified.