Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260803e-purpose-index/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260803e-purpose-index
statusfrozen
created2026-08-03
updated2026-08-03
sensesnaturalness, perceived-source-carriage, style-correspondence
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-fluency-record.md, wiki/base/anchors/A-english-tale-register/A-english-tale-register.md, wiki/base/anchors/A-mchugh-presence/A-mchugh-presence.md, wiki/base/sources/S-arnold-newman-homer.md, wiki/goodness-senses.md, workshop/translations/kachikachiyama/R10t-v1/translation.md, workshop/translations/kachikachiyama/R10u-v1/translation.md, workshop/translations/kachikachiyama/propositions.md, config/models.md, config/budget.md

E-20260803e — is "natural" one judgment or two? Corpus-indexed against purpose-indexed markedness

ARM-fluency-record step 3 of 3 (T4). Frozen before any dispatch.

0. The subject-rule sentence (continue-prompt.md §4.5), written before the unit was designed

What does this unit teach about translating literature or evaluating translations? — It teaches whether a reader's sense of "this English reads naturally" is one judgment or two: whether the same sentence can be at once inside the repertoire a reader brings to the kind of writing it is, and marked against present-day standard English. If it can, then a translator working in a genre-marked register is not failing naturalness but satisfying a different index of it, and an evaluation that scores such a translation against a period corpus is scoring the wrong thing.

That sentence is about readers and translations. The consequence for this project's sense list follows from the finding; it is not the question.

1. What is asked, and why this arm owes it

wiki/goodness-senses.md, naturalness, since S097: Arnold's 1861 established-possession test — "whether a diction is antiquated for that particular purpose for which it is employed" (S-arnold-newman-homer §2) — is recorded as a candidate refinement and NOT adopted, because "adopting it needs a measurement ARM-fluency-record is constituted to make and has not made." This is that measurement.

naturalness currently indexes markedness to three period corpora. Arnold's test indexes it to purpose and genre. The empirical question on which adoption turns is whether those two indices are actually different — whether they can be brought to disagree, and in which directions.

2. Materials

The translation limb, frozen and committed before this design was written. A matched R10 pair on 楠山正雄「かちかち山」 (Aozora 000329, PD, 4,177 non-whitespace characters, complete):

Same source, same translator, same session; only the target catalogue varies (R10). Contamination none, measured on a pre-selection gate before either translation began: 0 shared 7-grams, longest common run 4 tokens against Ozaki 1908, the freely reachable published English rendering of this tale, which was extracted to disk and never read.

The item set, materials/items.json, 34 items (amended, §4a), six classes:

class n what it is provenance
P 8 a repertoire feature, attested ≥ 3 times in jacobs-1890-tales.txt (51,154 words) verbatim from R10t
M 8 a contemporary-standard feature, 0 attestations in the anchor corpus, unmarked today, in its own contemporary-register context verbatim from R10u
M2 4 four of the same contemporary features, in tale-register sentences — the context-independence control CONSTRUCTED into R10t sentences
N 6 a span on which the two renderings agree verbatim and which carries no oral-formulaic or genre marking verbatim; context from R10t ×3, R10u ×3
X 6 an archaism with 0 attestations in Jacobs AND 0 in Kipling — eftsoons, certes, withouten, anon, erelong, natheless CONSTRUCTED, substituted into R10t sentences
A 2 ungrammatical by agreement and morphology, not by word order CONSTRUCTED

Twenty-two of the thirty-four items are verbatim from a translation frozen before this design existed, and that is checked mechanically by analysis/verify.py. The twelve constructed items are declared as constructed here and on the result page. Class X is Arnold's own negative list and its kin: the words he says are not an established possession — and the anchor corpus, which is saturated with quoth (22) and whereupon (6), contains none of them.

The marked element of each item is delimited by «guillemets». Every item is 1–3 sentences.

3. The manipulation — one text, two questions

Nothing about the material changes between conditions. Only the question does. Each seat sees the same 34 items twice, in two independent API calls with no shared context (so there is no carryover by construction). Both prompts open with the identical material description — "These passages are all from an English translation of a traditional Japanese folk tale — the kind of story that is handed down and read aloud" — so the genre frame is not what varies (amendment A1). What varies is:

Both scales run in the same direction, so a per-item gap = CORPUS − PURPOSE is meaningful. The purpose prompt's illustrative sentence was removed by amendment A1: it said a feature could be old and still be an established possession, which is a licence to downscore archaism present in one condition and absent from the other, and P1 could have been produced by it alone. Arnold's bare question now does the work.

Every item is truthfully described in both conditions. Every excerpt is from a translation of a traditional folk tale. No seat is told anything false.

4. Seats and dispatch

Jurors P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5, P5 deepseek/deepseek-v4-pro — the three non-Anthropic seats used as coders at S097, and here again coders of a descriptive property, not judges of quality. No translation is evaluated for quality anywhere in this run; Tier D is NOT PASSED and nothing here needs it.

3 seats × 2 indices × 2 item orders (registered forward order and its exact reverse) = 12 calls, 34 items each. The order arm is a robustness control, not a replication for power. Each call is stateless. The frozen order is materials/order.json, generated before dispatch.

Pre-run critic: P4 moonshotai/kimi-k3, which grades nothing in this experiment, with reasoning: {"effort": "low"} on the first dispatch — note (b)'s standing amendment, the one S097 failed to apply on this exact seat at a cost of $0.260301.

4a. AMENDED after pre-run critic pass 1 — read §10 before §5

Pass 1 returned NEEDS-AMENDMENT, seven findings, four BLOCKING. All seven were accepted and are applied below and in §10. Nothing had been dispatched. The item set went from 30 items to 34 (a new class M2), two prompt sentences changed, four item spans were re-cut, three neutral controls and both attention items were replaced, and two criteria were repaired. §5 as it now stands is the amended version and is what was frozen for dispatch.

5. Registered predictions and criteria

Let gap(item) = mean CORPUS − mean PURPOSE over the 6 (seat × order) cells.

PRIMARY — the purpose index is a DISTINCT INDEX, not a leniency adjustment toward old words, iff all FOUR of these hold (P4 added by amendment A2):

CONTROLS — gates on interpretability. If any fails, the primary is withheld and reported as uninterpretable.

Registered secondary, reported as context and NOT gating: per-item hit rate — the proportion of class-P and class-M items whose individual gap has the predicted sign.

6. Result → option map, fixed before dispatch

outcome what the arm does
P1, P2, P3, P4 all met; controls clean Open a motion to adopt the established-possession test as a fourth register anchor on naturalness, indexed to purpose, with A-english-tale-register as its Tier 1 evidence. Arm closes resolved. Amendment A8: this branch adopts the test as an index of the frame the rater applies; the reader-repertoire reading is bounded by limit 8, and the motion's language may not outrun it.
P1, P2, P3 met but P4 NOT met Open a motion recording the effect as register-congruence, not reader-repertoire: a contemporary feature is marked when it clashes with the sentence it sits in, which is not Arnold's test. No fourth anchor. Arm closes resolved. (Branch added by amendment A2.)
P1 met, P3 met, P2 NOT met Open a motion recording it as a leniency modifier, not an index: markedness of archaism is scored against purpose, but the register anchors stay corpus-indexed and no fourth anchor is added. Arm closes resolved.
P1 met, P3 NOT met (whatever P2 and P4 do — stated explicitly, amendment A6c) Not adoptable, no motion. The test would be indistinguishable from "be lenient to old words", which is not what Arnold's test claims to be. Arm closes resolved with that written.
P1 NOT met Not adoptable, no motion. Arm closes resolved with that written.
any control fails Primary withheld, run reported as uninterpretable, arm closes resolved with the failure written.

In every branch the arm closes at 3 of its 3 declared sessions. It does not extend. This is recorded now so that no outcome can be read as licensing another Homer step (ARM-fluency-record, "Do not extend the arm to translate more Homer").

7. What this design cannot show, registered as limits before the numbers exist

  1. One genre, one direction. The purpose tested is traditional tale. Whether purpose-indexing behaves the same for epic, for legal prose, for a modern realist novel, is untested. Arnold's own case was epic verse and this run does not reach it.
  2. The raters are language models, and the question is about a reader's repertoire. They are used here as coders of a descriptive property, which RS-20260725-anchor-verification supports for factual/linguistic coding, and not as a reception record. No claim about human readers follows.
  3. Class X is constructed and its items are rare words. A seat may score them high because they are unfamiliar rather than because they are unpossessed-for-the-purpose. The design cannot separate those, and P3 is therefore a weak discriminator in the passing direction and a strong one in the failing direction.
  4. The lead wrote both renderings and assigned every class. Class assignment for P and M is anchored by mechanical attestation counts; class N is anchored by verbatim agreement between two independently-targeted renderings; class X and A are not anchored by anything but the lead's construction.
  5. A tale register is still a target-culture norm. Nothing here bears on whether rendering 昔話 into English nursery idiom is a good idea; A-english-tale-register says so in its own text.
  6. Authority cue and prompt asymmetry, residual after A1 (critic finding 7a). The purpose prompt still names Arnold and is still the longer of the two. Naming an authority shifts ratings on its own. A1 removed the substantive asymmetry — the illustrative licence and the differential genre frame — and this residual is declared rather than removed, because the test cannot be stated without stating whose test it is.
  7. Within-call anchoring, uncontrolled (critic finding 7b). Every call contains all 34 items, so the presence of eftsoons in the list recalibrates what whereupon looks like. Class composition is constant across the order arm, so the order control does not touch this. Every number here is relative to the item pool it was rated in.
  8. The memorisation pathway, and the demand pathway beside it (pass 1 finding 7c; pass 2 finding 1). A seat asked whether something is in a present-day reader's repertoire has no reader's repertoire; it may instead be pattern-matching to 19th-century tale collections in its training data — of which Jacobs 1890 is certainly one. On that reading P1 measures familiarity with the Jacobs-flavoured corpus, not an index. The design cannot exclude it, and P4 and the M2 class are the only parts of the run that bear on it at all. Pass 2 sharpened this and the sharpening is accepted: M2 controls the local register-clash route but not the demand route — a seat told "traditional tale" and asked about a reader's repertoire may score contemporary phrases as outside it wherever they sit. So a P1–P4 pass licenses "the two indices behave differently when a rater is asked to apply them", and not "English readers possess this repertoire".
  9. The shared genre frame can only attenuate P1, never create it (pass 2 finding 4, registered so a marginal result cannot be re-litigated after the fact). Both prompts now name the folk-tale purpose, so some frame will bleed into the CORPUS condition. Bleed makes the corpus condition more lenient to archaism, which lowers class-P corpus scores and compresses the P1 gap. For class M there is no archaism to be lenient toward, so P2 is roughly untouched. A P1 pass is therefore harder to obtain than the design intends, and a P1 near-miss is a weaker result than it looks. This is the trade amendment A1 deliberately made: a prompt asymmetry that could manufacture P1 was exchanged for a frame bleed that suppresses it.

10. Amendment record — pre-run critic pass 1

Seat P4 moonshotai/kimi-k3, effort: low on the first dispatch, $0.073764, finish_reason: stop. Verdict NEEDS-AMENDMENT, seven findings, four BLOCKING, ALL SEVEN ACCEPTED. Note (rr), forty-first consecutive session with an independent pre-run critic pass.

# finding severity amendment
1 The purpose prompt's sentence "a feature can be old and still be … an ESTABLISHED POSSESSION" is a licence to downscore archaism present in one condition and absent from the other. P1 could have been produced by it alone. BLOCKING A1 — sentence removed; the material description made identical in both prompts, so the genre frame no longer varies either
2 P2's reversal could be register-congruence (a modern phrase clashing with its own sentence) rather than reader-repertoire, and every M item sits in contemporary context, so the frame must be imported from the prompt. The design cannot separate the constructs. BLOCKING A2 — new class M2: the same four contemporary features constructed into tale-register sentences; new criterion P4; new branch in the option map; limit 8
3 Class N is not genre-neutral: thump, thump is onomatopoeia, one, two, three a counting formula, the tripled adjective an oral cumulative. Verbatim agreement anchors provenance, not neutrality. F1 was a contaminated gate in both directions. BLOCKING A3 — all three replaced with plain narrative description, still verbatim-shared
4 Item spans containing more than one feature (i01, i09); like it for as if is not unmarked today (i14); i04's historic present may score ~0 on corpus and compress P1 NON-BLOCKING A4 — i01 cut to «Once upon a time», i09 cut to «wrecked», i14 replaced; i04 kept and the compression noted as working against P1, i.e. conservative. Per-item predicted signs registered in items.json
5 The attention items were scrambled word order, and the same prompt teaches that inversion is repertoire — so F2 could fail asymmetrically, killing exactly the runs where the seat took the purpose frame most seriously BLOCKING A5 — both replaced with agreement/morphology errors no register licenses
6 (a) P3's |gap| ≤ 0.50 passes if both indices score X low; (b) F4 withholds on sign flips that are noise; (c) the map's P1 met, P2 met, P3 not met cell is implicit NON-BLOCKING A6a/b/c — P3 now requires both means ≥ 2.50; F4 applies only where forward
7 Unlisted interpretability risks: authority cue, within-call anchoring, the memorisation pathway NON-BLOCKING A7 — added as limits 6, 7, 8

Nothing was dispatched before these amendments. The amended design, item set and both assembled prompts went to a second critic pass before the scoring stage.

8. Verification

analysis/verify.py recomputes every reported number from the stored raw JSON bodies with a parser independent of the scorer, re-checks all 22 verbatim item provenances against the frozen translations, re-checks all attestation counts against the stored anchor corpora, re-sums the billed cost, and runs mutation tests each asserting that the bytes on disk changed and each restoring the mutated file (notes (bgu), (bhd)).

9. Budget

Declared worst case built from max_tokens (note (abc)), not from expected output:

stage calls max_tokens worst case
pre-run critic (P4, effort low) 1–2 10,000 $0.36
scoring P1 4 6,000 $0.20
scoring P3 4 6,000 $0.18
scoring P5 4 6,000 $0.20 (priced at 4× list for routing, note on config/models.md)
retry reserve (note (bfc)) — — $0.20
total $1.14

Against $2.84 of headroom remaining on UTC 2026-08-03 after this session's ratification gate ($0.04527425). A run that does not fit is split or deferred.

11. Amendment record — pre-run critic pass 2

Same seat (P4, effort: low), $0.102372, finish_reason: stop. Verdict OK-TO-RUN. It was shown pass 1's report and asked to check closure rather than re-report. All four of pass 1's BLOCKING findings CLOSED, checked against the assembled prompts and item strings rather than against the design's claims about them. Four new NON-BLOCKING findings, all four accepted:

# finding amendment
1 M2 controls the local register-clash route but not the demand route; the map's first row overclaims A8 — caveat added to the map's first row; limit 8 rewritten to bound what a pass licenses
2 The amendments created near-duplicate item pairs inside every call (i32/i18, i29/i07), which cue the manipulation and are untouched by the order arm A9 — all four M2 carriers rebuilt; i18, i23 and i24 re-carried. A mechanical check now asserts no two items share a 10-word run: 3 pairs before, 0 after
3 The M2 carriers were pastiche, and i30's cleft ("it is not helping me you would be") is Irish-English idiom, in no catalogue here A10 — every M2 carrier is now verbatim-traceable to R10t, and verify.py asserts it
4 With both prompts naming the purpose, frame bleed into the CORPUS condition is likely; its direction is asymmetric across the predictions A11 — registered as limit 9 before the numbers exist

Two critic passes, eleven findings, all eleven accepted. Note (rr) holds. Combined critic cost $0.176136 against the declared $0.36.