Repository path: workshop/experiments/E-20260803e-purpose-index/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260803e-purpose-index |
| status | frozen |
| created | 2026-08-03 |
| updated | 2026-08-03 |
| senses | naturalness, perceived-source-carriage, style-correspondence |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-fluency-record.md, wiki/base/anchors/A-english-tale-register/A-english-tale-register.md, wiki/base/anchors/A-mchugh-presence/A-mchugh-presence.md, wiki/base/sources/S-arnold-newman-homer.md, wiki/goodness-senses.md, workshop/translations/kachikachiyama/R10t-v1/translation.md, workshop/translations/kachikachiyama/R10u-v1/translation.md, workshop/translations/kachikachiyama/propositions.md, config/models.md, config/budget.md |
E-20260803e — is "natural" one judgment or two? Corpus-indexed against purpose-indexed markedness
ARM-fluency-record step 3 of 3 (T4). Frozen before any dispatch.
0. The subject-rule sentence (continue-prompt.md §4.5), written before the unit was designed
What does this unit teach about translating literature or evaluating translations? — It teaches whether a reader's sense of "this English reads naturally" is one judgment or two: whether the same sentence can be at once inside the repertoire a reader brings to the kind of writing it is, and marked against present-day standard English. If it can, then a translator working in a genre-marked register is not failing naturalness but satisfying a different index of it, and an evaluation that scores such a translation against a period corpus is scoring the wrong thing.
That sentence is about readers and translations. The consequence for this project's sense list follows from the finding; it is not the question.
1. What is asked, and why this arm owes it
wiki/goodness-senses.md, naturalness, since S097: Arnold's 1861 established-possession test —
"whether a diction is antiquated for that particular purpose for which it is employed"
(S-arnold-newman-homer §2) — is recorded as a candidate refinement and NOT adopted, because
"adopting it needs a measurement ARM-fluency-record is constituted to make and has not made."
This is that measurement.
naturalness currently indexes markedness to three period corpora. Arnold's test indexes it to
purpose and genre. The empirical question on which adoption turns is whether those two indices
are actually different — whether they can be brought to disagree, and in which directions.
2. Materials
The translation limb, frozen and committed before this design was written. A matched R10 pair
on 楠山正雄「かちかち山」 (Aozora 000329, PD, 4,177 non-whitespace characters, complete):
T-kachikachiyama-R10t-v1— targetA-english-tale-register(purpose-indexed), 2,134 words.T-kachikachiyama-R10u-v1— targetA-mchugh-presence(corpus-indexed, unmarked/literary- contemporary), 1,734 words.
Same source, same translator, same session; only the target catalogue varies (R10). Contamination
none, measured on a pre-selection gate before either translation began: 0 shared 7-grams, longest
common run 4 tokens against Ozaki 1908, the freely reachable published English rendering of this
tale, which was extracted to disk and never read.
The item set, materials/items.json, 34 items (amended, §4a), six classes:
| class | n | what it is | provenance |
|---|---|---|---|
| P | 8 | a repertoire feature, attested ≥ 3 times in jacobs-1890-tales.txt (51,154 words) |
verbatim from R10t |
| M | 8 | a contemporary-standard feature, 0 attestations in the anchor corpus, unmarked today, in its own contemporary-register context | verbatim from R10u |
| M2 | 4 | four of the same contemporary features, in tale-register sentences — the context-independence control | CONSTRUCTED into R10t sentences |
| N | 6 | a span on which the two renderings agree verbatim and which carries no oral-formulaic or genre marking | verbatim; context from R10t ×3, R10u ×3 |
| X | 6 | an archaism with 0 attestations in Jacobs AND 0 in Kipling — eftsoons, certes, withouten, anon, erelong, natheless |
CONSTRUCTED, substituted into R10t sentences |
| A | 2 | ungrammatical by agreement and morphology, not by word order | CONSTRUCTED |
Twenty-two of the thirty-four items are verbatim from a translation frozen before this design
existed, and that is checked mechanically by analysis/verify.py. The twelve constructed items are
declared as constructed here and on the result page. Class X is Arnold's own negative list and its kin: the
words he says are not an established possession — and the anchor corpus, which is saturated with
quoth (22) and whereupon (6), contains none of them.
The marked element of each item is delimited by «guillemets». Every item is 1–3 sentences.
3. The manipulation — one text, two questions
Nothing about the material changes between conditions. Only the question does. Each seat sees the same 34 items twice, in two independent API calls with no shared context (so there is no carryover by construction). Both prompts open with the identical material description — "These passages are all from an English translation of a traditional Japanese folk tale — the kind of story that is handed down and read aloud" — so the genre frame is not what varies (amendment A1). What varies is:
- CORPUS index. "Against present-day unmarked literary English prose — the ordinary register of contemporary literary fiction, whatever the writing is for — how MARKED is the delimited element? 0 = wholly unmarked, 4 = maximally marked."
- PURPOSE index. "Matthew Arnold, writing in 1861 …, proposed this test: the question is 'whether a diction is antiquated for that particular purpose for which it is employed'. Apply that test to the purpose these passages are written for. Is the delimited element within the established possession of a present-day English reader FOR THAT PURPOSE? 0 = wholly within, 4 = wholly outside."
Both scales run in the same direction, so a per-item gap = CORPUS − PURPOSE is meaningful. The purpose prompt's illustrative sentence was removed by amendment A1: it said a feature could be old and still be an established possession, which is a licence to downscore archaism present in one condition and absent from the other, and P1 could have been produced by it alone. Arnold's bare question now does the work.
Every item is truthfully described in both conditions. Every excerpt is from a translation of a traditional folk tale. No seat is told anything false.
4. Seats and dispatch
Jurors P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5, P5 deepseek/deepseek-v4-pro — the
three non-Anthropic seats used as coders at S097, and here again coders of a descriptive property,
not judges of quality. No translation is evaluated for quality anywhere in this run; Tier D is NOT
PASSED and nothing here needs it.
3 seats × 2 indices × 2 item orders (registered forward order and its exact reverse) = 12 calls, 34 items each.
The order arm is a robustness control, not a replication for power. Each call is stateless. The
frozen order is materials/order.json, generated before dispatch.
Pre-run critic: P4 moonshotai/kimi-k3, which grades nothing in this experiment, with
reasoning: {"effort": "low"} on the first dispatch — note (b)'s standing amendment, the one
S097 failed to apply on this exact seat at a cost of $0.260301.
4a. AMENDED after pre-run critic pass 1 — read §10 before §5
Pass 1 returned NEEDS-AMENDMENT, seven findings, four BLOCKING. All seven were accepted and are
applied below and in §10. Nothing had been dispatched. The item set went from 30 items to 34
(a new class M2), two prompt sentences changed, four item spans were re-cut, three neutral
controls and both attention items were replaced, and two criteria were repaired. §5 as it now stands
is the amended version and is what was frozen for dispatch.
5. Registered predictions and criteria
Let gap(item) = mean CORPUS − mean PURPOSE over the 6 (seat × order) cells.
PRIMARY — the purpose index is a DISTINCT INDEX, not a leniency adjustment toward old words, iff all FOUR of these hold (P4 added by amendment A2):
- P1 — the expected direction. Class P mean gap ≥ +1.00. (Repertoire features are less marked when judged for the purpose.)
- P2 — THE REVERSAL, and the discriminating test. Class M mean gap ≤ −1.00 — i.e. contemporary-standard features are judged more marked under the purpose index than under the corpus index. A single scale cannot produce this; only a second index can.
- P3 — discrimination WITHIN archaism. Class X mean |gap| ≤ 0.50, with class X CORPUS
mean ≥ 2.50 AND class X PURPOSE mean ≥ 2.50. (Old words that are not in the repertoire stay
marked under both indices.) Amended A6a: as first written, |gap| ≤ 0.50 could be satisfied by
both indices scoring X low — a seat that knows
anonfrom Shakespeare and rates it 1 twice would have passed a criterion named "these words are not possessed". Both means must now be high. - P4 — context independence (new, amendment A2, critic finding 2). Class M2 mean gap ≤ −0.50 — the same four contemporary features as four of the M items, constructed into tale-register sentences. If M's reversal is driven by a contemporary feature clashing with its own contemporary surroundings, M2 should not reverse. P4 is required for the "distinct index" branch of the map; without it the run cannot separate outside the reader's repertoire for this purpose from incongruent with the register of the sentence it sits in.
CONTROLS — gates on interpretability. If any fails, the primary is withheld and reported as uninterpretable.
- F1 — no global scale shift. Class N mean |gap| ≤ 0.50. Amended A3: the three neutral items
the critic identified as oral-formulaic —
thump, thump,one, two, three, and a tripled-adjective span — are removed. Verbatim agreement between two independently-targeted renderings anchors provenance, not genre-neutrality, and two renderings agreeing on onomatopoeia does not make onomatopoeia neutral. All six N items are now plain narrative description carrying no oral-formulaic or genre marking, still verbatim-shared. - F2 — attention. Both class A items score ≥ 3.0 under both indices in ≥ 5 of 6
(seat × order) cells. Amended A5: the attention items were scrambled word order, and the same
prompt teaches the seat that inversion (
Up jumped the old man,said he) is repertoire — so a diligent purpose-index seat could read the scramble as extreme archaic inversion and score it low. The failure would have been asymmetric, more likely in exactly the runs where the seat took the purpose frame most seriously. Both items are now ungrammatical by agreement and morphology (the hare have took … and cutted,his legs was gave way under of him), which no register of English licenses. - F3 — completeness. Every call returns 34 parseable ratings. A body with fewer is a seat
failure in the runner, whatever
finish_reasonsays. - F4 — order robustness. The sign of each class-level gap (P, M, M2, X, N) is the same in the forward and reversed orders, for any class whose forward |gap| ≥ 0.50. Amended A6b: as first written, a class gap of +0.2 flipping to −0.1 — noise either side of zero — would have failed F4 and withheld the primary.
Registered secondary, reported as context and NOT gating: per-item hit rate — the proportion of class-P and class-M items whose individual gap has the predicted sign.
6. Result → option map, fixed before dispatch
| outcome | what the arm does |
|---|---|
| P1, P2, P3, P4 all met; controls clean | Open a motion to adopt the established-possession test as a fourth register anchor on naturalness, indexed to purpose, with A-english-tale-register as its Tier 1 evidence. Arm closes resolved. Amendment A8: this branch adopts the test as an index of the frame the rater applies; the reader-repertoire reading is bounded by limit 8, and the motion's language may not outrun it. |
| P1, P2, P3 met but P4 NOT met | Open a motion recording the effect as register-congruence, not reader-repertoire: a contemporary feature is marked when it clashes with the sentence it sits in, which is not Arnold's test. No fourth anchor. Arm closes resolved. (Branch added by amendment A2.) |
| P1 met, P3 met, P2 NOT met | Open a motion recording it as a leniency modifier, not an index: markedness of archaism is scored against purpose, but the register anchors stay corpus-indexed and no fourth anchor is added. Arm closes resolved. |
| P1 met, P3 NOT met (whatever P2 and P4 do — stated explicitly, amendment A6c) | Not adoptable, no motion. The test would be indistinguishable from "be lenient to old words", which is not what Arnold's test claims to be. Arm closes resolved with that written. |
| P1 NOT met | Not adoptable, no motion. Arm closes resolved with that written. |
| any control fails | Primary withheld, run reported as uninterpretable, arm closes resolved with the failure written. |
In every branch the arm closes at 3 of its 3 declared sessions. It does not extend. This is
recorded now so that no outcome can be read as licensing another Homer step (ARM-fluency-record,
"Do not extend the arm to translate more Homer").
7. What this design cannot show, registered as limits before the numbers exist
- One genre, one direction. The purpose tested is traditional tale. Whether purpose-indexing behaves the same for epic, for legal prose, for a modern realist novel, is untested. Arnold's own case was epic verse and this run does not reach it.
- The raters are language models, and the question is about a reader's repertoire. They are
used here as coders of a descriptive property, which
RS-20260725-anchor-verificationsupports for factual/linguistic coding, and not as a reception record. No claim about human readers follows. - Class X is constructed and its items are rare words. A seat may score them high because they are unfamiliar rather than because they are unpossessed-for-the-purpose. The design cannot separate those, and P3 is therefore a weak discriminator in the passing direction and a strong one in the failing direction.
- The lead wrote both renderings and assigned every class. Class assignment for P and M is anchored by mechanical attestation counts; class N is anchored by verbatim agreement between two independently-targeted renderings; class X and A are not anchored by anything but the lead's construction.
- A tale register is still a target-culture norm. Nothing here bears on whether rendering 昔話
into English nursery idiom is a good idea;
A-english-tale-registersays so in its own text. - Authority cue and prompt asymmetry, residual after A1 (critic finding 7a). The purpose prompt still names Arnold and is still the longer of the two. Naming an authority shifts ratings on its own. A1 removed the substantive asymmetry — the illustrative licence and the differential genre frame — and this residual is declared rather than removed, because the test cannot be stated without stating whose test it is.
- Within-call anchoring, uncontrolled (critic finding 7b). Every call contains all 34 items, so
the presence of
eftsoonsin the list recalibrates whatwhereuponlooks like. Class composition is constant across the order arm, so the order control does not touch this. Every number here is relative to the item pool it was rated in. - The memorisation pathway, and the demand pathway beside it (pass 1 finding 7c; pass 2 finding 1). A seat asked whether something is in a present-day reader's repertoire has no reader's repertoire; it may instead be pattern-matching to 19th-century tale collections in its training data — of which Jacobs 1890 is certainly one. On that reading P1 measures familiarity with the Jacobs-flavoured corpus, not an index. The design cannot exclude it, and P4 and the M2 class are the only parts of the run that bear on it at all. Pass 2 sharpened this and the sharpening is accepted: M2 controls the local register-clash route but not the demand route — a seat told "traditional tale" and asked about a reader's repertoire may score contemporary phrases as outside it wherever they sit. So a P1–P4 pass licenses "the two indices behave differently when a rater is asked to apply them", and not "English readers possess this repertoire".
- The shared genre frame can only attenuate P1, never create it (pass 2 finding 4, registered so a marginal result cannot be re-litigated after the fact). Both prompts now name the folk-tale purpose, so some frame will bleed into the CORPUS condition. Bleed makes the corpus condition more lenient to archaism, which lowers class-P corpus scores and compresses the P1 gap. For class M there is no archaism to be lenient toward, so P2 is roughly untouched. A P1 pass is therefore harder to obtain than the design intends, and a P1 near-miss is a weaker result than it looks. This is the trade amendment A1 deliberately made: a prompt asymmetry that could manufacture P1 was exchanged for a frame bleed that suppresses it.
10. Amendment record — pre-run critic pass 1
Seat P4 moonshotai/kimi-k3, effort: low on the first dispatch, $0.073764, finish_reason:
stop. Verdict NEEDS-AMENDMENT, seven findings, four BLOCKING, ALL SEVEN ACCEPTED. Note (rr),
forty-first consecutive session with an independent pre-run critic pass.
| # | finding | severity | amendment |
|---|---|---|---|
| 1 | The purpose prompt's sentence "a feature can be old and still be … an ESTABLISHED POSSESSION" is a licence to downscore archaism present in one condition and absent from the other. P1 could have been produced by it alone. | BLOCKING | A1 — sentence removed; the material description made identical in both prompts, so the genre frame no longer varies either |
| 2 | P2's reversal could be register-congruence (a modern phrase clashing with its own sentence) rather than reader-repertoire, and every M item sits in contemporary context, so the frame must be imported from the prompt. The design cannot separate the constructs. | BLOCKING | A2 — new class M2: the same four contemporary features constructed into tale-register sentences; new criterion P4; new branch in the option map; limit 8 |
| 3 | Class N is not genre-neutral: thump, thump is onomatopoeia, one, two, three a counting formula, the tripled adjective an oral cumulative. Verbatim agreement anchors provenance, not neutrality. F1 was a contaminated gate in both directions. |
BLOCKING | A3 — all three replaced with plain narrative description, still verbatim-shared |
| 4 | Item spans containing more than one feature (i01, i09); like it for as if is not unmarked today (i14); i04's historic present may score ~0 on corpus and compress P1 |
NON-BLOCKING | A4 — i01 cut to «Once upon a time», i09 cut to «wrecked», i14 replaced; i04 kept and the compression noted as working against P1, i.e. conservative. Per-item predicted signs registered in items.json |
| 5 | The attention items were scrambled word order, and the same prompt teaches that inversion is repertoire — so F2 could fail asymmetrically, killing exactly the runs where the seat took the purpose frame most seriously | BLOCKING | A5 — both replaced with agreement/morphology errors no register licenses |
| 6 | (a) P3's |gap| ≤ 0.50 passes if both indices score X low; (b) F4 withholds on sign flips that are noise; (c) the map's P1 met, P2 met, P3 not met cell is implicit |
NON-BLOCKING | A6a/b/c — P3 now requires both means ≥ 2.50; F4 applies only where forward |
| 7 | Unlisted interpretability risks: authority cue, within-call anchoring, the memorisation pathway | NON-BLOCKING | A7 — added as limits 6, 7, 8 |
Nothing was dispatched before these amendments. The amended design, item set and both assembled prompts went to a second critic pass before the scoring stage.
8. Verification
analysis/verify.py recomputes every reported number from the stored raw JSON bodies with a parser
independent of the scorer, re-checks all 22 verbatim item provenances against the frozen
translations, re-checks all attestation counts against the stored anchor corpora, re-sums the billed
cost, and runs mutation tests each asserting that the bytes on disk changed and each restoring the
mutated file (notes (bgu), (bhd)).
9. Budget
Declared worst case built from max_tokens (note (abc)), not from expected output:
| stage | calls | max_tokens | worst case |
|---|---|---|---|
| pre-run critic (P4, effort low) | 1–2 | 10,000 | $0.36 |
| scoring P1 | 4 | 6,000 | $0.20 |
| scoring P3 | 4 | 6,000 | $0.18 |
| scoring P5 | 4 | 6,000 | $0.20 (priced at 4× list for routing, note on config/models.md) |
| retry reserve (note (bfc)) | — | — | $0.20 |
| total | $1.14 |
Against $2.84 of headroom remaining on UTC 2026-08-03 after this session's ratification gate ($0.04527425). A run that does not fit is split or deferred.
11. Amendment record — pre-run critic pass 2
Same seat (P4, effort: low), $0.102372, finish_reason: stop. Verdict OK-TO-RUN. It was shown
pass 1's report and asked to check closure rather than re-report. All four of pass 1's BLOCKING
findings CLOSED, checked against the assembled prompts and item strings rather than against the
design's claims about them. Four new NON-BLOCKING findings, all four accepted:
| # | finding | amendment |
|---|---|---|
| 1 | M2 controls the local register-clash route but not the demand route; the map's first row overclaims | A8 — caveat added to the map's first row; limit 8 rewritten to bound what a pass licenses |
| 2 | The amendments created near-duplicate item pairs inside every call (i32/i18, i29/i07), which cue the manipulation and are untouched by the order arm | A9 — all four M2 carriers rebuilt; i18, i23 and i24 re-carried. A mechanical check now asserts no two items share a 10-word run: 3 pairs before, 0 after |
| 3 | The M2 carriers were pastiche, and i30's cleft ("it is not helping me you would be") is Irish-English idiom, in no catalogue here | A10 — every M2 carrier is now verbatim-traceable to R10t, and verify.py asserts it |
| 4 | With both prompts naming the purpose, frame bleed into the CORPUS condition is likely; its direction is asymmetric across the predictions | A11 — registered as limit 9 before the numbers exist |
Two critic passes, eleven findings, all eleven accepted. Note (rr) holds. Combined critic cost $0.176136 against the declared $0.36.