Repository path: workshop/experiments/E-20260801g-purpose-row/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260801g-purpose-row |
| status | frozen |
| created | 2026-08-01 |
| updated | 2026-08-01 |
| senses | purpose-fit, accuracy, naturalness, voice, style-correspondence, affect, literary-quality, cultural-mediation, consistency |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-typology-derivation.md, wiki/goodness-senses.md, workshop/regimes/R11-declared-purpose.md, workshop/experiments/E-20260801g-purpose-row/material/purpose-specs.md, workshop/experiments/E-20260727-log-decision-coding/codes.md, workshop/experiments/E-20260727-log-decision-coding/decisions.tsv, wiki/decisions/resolved/D-20260725-05-typology-under-slate-f.md, wiki/findings/results/RS-20260727-log-typology.md |
E-20260801g — is purpose-fit a row in the typology or a parameter on it?
ARM-typology-derivation step 1. Frozen before the pair's locus was selected and before a word of either arm was written. The purpose specifications (material/purpose-specs.md) were frozen before the source was read for translation, per R11 §1.
The subject-rule sentence (continue-prompt.md §4.5, wiki/tracks.md): this unit teaches whether serving a declared purpose is itself a way a literary translation can be good, or only a re-weighting of the other ways it can be good. That is a question about what "good" means in literary translation — the charter's named central object — not about the project's instruments.
1. The question, and why it is answerable now
wiki/goodness-senses.md lists purpose-fit as one of nine senses and defines it as "The meta-sense: a declared purpose re-weights all the other senses." Two things about that entry have stood unexamined for eighty-three sessions.
First, the project already has a ratified test for exactly this distinction, and has never applied it here. D-20260725-05 Q3 ruled — unanimously, review and vote agreeing for once — that "capability as a parameter" is not a row in the list, on the ground that it "answers 'under what conditions can this method be good?', not 'in what way is this translation good?' — a parameter of the framework, not a row here." purpose-fit's own definition answers the first question in its own words: a purpose is a condition, and what it does is re-weight. The criterion has been on the books since S022 and was never turned on the row that most obviously invites it.
Second, the one place the project reached for purpose-fit on evidence, it recorded a stretch. In E-20260727-log-decision-coding's frozen class→sense mapping, over 380 coded decisions from twenty-one translator's logs, purpose-fit appears in exactly two rows and is the sole home of none:
| class | senses | verdict |
|---|---|---|
| C7 register placement | naturalness, purpose-fit |
clean — "naturalness is already a measurable register variable; a declared purpose re-weights it" |
| C14 reader-knowledge management | cultural-mediation, purpose-fit |
stretch |
In both, purpose-fit rides alongside a sense that is doing the naming. That is what a parameter looks like in a coding table. It is not proof, because purpose has been an unstated constant across every artifact this project has ever produced — a competent general rendering, never declared, never varied. A variable held fixed cannot show what it does.
R11 (created for this unit) makes it a variable. That is the whole reason a translation limb is required here and could not be replaced by more reading.
2. Design
A matched R11 pair: one source unit, one translator, one session, two declared purposes — P1 the crib and P2 the reading edition, both drawn from purpose-fit's own definitional list of legitimately differing purposes ("a scholar's crib, a general reading edition, a performance script, a children's edition differ legitimately").
Every logged decision in either arm is coded twice:
- a purpose code (
A/X/N) perR11§5, assigned at the moment of writing; - a sense home: the id of any of the other eight senses whose definition text on
wiki/goodness-senses.mdnames the consideration the decision turned on, with the quoted clause recorded. If no sense's definition text reaches it → residue.
The primary quantity is the count of A-coded decisions whose sense home is residue.
2.1 What the two hypotheses predict
| prediction | |
|---|---|
H-parameter — purpose-fit is a condition that re-weights the other eight |
residue among A-coded decisions = 0. Purpose decided which of the eight to prefer at each site; it named nothing the eight do not. |
H-row — purpose-fit is a distinct way a translation can be good |
residue among A-coded decisions ≥ 1, and the residue is not of the four RS-20260727-log-typology residue kinds (C11 which source, C12 how much of it, C13 what the translator had read, C4b what the target's grammar compels — the first three being conditions of production and the fourth having been absorbed into accuracy by D-20260727-08). |
The null is registered in the direction against the session's own motion. This session intends to propose that purpose-fit is a parameter. Under H-parameter the prediction is exactly zero, which any single residue decision refutes. Confirming the motion therefore requires finding nothing, and finding nothing is the outcome a motivated coder would have to manufacture rather than merely fail to look for. That asymmetry is the only leak control available on a limb the lead both writes and codes, and it is stated as weaker than an independent coder.
2.2 Predictions that are not the primary, registered anyway
- D1. The two arms diverge most on clauses P1.2/P1.4 against P2.2/P2.5 — the reader-knowledge clauses — because that is where
E-20260727's C14 stretch already sat. - D2.
X-coded decisions (a clause bore and was overridden) occur in both arms. A purpose spec that is never overridden is a spec the translator was free to obey costlessly, which would make the pair a weaker instrument than it looks. - D3. At least one
A-coded decision in each arm has two or more sense homes. If purpose never touches more than one sense at a site, "re-weights all the other senses" is overstated even under H-parameter.
3. Controls
Three, all frozen here, all $0. The primary is void if either coding control fails.
3.1 PC — can the coding rule return residue at all? (must fire)
A rule that cannot return residue makes H-parameter unfalsifiable. Tested on frozen, externally-labelled material: the 17 rows of E-20260727-log-decision-coding/decisions.tsv coded C11 (source-text uncertainty, 12) or C12 (excerpt artefact, 5), which that experiment's frozen mapping puts at residue on the nine senses.
Pass: the rule returns residue for ≥ 80% (≥ 14 of 17). Fail → primary void.
3.2 NC — does the rule over-produce residue? (must not fire)
The mirror. Tested on the 25 rows coded C2 (source-internal repetition), which the same frozen mapping puts cleanly at style-correspondence.
Pass: the rule returns residue for ≤ 20% (≤ 5 of 25). Fail → primary void.
3.3 The controls are coded blind
analysis/controls.py emits the 42 control decision texts shuffled, with the class column stripped, to analysis/control_items.txt. They are coded from that file alone, into analysis/control_codes.tsv, and the script joins the labels back afterwards. The lead cannot see, while coding, which items are meant to come out residue.
The blind is real but partial and is stated as such: the lead wrote most of the underlying logs, and the labels — though assigned by an earlier session under a design blind to this hypothesis — were assigned by the same agent.
3.4 The contamination gate (note (bcd)) — runs before the locus is selected
Order per S050: a gate unit is translated first, measured against the whole comparator, and only then is the pair's locus chosen, beyond the gate unit's span.
- Source: Selma Lagerlöf, «En julgäst», Osynliga länkar (1894), Project Runeberg — 2,605 Swedish words.
- Comparator: Invisible Links, tr. Pauline Bancroft Flach (1899), Project Gutenberg #14273, "A Christmas Guest" — 2,873 English words, body never read; extraction by script, counts only printed (note (bdn)).
- Instrument:
tools/dependence_check.py, unmodified, against the whole comparator story.
V3 — discard rule, frozen: if the gate returns a longest common run ≥ 12 tokens or ≥ 1 shared 12-gram, the work is discarded as a limb, the session reports the gate result and nothing else from this design.
4. Void conditions
| id | condition | consequence |
|---|---|---|
| V1 | fewer than 10 A-coded decisions across the two arms |
underpowered; report the count, no verdict on the primary |
| V2 | the two arms' English is too similar for purpose to have been operationalised — normalised token-level edit distance < 0.15 between the two arms, computed by analysis/divergence.py |
void; the specs failed to produce a contrast and the primary means nothing |
| V3 | the contamination gate fires (§3.4) | the unit is discarded before the pair is written |
| V4 | PC or NC fails (§3.1, §3.2) | the coding rule is not fit to answer the primary; report the controls only |
V2 is the condition this design is most likely to fail, and it is registered because failing it is the informative outcome it looks least like: two purposes that far apart producing near-identical English would be evidence that purpose does very little, which bears on the same question from the other side.
5. What is not being measured, and why
No quality judgment about either arm is made anywhere in this design. Charter §5 forbids the lead judging its own translation, and nothing here needs it: the primary quantity is whether a sense's definition text reaches a decision, which is a question about the definition, not about the rendering. Which arm is better, whether either serves its purpose, and whether the crib is a good crib are all outside this unit and would need a jury this project does not yet have (Tier D is not passed).
This also rules out the design the question first suggests — score both arms under both purposes and see whether purpose-fit predicts anything the other eight do not. That design is correct and unavailable.
6. Carryover, declared
R11 §6. The second arm is written by a translator that has just rendered the same source. The two arms are not independent renderings; note (bhb) records the lead matching itself at up to 37 contiguous tokens across sessions, above its record against any published human translation. Within one session the carryover can only be larger.
Consequence for this design, stated before the run: the pair's similarity is uninformative — V2's threshold is a floor on divergence, never a ceiling — and any figure of the form "the arms share N tokens" will not be reported as evidence of anything. What carryover cannot manufacture is a divergence: a site where the two arms differ is a site where the purpose overrode the pull of the text already written. Arm order (P1 first) is recorded on both artifacts.
7. Procedure
- Freeze the purpose specs and opportunity list; run
material/spec_check.py. Done before the source was read for translation. - Freeze this design; independent pre-run critic pass; amendments recorded here before anything is translated.
- Translate the gate unit (the story's opening, ~180 Swedish words) under
R06; freeze; run the contamination gate against the whole comparator. Apply V3. - Select the pair's locus beyond the gate unit's span; translate arm
P1, freeze it in its own commit; then armP2, freeze it. - Code the controls blind (§3.3); then code the pair's decisions.
analysis/verify.pyrecomputes every reported number from the frozen artifacts, with mutation tests.
8. Cost
$0 for everything except the pre-run critic. All translation is lead work (charter §3, A4, never ledgered); both texts are free; all analysis is local computation. The critic is the design's only API call — reserved at $0.16 worst case from max_tokens 12,000, note (abc) — against a UTC-day headroom of $1.4966.
9. Amendments after the pre-run critic
qwen/qwen3.7-max, probed-but-not-selected (config/models.md), 2026-08-01, $0.03541475, 5,638 in / 6,124 out of which 4,554 reasoning, provider Alibaba, stop, 147.8 s. Verdict NEEDS-REDESIGN, six findings, four BLOCKING. All six accepted in substance; one prescribed remedy declined with a written reason; two accepted in a stronger form than prescribed. Nothing had been translated when this ran, so the amendments below are amendments to a design, not to a result. Full text: critic.txt.
The verdict was NEEDS-REDESIGN rather than NEEDS-AMENDMENT, and the changes below are on that scale: a co-primary added, the primary's licence cut, two of three controls replaced, a frozen spec clause rewritten, and two claims withdrawn.
A1 (finding 1, BLOCKING) — the primary is underpowered in the confirming direction, and now says so
"Residue=0 is nearly guaranteed by the elasticity of the eight senses, not by the ontological status of purpose-fit."
Accepted, and it is the finding that matters most. accuracy, naturalness and cultural-mediation are worded broadly enough to absorb almost any rendering decision, so a coder looking for some home will find one. §2.1's table implied that residue = 0 would support H-parameter. It does not, and cannot.
The prescribed remedy — forced choice of the single most specific sense — is declined, with a reason: forcing one label does not narrow an elastic label, it only hides which others were reachable. What is adopted instead is stronger.
- Primary A (coverage residue) keeps its form and loses its licence. Residue ≥ 1 refutes H-parameter. Residue = 0 is consistent with H-parameter and establishes nothing, and will be reported in those words. The test is deliberately asymmetric and the asymmetry is now stated where the prediction is made.
- Primary B — the trade test — is added, and it is the load-bearing one. H-parameter's actual content is that purpose chose between considerations the eight already name. That entails a two-sided structure at each divergence: senses X ≠ Y among the eight such that the crib's choice serves X at Y's expense and the reading edition's choice serves Y at X's expense. Each divergence is therefore classified: - trade — two distinct senses, opposed directions. Consistent with H-parameter. - one-sided — one arm's choice is named by a sense and the other arm's is named by nothing among the eight, so the second arm's choice has no home but the purpose itself. Evidence for H-row. - unnamed — neither arm's choice is named by any sense. Evidence for H-row.
Elasticity works against finding one-sided and unnamed sites, which is the direction that makes a positive finding meaningful. Registered: under H-parameter, ≥ 90% of divergences are trades. 3. Primary B does not depend on PC and is not voided by it.
A2 (finding 2, BLOCKING) — the four whole-text senses, and what a site-level log cannot reach
"Voice, affect, literary-quality, and consistency are macro-level, whole-text properties. A site-level decision log … lacks the capacity to surface these."
Accepted, and independently corroborated from inside the repo: in E-20260727-log-decision-coding's frozen class→sense mapping, over 380 decisions from twenty-one logs, voice, affect, literary-quality and consistency are the home of zero of the fourteen classes. The critic reached this from the design text alone, without seeing that table.
The prescribed remedy — drop the four macro-senses from the test — is declined: dropping them would assume the conclusion this session's dossier is being written to examine. Adopted instead, in a stronger form:
- Each arm's log opens with its whole-text stance decisions — the stance taken across the whole span, recorded and coded exactly like any site decision.
R11's purpose specs are global by construction, so these decisions are real, not manufactured; what changes is that they are now recorded rather than left implicit. This is the analogue ofR10§6's opportunity list. - The primary is reported twice: over all eight senses, and over the four site-level senses (
accuracy,naturalness,style-correspondence,cultural-mediation) alone. If a whole-text sense absorbs a divergence, the second figure shows it. - New registered prediction D4: the four whole-text senses appear as sense homes at a lower rate than the four site-level ones. If they appear at zero, Primary A and Primary B are reported as bearing on the site-level senses only, and that limit is a headline of the result rather than a caveat in it.
A3 (finding 3, BLOCKING) — the controls tested the wrong kind of object; all three replaced
"C11 and C12 are meta-decisions about the production process … does not prove it can identify rendering decisions as residue."
Accepted without qualification. §3.1's control was on the wrong object. The control set is rebuilt from the same frozen file, and the new gate is on rendering decisions:
| id | material | frozen label | required | if it fails |
|---|---|---|---|---|
| PC | the 13 rows coded C13 (prior-rendering pressure) — rendering decisions whose pressure no sense names | residue | ≥ 11 of 13 residue | Primary A void |
| NC | the 25 rows coded C2 (source-internal repetition) | clean → style-correspondence |
≤ 5 of 25 residue | Primary A void |
| AC | the 14 rows coded C4b (determinacy the target compels) | residue when coded, and non-residue now: D-20260727-08 absorbed them into accuracy's compelled-specification clause |
≥ 11 of 14 non-residue | the rule is reading the page as sketched rather than as ratified; Primary A void |
| (secondary, reported, not gating) | the 17 rows coded C11/C12 | residue | — | — |
AC is new and is the control the design lacked: a set of rendering decisions whose correct coding changed by ratification, so a rule that returns the old answer is detectably out of date. It costs nothing and it checks the one thing the other two cannot.
A PC failure is not a null result. If a rule applying these definitions cannot return residue on rendering decisions that an earlier design labelled residue, then coverage coding cannot discriminate between senses of this typology at all — which would retire the method RS-20260727-log-typology used and would bear directly on how that result's headline "80.6% clean" may be read. Registered as an outcome, not as a failure to report.
A4 (finding 4, not blocking) — the blind is withdrawn as a control
"It is not a blind; it is just a shuffled list. The agent can trivially reconstruct the labels from its memory of the generation process."
Accepted, and the claim is withdrawn rather than weakened. §3.3 said the blind was "real but partial". It is not real. The shuffle is kept because it removes ordering cues at no cost, but the controls are reported as lead-coded and unblinded, with no independence claimed.
The prescribed remedy — an independent agent instance coding the controls — is declined, with a reason. Charter §4 forbids reading panel agreement as validation, so an independent coder is evidence in the failing direction only; RS-20260727-log-typology §4 already ran exactly that check on exactly this scheme (83.3% / 75.0% against a 6.7% chance rate, the two coders agreeing with each other more than with the lead); and re-running it would be a measurement of the project's own instrument, which the subject rule (wiki/tracks.md) makes a gate at best and never a use of a session's budget. The honest move is to stop calling it a blind.
A5 (finding 5, BLOCKING) — P1.5 is rewritten; the vocabulary check is downgraded to what it is
"P1 is a caricature of a crib. Explicitly stating that the English need not be 'mistakable for something written in English' actively encourages ungrammatical or calque-heavy translation … This tests 'do I write in English or Swenglish?'"
Accepted. As frozen, P1.5 licensed broken English, which would have manufactured a trivial two-sided trade at every site and made Primary B vacuous. Rewritten before any translation, per R11 §1 (a clause is frozen against rewording once translating has begun; the pre-run critic sits before that line, which is why it sits there):
- was — "The English is not required to be read aloud, to be enjoyable, or to be mistakable for something written in English."
- now — "The English must be well-formed English: the reader cannot check their Swedish against a sentence that does not parse, and a broken crib teaches nothing. What the reader does not need from the page is the pleasure of reading it."
The second half of the finding is also accepted: the vocabulary check cannot stop paraphrase, and material/spec_check.py is a hygiene check on literal wording, not a guarantee of independence from the sense list. R11's section on the constraint is amended to say so. P1.1 is not rewritten, with a reason: "the reader is looking at the Swedish; the English exists to be mapped onto it" specifies the reader's activity — a crib's reader is comparing two pages, and any honest spec must say so — where P1.5 specified a licence.
A6 (finding 6, not blocking) — the V2 sentence is withdrawn
"You cannot infer the null hypothesis … from a failed manipulation check."
Accepted. §4's sentence — that near-identical arms "would be evidence that purpose does very little, which bears on the same question from the other side" — is withdrawn. V2 is a failed manipulation check and nothing else: it voids the unit and licenses no inference in any direction.
What the amendments cost and did not cost
No amendment moves a threshold after seeing data, because there is no data: the pair had not been translated and the locus had not been chosen. Two claims the design would otherwise have published are gone (A4's blind, A6's V2 reading), one licence is cut to nothing (A1), and the control set that was going to gate the primary would not have gated it (A3).