Repository path: framework/closure.md · rendered 2026-09-09
Page metadata (front matter)
| type | ledger |
|---|---|
| id | framework-closure |
| status | active |
| created | 2026-07-28 |
| updated | 2026-08-02 |
| senses | accuracy, naturalness, style-correspondence, voice, cultural-mediation, consistency, purpose-fit |
| provisional | true |
| internal-judgment-only | true |
| links | framework/README.md, wiki/findings/results/RS-20260730g-nonlead-decisions.md, wiki/arms/ARM-nonlead-log.md, wiki/findings/results/RS-20260729d-decision-grain.md, wiki/arms/ARM-decision-grain.md, framework/traceability-inventory.md, framework/control-arm-spec.md, wiki/arms/ARM-framework.md, wiki/findings/results/RS-20260728i-coverage-independent.md, wiki/findings/results/RS-20260726e-framework-coverage.md, wiki/findings/results/RS-20260728c-length-matching.md, wiki/findings/results/RS-20260726d-tierD-heldout.md, PROJECT.md, config/models.md |
What a framework release is short of — the closure statement
What this is. ARM-framework step 4, and the arm's declared ending. The completion criterion set at S035, sixteen sessions before this page: either framework/v0.1/ exists, or the arm closes with a written statement of what evidence a release is short of — a statement the inventory can now make in numbers rather than in prose. framework/v0.1/ does not exist. This is the other ending.
On the status word. S035's criterion said the arm would close retired. It closes resolved, and the difference is not a softening: under CLAUDE.md's vocabulary retired means abandoned with a written reason and resolved means complete. The arm reached a completion criterion it declared for itself and produced the artifact that criterion names. S035 reached for retired because it read closure-without-release as a failure to release. It is not: it is the answer to the arm's own question, what does the project actually recommend doing, for which pairs, on what evidence? The answer is almost nothing, and here is the arithmetic.
Standing. internal-judgment-only and provisional. Every count below is recomputable from the inventory table and from RS-20260728i's stored outputs; every judgment about what a release should do is the lead's.
1. The counts
1.1 Fourteen candidates, by shape
| shape | n | which | standing |
|---|---|---|---|
| prescriptive about translating — do X rather than Y, addressed to a translator | 1 | C12 | inadmissible, on two independent grounds (§2) |
| prescriptive about method — addressed to the project | 4 | C10, C11, C13, C14 | sound, X2, verified, replicated — and not framework content |
| procedural, diagnostic, taxonomic or descriptive | 9 | C1–C9 less C12 | a vocabulary and a set of questions |
1.2 The same fourteen, by evidence class after Tier D's failure
| class | n | standing |
|---|---|---|
| X1a externally anchored, second-read | 6 (C1, C2, C3, C5, C7, C9) | intact |
| X1b externally anchored, single-reader | 3 (C4, C6, C8) | intact but unverified |
| X2 machine-measured | 4 (C10, C11, C13, C14) | intact |
| X3 panel-scored | 1 (C12) | INADMISSIBLE — config/models.md NOT CALIBRATED |
1.3 What a release could carry, addressed to a translator
Zero prescriptive recommendations. One candidate has that shape and it is inadmissible. Charter §3 requires recommendations "stated so a user (or an autonomous run) can follow them"; "here are eight things translators do" is not one.
1.4 The number this session added: how many candidates a real decision ever reaches
RS-20260728i: two independent readers, 63 logged translation decisions from two languages, 126 classifications.
DECIDES: 0 of 126. Not one classification says a candidate determines an answer — including C12, which was supplied with its text as written and no note that it is inadmissible.
⚠ CORRECTED 2026-07-30 (S068), and the correction is to the FORM of this number, not to its substance (
RS-20260730i-candidate-reach,ARM-candidate-reachstep 1). This section said the zero was "a fact about the candidates". It is two claims and only one of them is. A third language pair (PT→EN, 45 sites, mean 3.98 live renderings per site, three raters) was put to the panel twice: once with the options the translator actually had, once with two. A bearing candidate excludes 0.1395 options at four and 0.1538 at two — invariant. That same behaviour readsDECIDES= 0.000 at four options and 0.154 at two. So the magnitude is a fact about the candidates and the code is a fact about how long the translator's option list was: a translator keeping a two-option log would have published a non-zero rate from the same evidence base and the same candidate behaviour. What to report instead isk, the absolute number of renderings a candidate rules out — the only quantity that did not move when the list was halved. It is 0.14. The normalisedE = k/nis not invariant (0.0368 → 0.0769) and must not be used. The zero is not withdrawn and nothing here makes the evidence base look better. At two options the candidates still exclude a seventh of an option.⚠ TESTED AGAINST NON-LEAD CENSUSES 2026-07-31 (S073), and this time the correction goes the other way (
RS-20260731e-option-census,ARM-option-censusstep 1). The box above says theDECIDEScode is "a fact about how long the translator's option list was", and every option list this project holds was written by the lead. Three independent subjects were given 24 loci from a fresh Polish translation and asked to enumerate every live rendering, under three preambles differing in nothing but content: none, a length-matched sham of platitudes, and the fourteen rows verbatim. - The fourteen rows do not move the length of the list. FRAMEWORK − NONE = +0.292, −0.667, +0.417; mean +0.014 against a same-day byte-identical noise floor of 0.375. The registered prediction failed on its floor clause. -frac2, the fraction of sites with ≤ 2 live renderings — the only condition under whichDECIDEScan be non-zero — is 0.000 for the lead, 0.000 for all three seats unprimed, and 0.000 for all three seats framework-primed. Nine of nine cells. - So §1.4's zero is not an artifact of the lead writing long option lists. Three other subjects write lists of the same shape at the same loci. The zero survives its first test against a non-lead census, andARM-candidate-reach's warning that a correction here would flatter the project did not have to be applied. - The one arm that madeDECIDESreachable is the SHAM: 0.375 / 0.083 / 0.042. Whether a coverage statistic can fire at all turned on what was pasted above the task, and it was the preamble with no warrant behind it that moved it.And one thing gets worse rather than better, about
kand not about the zero. The same three subjects share only 0.170 of their option sets at the same loci (three-way Jaccard, frozen matcher), and each recovers 0.22–0.26 of the lead's own list. The size of a census reproduces across subjects; its contents do not.k— the renderings a candidate rules out, which S068 established as the only option-count-invariant statistic and which this page now tells sessions to quote — counts exclusions from a set three competent subjects do not agree on. Invariance to how long the list is does not buy invariance to whose list it is, and nothing before S073 distinguished the two.ARM-option-censusstep 2 owns it.⚠ STEP 2 RAN AT S078 AND THE ANSWER IS THAT THE INSTRUMENT CANNOT SAY (
RS-20260801b-census-author;ARM-option-censusclosedresolvedat 2 of 2). The four censuses — the lead's and the three seats' — were coded against the fourteen candidates on the identical 24 loci, in a four-payload rotation showing each rater each site once. Failure criterion F1 fired, under both readings of the one control whose criterion is ambiguous, so nokcomparison from that run is an estimate. What it did establish: - The control block does not reproduce. The same six items, the same prompt, two of the same three seats:E-20260730i's 6/6 and 6/6 became 3/6 and 4/6. Note (bgr). - The instrument's own noise floor is the size of the effect. A byte-identical repeat movedkby 0.333–0.667 in every arm, all in one direction, against a four-arm range of 0.714 — and 1.400 against 0.667 when both are computed on the same subsample. - Within one rater on an unchanged request,BEARSreproduced at 30 of 30 cells and the exclusion count did not (0.633 exact, +0.467 options per cell on the second pass).kis a mean of the unstable judgment over a set fixed by the stable one. Note (bgs). - Descriptively, and quotable as nothing more:k= 0.14 did not reproduce anywhere. The lead's own moment-of-decision census on Polish loci returned 1.0000; the other five arms 0.29–0.64.What this page must therefore stop saying. "What to report instead is
k" stands as a preference overDECIDESandE, both demonstrably option-count dependent. It does not stand as a licence to quote 0.14 as a property of the candidates. That figure is one run, one language pair, one translator's census, under a control block a second run could not reproduce, and it moves by a factor of seven on different material. Quote it only as 0.14 options per bearing candidate on the 45 PT→EN sites ofE-20260730i— evidence base named, never as a constant.~~And one thing about
E-20260730iitself is now open (RS-20260801b§2). That design states itsCTRL-NEGcriterion as "no candidate bears" and its analysis scored a body in which a candidate did bear as a pass. Under the stated criterion no rater there passed all six controls and that run's own F2 would have fired, which would makeRS-20260730idescriptive only — including the 0.14. S078 did not adjudicate it, having an interest in the answer; it is a gate on the next session, with the evidence inRS-20260801b§2.~~⚠ ADJUDICATED 2026-08-01 (S079), AND IT WENT THE WAY THAT COSTS THIS SECTION ITS NUMBER. Three independent non-Anthropic seats, blind to one another, on a pack sliced verbatim from the frozen files: all three read the design text alone as requiring no candidate bears, the dissenter included, and the forced choice went 2 of 3 for the stated criterion.
E-20260730i's F2 therefore fires.RS-20260730iis descriptive only.
k= 0.14 IS NOT AN ESTIMATE AND MAY NOT BE QUOTED AS ONE — not even in the narrowed form the paragraph above prescribes. The narrowing ("0.14 options per bearing candidate on the 45 PT→EN sites ofE-20260730i") named the right evidence base and the wrong standing: that run's own pre-committed failure criterion had fired, so the figure is a description of three rater bodies whose control block failed, and nothing more. A session that wants a number here has to run one. The instruction the earlier boxes gave — reportkrather thanDECIDESorE— stands as a preference between statistics and is now a preference with no measured value attached to it.What this does NOT do. It does not withdraw the zero, which
RS-20260728iandRS-20260731esupport independently. It does not touch the ladder's qualitative point — that the same behaviour reads 0.000 at four options and 0.154 at two — except to say it too is description from a failed control block. And it does not reopenRS-20260801b, whose own F1 fired independently. - 6 of the 14 candidates were reached by any decision in either log, and they are exactly the six that are admissible, about translating rather than about method, and carry a claim page: C1, C2, C3, C5, C6, C7. A seventh, C9, was invoked once and is a refused candidate. - 7 of 14 were never invoked at all: C4, C8, C10, C11, C12, C13, C14. Four of those seven are the method candidates the inventory calls "the strongest statements in the repository." S068 takes this to TEN of 14 on a third pair — only C1, C2, C3 and C5 bear on anything in 45 PT→EN decisions, C12 among the ten that do not, and again no method candidate anywhere (RS-20260730i§2). - C3 alone carries 33 of the 74 invocations — 45%. In use, this evidence base is one taxonomy of realia handling, plus a register instruction, plus two questions. On the third pair the concentration is worse and the leader changes: C1 carries 31 of 43, 72%. - And the coverage rate itself is now known not to reproduce, which the zero does (RS-20260730i§4). Pairwise rater agreement on the positive cells of the "does this candidate bear?" judgment is 0.21–0.51; the same three answer sheets give a site-level rate of 0.867 over three raters and 0.289 over the two that passed all six controls. Quote the zero; do not quote a coverage proportion from this instrument without its rater pool beside it.
2. What the one blocked candidate would need, in numbers
C12 — one self-revision pass buys naturalness without moving accuracy — is blocked twice over, and neither block is removed by the other.
- Calibration. It is X3 and Tier D is NOT PASSED (
RS-20260726d-tierD-heldout). A repaired Tier D would restore admissibility and nothing else. - Its own control does not support the inference drawn from it (
RS-20260727c§4). The temperature arm's pooled ≈0.53 is the average of two opposite-signed, length-driven halves; a calibrated jury re-run on the same design inherits it unchanged. - The repair is a cost statement, not a method statement (
RS-20260728c-length-matching). Matching produces a third text (0 REVERT / 4 NEW / 2 DEPART against a 0.0665 coincidence floor). Stratifying is not an estimator at this n: interval width fails on 3 of 5 senses at 6 items and 5 of 5 at 10, and the computed requirement is 22–30 items per length-sign stratum, roughly 45–60 pairs, against S010's 12. Four to five times the pairs. - And now a fourth thing, which is new and cuts deeper than the other three. In 126 classifications C12 was invoked zero times, on a log containing fifteen revision decisions. C12 is a claim about what a pass buys; no individual decision is a decision about whether to make a pass. So even a fully revived C12 fires on 0 of the 63 decisions this project has on record. Reviving it would make the inventory admissible and leave a translator with exactly what they have now.
3. The gate is unmet on both limbs, not one
Charter A7 and framework/README.md: no release before Tier D has run and at least one disciplined regime comparison has results.
- Tier D: NOT PASSED, on two independent grounds, with all three mandatory controls on the same materials (
RS-20260726d). Its stage-1 scale-usage gate was subsequently declared unrepairable (RS-20260728g), and its replacement is undecided. - Regime comparisons with admissible results: zero. S010's is the project's only scored comparison, and it is panel-scored on an uncalibrated jury with a control whose inference is withdrawn. S041 built a paired R06/R04 comparison (Sōseki) and deliberately did not score it; S051 built a second (Ōgai,
T-takasebune-R06-v1/-R04-v1). Both wait on the same gate.
Stating it as "blocked on Tier D" was always half the picture. The second limb has never been satisfied either, and the two are not independent: the comparisons that exist cannot be scored until the jury is calibrated, so passing Tier D is what would convert two already-built pairs into the missing evidence. The cheapest route to a release is not more translating and not more inventory work. It is Tier D.
4. framework/control-arm-spec.md and ratification
The spec was adopted by S046 on measurement and is not a ratified decision — no independent vote has been routed through it, and it binds designs only by being cited.
Recommendation, internal-judgment-only: it should go to ratification, but not as a session's principal unit and not on its own. It is a specification about how paired comparisons are analysed; a vote taken with no live design in front of it would be reviewing prose. The trigger to put it to the protocol is the next design that builds a paired comparison — which, given §3, is whatever repairs Tier D. Filed to wiki/backlog.md with that trigger rather than as a dated action.
5. The check the arm owed, and its result
ARM-framework step 4 absorbed a backlog item at S046: a second reader on RS-20260726e's decision-to-claim mapping, which that page names as its own weakest link and its revision trigger 2. Discharged by RS-20260728i §2:
- The mapping held. κ 0.905 against reader 1 and 0.715 against reader 2; readers agree with each other at κ 0.809.
- Trigger 2 fires, on one reader by one decision (3 disagreements against a threshold of more than 2; the other reader disagrees on 1).
- The one agreed error is A3, and correcting it moves the published figure to 11 of 21 — up, against the direction the project's thesis would have preferred.
- The lead predicted the mapping would fail this check badly, on the basis of S048's 0.483 classifier agreement. It did not.
This is the component of the completion criterion that could have gone the other way, and the arm is closed on it having gone the reassuring way.
6. What a release would need, in order of how much it would move this page
- A repaired Tier D that passes on at least one sense. It converts C12 from inadmissible to admissible, unblocks
CJ-20260724-scaffolded-borrowing, and turns two already-built and unscored regime pairs into the second limb of the release gate. Nothing else on this list does more than one of those. - 45–60 pairs for the self-revision comparison (§2.3) — and §2.4 says that even then a translator gets nothing at the decision level.
- ~~A prescriptive candidate that is not about a whole pass.~~ SUPPLIED 2026-08-02 (S091):
R1, displaced marking, is inframework/v0.1/. This item's own closing question — "whether the trade is a property of warrant or of this rule's wording" — is answered, and the answer is wording.R1carries a warrant (two published human translators, in two pairs, making the move independently) and is applied alike by three independent readers at κ 0.630–0.774, againstC15's 0.452. A prescriptive rule at decision grain can be both. The S056 record below stands as written and is the reason the bar was set where it was. ~~TESTED 2026-07-29 (S056), and the gap is not simply fillable~~ (historical, retained) (RS-20260729d-decision-grain,ARM-decision-grain). §1.4's zero is the sharpest thing this page knows, and it stated the item as an absence to be supplied: the project has never produced a statement of the form at a decision like this one, do X. One was written —C15, four ordered tests over the eight-wide handling set, every clause traceable to a site in a second-read anchor — frozen before any source text was opened, and put to the same two readers asRS-20260728ibeside a shape-matched sham with no warrant at all. Three things came back and all three bear on this list. - The zero is not a reader artifact. A positive control — a rule that decides with no judgment — returnedDECIDESon 58 and 12 of 63 entries. §1.4's 0 of 126 is a fact about the candidates. This page could not previously say that. -C15is recognised as deciding and is not reproducible enough to be a recommendation. Two readers applying it to 23 frozen sites agree at raw 0.609 / κ 0.452; the sham they agree at 0.957 / 0.933.E-20260729d's failure criterion F3 fired, andC15does not enterframework/traceability-inventory.md. - So the obstacle is not shape. It is that the conditions carrying the evidence — load-bearing or furniture, an exact equivalent — are precisely the conditions two readers read differently. Eight of the nine disagreements are one clause. Charter §3 wants recommendations "stated so a user (or an autonomous run) can follow them"; two readers following this one do not arrive at the same place.
What this item now says. A release needs a prescriptive candidate that is both warranted and reproducible, and the project has now built one of each and not one that is both. Whether the trade is a property of warrant or of this rule's wording is ARM-decision-grain step 2, and it is answerable for about $0.09.
4. Second-read the two intralingual anchors (A-yosano-yomogiu, A-beowulf-ingeld) — they carry C6 at X1b, and the one verification run this project has ever done moved four claims.
5. ~~A log from a translator who is not the lead.~~ MEASURED 2026-07-30 (S066), and the measurement is worse than this item assumed (RS-20260730g-nonlead-decisions, ARM-nonlead-log step 1). Every coverage number the project has, including §1.4's, is the lead's self-report about the lead — that part stands. What has changed is the premise underneath it, which was that a non-lead decision list exists somewhere and needs finding.
- Two published translators' own accounts of their own method, read in the original, classified by three independent readers on a frozen scheme with positive controls firing in English, classical Chinese and Meiji Japanese, yield ONE site-level decision in 23 units. 嚴復's 譯例言 §4 (卮言 → 懸談 → 導言, three candidates, two named objectors) is the one. 二葉亭四迷's 「余が翻訳の標準」 yields none — and its most site-specific passages describe Zhukovsky's choices, not his own. The project's own logs score 8 of 8 on the same instrument in the same call.
- So a non-lead coverage denominator cannot be built from translators' essays about method. They are apologiae; 二葉亭's is explicitly a defence of a method he had abandoned. The place to look is prefaces and notes that enumerate choices, and that is now ARM-nonlead-log's live path rather than its retirement path.
- And the item gains a route it did not have, which needs no self-report at all. 二葉亭 states a rule that is arithmetic — preserve the source's comma and full-stop counts — and his 1888 execution of it survives. A stated rule plus a surviving text is a checkable decision record. E-20260730g condition B attempted it and was voided by its own registered alignment criterion; the repair is free and is the arm's step 2.
- What a release can now say that it could not: the gap between what a translator reports and what a translator does is not a suspicion about the lead. It is measured, on somebody else, at 1 of 23.
- And the baseline that gap would be measured against is now on this page too, absorbed from
wiki/backlog.md at S067. RS-20260729e §3: two of four translator's logs account for 100% of
their own textual changes, and R04's own warning that a log under-records is real and bounded at
about one edit in ten. It is four pairs, one translator, one regime — and it is what an external
log would have to be compared against. The 1-of-23 figure and the 100%-of-4 figure are the two
halves of one comparison and must not be quoted apart.
7. What this page does not say
It does not say the evidence base is bad. X2 method work, X1a anchors and the claim pages are sound, verified and in several cases replicated; the inventory's §3 item 2 has said since S041 that this arm can go on producing sound method claims indefinitely without getting one step closer to a release, and it was right.
It does not say a release will never exist. It says the release is short of one calibrated sense and one scorable comparison, that both are the same blocker, and that no amount of further inventory work reaches it.
One sentence of §3 is now under-stated rather than wrong. "The cheapest route to a release is not more translating and not more inventory work. It is Tier D." That remains true of the release gate. It is not true of everything a release needs: §6 item 3 was reachable without a jury, was reached, and returned a result that changes what a release would have to contain. Recorded here because the sentence has been quoted as though it closed off all non-Tier-D work.
It does not say the coverage proportion is ~0.5. RS-20260728i §3 withdraws that inference; what is established is the zero, not the half. S068 strengthens this from an inference problem to a measurement problem: the coverage proportion does not reproduce across rater pools on identical data (0.867 against 0.289), while the zero reproduces on every pool and on a third language pair. This sentence was more right than the page knew when it wrote it.
8. Where the five items live, added 2026-07-30 (S066)
This page listed five things a release needs and named an owner for none of them, and that is why T5's count keeps rising while nothing on this list moves. The map, so that a session reaching for T5 because the count is high can see what it would actually be taking:
| item | owner | state |
|---|---|---|
| 1 — a repaired Tier D that passes on one sense | ~~nobody~~ ~~ARM-tierD-run~~ — the arm ran the gate and closed resolved 2026-08-02 (S086) |
ANSWERED, and the answer is no: TIER D NOT PASSED (RS-20260802-tierD-verdict). The repaired instrument was run to completion with all four controls; it fails on one pre-registered number, drop(naturalness) = +1.12 at the 8-site dose against a ≤ 0.75 condition, having fired detection at 12 of 12 units in both targeted cells within every juror separately. So this item is not blocked and not owned — it is decided against, and a v0.1 release does not open. What would change it is a decision about which dose is the primary, which needs Tom (ARM-tierD-run §Done when). ARM-first-judgment is unblocked provisional-only; ARM-framework-v01 stays blocked on this row |
| 2 — 45–60 pairs for the self-revision comparison | nobody — ARM-atelier-cycle may bank pairs as craft, but nothing owns the count |
and §2.4 says a translator gets nothing at the decision level even then |
| 3 — a prescriptive candidate both warranted and reproducible | ARM-decision-grain, closed retired S061 |
tested; the trade is unresolved and not currently decidable |
| 4 — second-read the two intralingual anchors | ARM-anchor-second-read (T4, budget 2, never worked) |
live, on another track |
| 5 — a log from a translator who is not the lead | ARM-nonlead-log (T1, budget 2, used 1) |
worked S066; see the item above |
Seventh row, added 2026-07-31 (S073). The denominator every figure on this page is computed
over — the option census — is owned by ARM-option-census (T5, budget 2, used 1, open). Step 1
is done and is the second boxed correction in §1.4; step 2 asks whether k survives a change of
census author, which step 1 showed is a different question from whether it survives a change of
option-list length. That is now three consecutive T5 units that were neither downstream of Tier D
nor a fourth container — ARM-candidate-reach and both steps of this arm — and the pattern is
worth naming: the numbers this page publishes are themselves T5 work, and there are more of them
than there are things a release needs.
Sixth row, added 2026-07-30 (S068), because this page's own headline number turned out to need an owner too. §1.4's zero is owned by ARM-candidate-reach (T5, budget 2, used 1) — the first arm this ledger has ever carried on T5 that is not a release push, constituted when check_balance.py named T5 for the third consecutive session at a count of 6. Step 1 is done and is the correction boxed in §1.4; step 2 is this table and the inventory. What it demonstrates for the next session that reaches for T5: there is T5 work that is neither downstream of Tier D nor a fourth container — auditing the numbers this page already publishes.
Three of the five are owned by arms on tracks other than T5, and the two that are not owned by anything are the two that are downstream of Tier D. That is the whole content of S051's standing sentence — the next T5 unit is downstream of Tier D, not parallel to it — restated as a table instead of as a warning, and it is why S066 took ARM-nonlead-log when the tool named T5. A track whose work is done on other tracks will keep reading as starved, and the honest remedy is this table rather than a fourth container.