Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: framework/closure.md · rendered 2026-09-09

Page metadata (front matter)
typeledger
idframework-closure
statusactive
created2026-07-28
updated2026-08-02
sensesaccuracy, naturalness, style-correspondence, voice, cultural-mediation, consistency, purpose-fit
provisionaltrue
internal-judgment-onlytrue
linksframework/README.md, wiki/findings/results/RS-20260730g-nonlead-decisions.md, wiki/arms/ARM-nonlead-log.md, wiki/findings/results/RS-20260729d-decision-grain.md, wiki/arms/ARM-decision-grain.md, framework/traceability-inventory.md, framework/control-arm-spec.md, wiki/arms/ARM-framework.md, wiki/findings/results/RS-20260728i-coverage-independent.md, wiki/findings/results/RS-20260726e-framework-coverage.md, wiki/findings/results/RS-20260728c-length-matching.md, wiki/findings/results/RS-20260726d-tierD-heldout.md, PROJECT.md, config/models.md

What a framework release is short of — the closure statement

What this is. ARM-framework step 4, and the arm's declared ending. The completion criterion set at S035, sixteen sessions before this page: either framework/v0.1/ exists, or the arm closes with a written statement of what evidence a release is short of — a statement the inventory can now make in numbers rather than in prose. framework/v0.1/ does not exist. This is the other ending.

On the status word. S035's criterion said the arm would close retired. It closes resolved, and the difference is not a softening: under CLAUDE.md's vocabulary retired means abandoned with a written reason and resolved means complete. The arm reached a completion criterion it declared for itself and produced the artifact that criterion names. S035 reached for retired because it read closure-without-release as a failure to release. It is not: it is the answer to the arm's own question, what does the project actually recommend doing, for which pairs, on what evidence? The answer is almost nothing, and here is the arithmetic.

Standing. internal-judgment-only and provisional. Every count below is recomputable from the inventory table and from RS-20260728i's stored outputs; every judgment about what a release should do is the lead's.


1. The counts

1.1 Fourteen candidates, by shape

shape n which standing
prescriptive about translating — do X rather than Y, addressed to a translator 1 C12 inadmissible, on two independent grounds (§2)
prescriptive about method — addressed to the project 4 C10, C11, C13, C14 sound, X2, verified, replicated — and not framework content
procedural, diagnostic, taxonomic or descriptive 9 C1–C9 less C12 a vocabulary and a set of questions

1.2 The same fourteen, by evidence class after Tier D's failure

class n standing
X1a externally anchored, second-read 6 (C1, C2, C3, C5, C7, C9) intact
X1b externally anchored, single-reader 3 (C4, C6, C8) intact but unverified
X2 machine-measured 4 (C10, C11, C13, C14) intact
X3 panel-scored 1 (C12) INADMISSIBLE — config/models.md NOT CALIBRATED

1.3 What a release could carry, addressed to a translator

Zero prescriptive recommendations. One candidate has that shape and it is inadmissible. Charter §3 requires recommendations "stated so a user (or an autonomous run) can follow them"; "here are eight things translators do" is not one.

1.4 The number this session added: how many candidates a real decision ever reaches

RS-20260728i: two independent readers, 63 logged translation decisions from two languages, 126 classifications.

⚠ CORRECTED 2026-07-30 (S068), and the correction is to the FORM of this number, not to its substance (RS-20260730i-candidate-reach, ARM-candidate-reach step 1). This section said the zero was "a fact about the candidates". It is two claims and only one of them is. A third language pair (PT→EN, 45 sites, mean 3.98 live renderings per site, three raters) was put to the panel twice: once with the options the translator actually had, once with two. A bearing candidate excludes 0.1395 options at four and 0.1538 at two — invariant. That same behaviour reads DECIDES = 0.000 at four options and 0.154 at two. So the magnitude is a fact about the candidates and the code is a fact about how long the translator's option list was: a translator keeping a two-option log would have published a non-zero rate from the same evidence base and the same candidate behaviour. What to report instead is k, the absolute number of renderings a candidate rules out — the only quantity that did not move when the list was halved. It is 0.14. The normalised E = k/n is not invariant (0.0368 → 0.0769) and must not be used. The zero is not withdrawn and nothing here makes the evidence base look better. At two options the candidates still exclude a seventh of an option.

⚠ TESTED AGAINST NON-LEAD CENSUSES 2026-07-31 (S073), and this time the correction goes the other way (RS-20260731e-option-census, ARM-option-census step 1). The box above says the DECIDES code is "a fact about how long the translator's option list was", and every option list this project holds was written by the lead. Three independent subjects were given 24 loci from a fresh Polish translation and asked to enumerate every live rendering, under three preambles differing in nothing but content: none, a length-matched sham of platitudes, and the fourteen rows verbatim. - The fourteen rows do not move the length of the list. FRAMEWORK − NONE = +0.292, −0.667, +0.417; mean +0.014 against a same-day byte-identical noise floor of 0.375. The registered prediction failed on its floor clause. - frac2, the fraction of sites with ≤ 2 live renderings — the only condition under which DECIDES can be non-zero — is 0.000 for the lead, 0.000 for all three seats unprimed, and 0.000 for all three seats framework-primed. Nine of nine cells. - So §1.4's zero is not an artifact of the lead writing long option lists. Three other subjects write lists of the same shape at the same loci. The zero survives its first test against a non-lead census, and ARM-candidate-reach's warning that a correction here would flatter the project did not have to be applied. - The one arm that made DECIDES reachable is the SHAM: 0.375 / 0.083 / 0.042. Whether a coverage statistic can fire at all turned on what was pasted above the task, and it was the preamble with no warrant behind it that moved it.

And one thing gets worse rather than better, about k and not about the zero. The same three subjects share only 0.170 of their option sets at the same loci (three-way Jaccard, frozen matcher), and each recovers 0.22–0.26 of the lead's own list. The size of a census reproduces across subjects; its contents do not. k — the renderings a candidate rules out, which S068 established as the only option-count-invariant statistic and which this page now tells sessions to quote — counts exclusions from a set three competent subjects do not agree on. Invariance to how long the list is does not buy invariance to whose list it is, and nothing before S073 distinguished the two. ARM-option-census step 2 owns it.

⚠ STEP 2 RAN AT S078 AND THE ANSWER IS THAT THE INSTRUMENT CANNOT SAY (RS-20260801b-census-author; ARM-option-census closed resolved at 2 of 2). The four censuses — the lead's and the three seats' — were coded against the fourteen candidates on the identical 24 loci, in a four-payload rotation showing each rater each site once. Failure criterion F1 fired, under both readings of the one control whose criterion is ambiguous, so no k comparison from that run is an estimate. What it did establish: - The control block does not reproduce. The same six items, the same prompt, two of the same three seats: E-20260730i's 6/6 and 6/6 became 3/6 and 4/6. Note (bgr). - The instrument's own noise floor is the size of the effect. A byte-identical repeat moved k by 0.333–0.667 in every arm, all in one direction, against a four-arm range of 0.714 — and 1.400 against 0.667 when both are computed on the same subsample. - Within one rater on an unchanged request, BEARS reproduced at 30 of 30 cells and the exclusion count did not (0.633 exact, +0.467 options per cell on the second pass). k is a mean of the unstable judgment over a set fixed by the stable one. Note (bgs). - Descriptively, and quotable as nothing more: k = 0.14 did not reproduce anywhere. The lead's own moment-of-decision census on Polish loci returned 1.0000; the other five arms 0.29–0.64.

What this page must therefore stop saying. "What to report instead is k" stands as a preference over DECIDES and E, both demonstrably option-count dependent. It does not stand as a licence to quote 0.14 as a property of the candidates. That figure is one run, one language pair, one translator's census, under a control block a second run could not reproduce, and it moves by a factor of seven on different material. Quote it only as 0.14 options per bearing candidate on the 45 PT→EN sites of E-20260730i — evidence base named, never as a constant.

~~And one thing about E-20260730i itself is now open (RS-20260801b §2). That design states its CTRL-NEG criterion as "no candidate bears" and its analysis scored a body in which a candidate did bear as a pass. Under the stated criterion no rater there passed all six controls and that run's own F2 would have fired, which would make RS-20260730i descriptive only — including the 0.14. S078 did not adjudicate it, having an interest in the answer; it is a gate on the next session, with the evidence in RS-20260801b §2.~~

⚠ ADJUDICATED 2026-08-01 (S079), AND IT WENT THE WAY THAT COSTS THIS SECTION ITS NUMBER. Three independent non-Anthropic seats, blind to one another, on a pack sliced verbatim from the frozen files: all three read the design text alone as requiring no candidate bears, the dissenter included, and the forced choice went 2 of 3 for the stated criterion. E-20260730i's F2 therefore fires. RS-20260730i is descriptive only.

k = 0.14 IS NOT AN ESTIMATE AND MAY NOT BE QUOTED AS ONE — not even in the narrowed form the paragraph above prescribes. The narrowing ("0.14 options per bearing candidate on the 45 PT→EN sites of E-20260730i") named the right evidence base and the wrong standing: that run's own pre-committed failure criterion had fired, so the figure is a description of three rater bodies whose control block failed, and nothing more. A session that wants a number here has to run one. The instruction the earlier boxes gave — report k rather than DECIDES or E — stands as a preference between statistics and is now a preference with no measured value attached to it.

What this does NOT do. It does not withdraw the zero, which RS-20260728i and RS-20260731e support independently. It does not touch the ladder's qualitative point — that the same behaviour reads 0.000 at four options and 0.154 at two — except to say it too is description from a failed control block. And it does not reopen RS-20260801b, whose own F1 fired independently. - 6 of the 14 candidates were reached by any decision in either log, and they are exactly the six that are admissible, about translating rather than about method, and carry a claim page: C1, C2, C3, C5, C6, C7. A seventh, C9, was invoked once and is a refused candidate. - 7 of 14 were never invoked at all: C4, C8, C10, C11, C12, C13, C14. Four of those seven are the method candidates the inventory calls "the strongest statements in the repository." S068 takes this to TEN of 14 on a third pair — only C1, C2, C3 and C5 bear on anything in 45 PT→EN decisions, C12 among the ten that do not, and again no method candidate anywhere (RS-20260730i §2). - C3 alone carries 33 of the 74 invocations — 45%. In use, this evidence base is one taxonomy of realia handling, plus a register instruction, plus two questions. On the third pair the concentration is worse and the leader changes: C1 carries 31 of 43, 72%. - And the coverage rate itself is now known not to reproduce, which the zero does (RS-20260730i §4). Pairwise rater agreement on the positive cells of the "does this candidate bear?" judgment is 0.21–0.51; the same three answer sheets give a site-level rate of 0.867 over three raters and 0.289 over the two that passed all six controls. Quote the zero; do not quote a coverage proportion from this instrument without its rater pool beside it.

2. What the one blocked candidate would need, in numbers

C12 — one self-revision pass buys naturalness without moving accuracy — is blocked twice over, and neither block is removed by the other.

  1. Calibration. It is X3 and Tier D is NOT PASSED (RS-20260726d-tierD-heldout). A repaired Tier D would restore admissibility and nothing else.
  2. Its own control does not support the inference drawn from it (RS-20260727c §4). The temperature arm's pooled ≈0.53 is the average of two opposite-signed, length-driven halves; a calibrated jury re-run on the same design inherits it unchanged.
  3. The repair is a cost statement, not a method statement (RS-20260728c-length-matching). Matching produces a third text (0 REVERT / 4 NEW / 2 DEPART against a 0.0665 coincidence floor). Stratifying is not an estimator at this n: interval width fails on 3 of 5 senses at 6 items and 5 of 5 at 10, and the computed requirement is 22–30 items per length-sign stratum, roughly 45–60 pairs, against S010's 12. Four to five times the pairs.
  4. And now a fourth thing, which is new and cuts deeper than the other three. In 126 classifications C12 was invoked zero times, on a log containing fifteen revision decisions. C12 is a claim about what a pass buys; no individual decision is a decision about whether to make a pass. So even a fully revived C12 fires on 0 of the 63 decisions this project has on record. Reviving it would make the inventory admissible and leave a translator with exactly what they have now.

3. The gate is unmet on both limbs, not one

Charter A7 and framework/README.md: no release before Tier D has run and at least one disciplined regime comparison has results.

Stating it as "blocked on Tier D" was always half the picture. The second limb has never been satisfied either, and the two are not independent: the comparisons that exist cannot be scored until the jury is calibrated, so passing Tier D is what would convert two already-built pairs into the missing evidence. The cheapest route to a release is not more translating and not more inventory work. It is Tier D.

4. framework/control-arm-spec.md and ratification

The spec was adopted by S046 on measurement and is not a ratified decision — no independent vote has been routed through it, and it binds designs only by being cited.

Recommendation, internal-judgment-only: it should go to ratification, but not as a session's principal unit and not on its own. It is a specification about how paired comparisons are analysed; a vote taken with no live design in front of it would be reviewing prose. The trigger to put it to the protocol is the next design that builds a paired comparison — which, given §3, is whatever repairs Tier D. Filed to wiki/backlog.md with that trigger rather than as a dated action.

5. The check the arm owed, and its result

ARM-framework step 4 absorbed a backlog item at S046: a second reader on RS-20260726e's decision-to-claim mapping, which that page names as its own weakest link and its revision trigger 2. Discharged by RS-20260728i §2:

This is the component of the completion criterion that could have gone the other way, and the arm is closed on it having gone the reassuring way.

6. What a release would need, in order of how much it would move this page

  1. A repaired Tier D that passes on at least one sense. It converts C12 from inadmissible to admissible, unblocks CJ-20260724-scaffolded-borrowing, and turns two already-built and unscored regime pairs into the second limb of the release gate. Nothing else on this list does more than one of those.
  2. 45–60 pairs for the self-revision comparison (§2.3) — and §2.4 says that even then a translator gets nothing at the decision level.
  3. ~~A prescriptive candidate that is not about a whole pass.~~ SUPPLIED 2026-08-02 (S091): R1, displaced marking, is in framework/v0.1/. This item's own closing question — "whether the trade is a property of warrant or of this rule's wording" — is answered, and the answer is wording. R1 carries a warrant (two published human translators, in two pairs, making the move independently) and is applied alike by three independent readers at κ 0.630–0.774, against C15's 0.452. A prescriptive rule at decision grain can be both. The S056 record below stands as written and is the reason the bar was set where it was. ~~TESTED 2026-07-29 (S056), and the gap is not simply fillable~~ (historical, retained) (RS-20260729d-decision-grain, ARM-decision-grain). §1.4's zero is the sharpest thing this page knows, and it stated the item as an absence to be supplied: the project has never produced a statement of the form at a decision like this one, do X. One was written — C15, four ordered tests over the eight-wide handling set, every clause traceable to a site in a second-read anchor — frozen before any source text was opened, and put to the same two readers as RS-20260728i beside a shape-matched sham with no warrant at all. Three things came back and all three bear on this list. - The zero is not a reader artifact. A positive control — a rule that decides with no judgment — returned DECIDES on 58 and 12 of 63 entries. §1.4's 0 of 126 is a fact about the candidates. This page could not previously say that. - C15 is recognised as deciding and is not reproducible enough to be a recommendation. Two readers applying it to 23 frozen sites agree at raw 0.609 / κ 0.452; the sham they agree at 0.957 / 0.933. E-20260729d's failure criterion F3 fired, and C15 does not enter framework/traceability-inventory.md. - So the obstacle is not shape. It is that the conditions carrying the evidence — load-bearing or furniture, an exact equivalent — are precisely the conditions two readers read differently. Eight of the nine disagreements are one clause. Charter §3 wants recommendations "stated so a user (or an autonomous run) can follow them"; two readers following this one do not arrive at the same place.

What this item now says. A release needs a prescriptive candidate that is both warranted and reproducible, and the project has now built one of each and not one that is both. Whether the trade is a property of warrant or of this rule's wording is ARM-decision-grain step 2, and it is answerable for about $0.09. 4. Second-read the two intralingual anchors (A-yosano-yomogiu, A-beowulf-ingeld) — they carry C6 at X1b, and the one verification run this project has ever done moved four claims. 5. ~~A log from a translator who is not the lead.~~ MEASURED 2026-07-30 (S066), and the measurement is worse than this item assumed (RS-20260730g-nonlead-decisions, ARM-nonlead-log step 1). Every coverage number the project has, including §1.4's, is the lead's self-report about the lead — that part stands. What has changed is the premise underneath it, which was that a non-lead decision list exists somewhere and needs finding. - Two published translators' own accounts of their own method, read in the original, classified by three independent readers on a frozen scheme with positive controls firing in English, classical Chinese and Meiji Japanese, yield ONE site-level decision in 23 units. 嚴復's 譯例言 §4 (卮言 → 懸談 → 導言, three candidates, two named objectors) is the one. 二葉亭四迷's 「余が翻訳の標準」 yields none — and its most site-specific passages describe Zhukovsky's choices, not his own. The project's own logs score 8 of 8 on the same instrument in the same call. - So a non-lead coverage denominator cannot be built from translators' essays about method. They are apologiae; 二葉亭's is explicitly a defence of a method he had abandoned. The place to look is prefaces and notes that enumerate choices, and that is now ARM-nonlead-log's live path rather than its retirement path. - And the item gains a route it did not have, which needs no self-report at all. 二葉亭 states a rule that is arithmetic — preserve the source's comma and full-stop counts — and his 1888 execution of it survives. A stated rule plus a surviving text is a checkable decision record. E-20260730g condition B attempted it and was voided by its own registered alignment criterion; the repair is free and is the arm's step 2. - What a release can now say that it could not: the gap between what a translator reports and what a translator does is not a suspicion about the lead. It is measured, on somebody else, at 1 of 23. - And the baseline that gap would be measured against is now on this page too, absorbed from wiki/backlog.md at S067. RS-20260729e §3: two of four translator's logs account for 100% of their own textual changes, and R04's own warning that a log under-records is real and bounded at about one edit in ten. It is four pairs, one translator, one regime — and it is what an external log would have to be compared against. The 1-of-23 figure and the 100%-of-4 figure are the two halves of one comparison and must not be quoted apart.

7. What this page does not say

It does not say the evidence base is bad. X2 method work, X1a anchors and the claim pages are sound, verified and in several cases replicated; the inventory's §3 item 2 has said since S041 that this arm can go on producing sound method claims indefinitely without getting one step closer to a release, and it was right.

It does not say a release will never exist. It says the release is short of one calibrated sense and one scorable comparison, that both are the same blocker, and that no amount of further inventory work reaches it.

One sentence of §3 is now under-stated rather than wrong. "The cheapest route to a release is not more translating and not more inventory work. It is Tier D." That remains true of the release gate. It is not true of everything a release needs: §6 item 3 was reachable without a jury, was reached, and returned a result that changes what a release would have to contain. Recorded here because the sentence has been quoted as though it closed off all non-Tier-D work.

It does not say the coverage proportion is ~0.5. RS-20260728i §3 withdraws that inference; what is established is the zero, not the half. S068 strengthens this from an inference problem to a measurement problem: the coverage proportion does not reproduce across rater pools on identical data (0.867 against 0.289), while the zero reproduces on every pool and on a third language pair. This sentence was more right than the page knew when it wrote it.

8. Where the five items live, added 2026-07-30 (S066)

This page listed five things a release needs and named an owner for none of them, and that is why T5's count keeps rising while nothing on this list moves. The map, so that a session reaching for T5 because the count is high can see what it would actually be taking:

item owner state
1 — a repaired Tier D that passes on one sense ~~nobody~~ ~~ARM-tierD-run~~ — the arm ran the gate and closed resolved 2026-08-02 (S086) ANSWERED, and the answer is no: TIER D NOT PASSED (RS-20260802-tierD-verdict). The repaired instrument was run to completion with all four controls; it fails on one pre-registered number, drop(naturalness) = +1.12 at the 8-site dose against a ≤ 0.75 condition, having fired detection at 12 of 12 units in both targeted cells within every juror separately. So this item is not blocked and not owned — it is decided against, and a v0.1 release does not open. What would change it is a decision about which dose is the primary, which needs Tom (ARM-tierD-run §Done when). ARM-first-judgment is unblocked provisional-only; ARM-framework-v01 stays blocked on this row
2 — 45–60 pairs for the self-revision comparison nobody — ARM-atelier-cycle may bank pairs as craft, but nothing owns the count and §2.4 says a translator gets nothing at the decision level even then
3 — a prescriptive candidate both warranted and reproducible ARM-decision-grain, closed retired S061 tested; the trade is unresolved and not currently decidable
4 — second-read the two intralingual anchors ARM-anchor-second-read (T4, budget 2, never worked) live, on another track
5 — a log from a translator who is not the lead ARM-nonlead-log (T1, budget 2, used 1) worked S066; see the item above

Seventh row, added 2026-07-31 (S073). The denominator every figure on this page is computed over — the option census — is owned by ARM-option-census (T5, budget 2, used 1, open). Step 1 is done and is the second boxed correction in §1.4; step 2 asks whether k survives a change of census author, which step 1 showed is a different question from whether it survives a change of option-list length. That is now three consecutive T5 units that were neither downstream of Tier D nor a fourth container — ARM-candidate-reach and both steps of this arm — and the pattern is worth naming: the numbers this page publishes are themselves T5 work, and there are more of them than there are things a release needs.

Sixth row, added 2026-07-30 (S068), because this page's own headline number turned out to need an owner too. §1.4's zero is owned by ARM-candidate-reach (T5, budget 2, used 1) — the first arm this ledger has ever carried on T5 that is not a release push, constituted when check_balance.py named T5 for the third consecutive session at a count of 6. Step 1 is done and is the correction boxed in §1.4; step 2 is this table and the inventory. What it demonstrates for the next session that reaches for T5: there is T5 work that is neither downstream of Tier D nor a fourth container — auditing the numbers this page already publishes.

Three of the five are owned by arms on tracks other than T5, and the two that are not owned by anything are the two that are downstream of Tier D. That is the whole content of S051's standing sentence — the next T5 unit is downstream of Tier D, not parallel to it — restated as a table instead of as a warning, and it is why S066 took ARM-nonlead-log when the tool named T5. A track whose work is done on other tracks will keep reading as starved, and the honest remedy is this table rather than a fourth container.