Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: framework/traceability-inventory.md · rendered 2026-09-09

Page metadata (front matter)
typeledger
idtraceability-inventory
statusactive
created2026-07-26
updated2026-07-29
sensesaccuracy, naturalness, voice, style-correspondence, affect, literary-quality, cultural-mediation, purpose-fit, consistency
provisionaltrue
linksframework/README.md, wiki/findings/results/RS-20260729d-decision-grain.md, wiki/arms/ARM-decision-grain.md, framework/closure.md, wiki/findings/results/RS-20260728i-coverage-independent.md, wiki/arms/ARM-framework.md, wiki/findings/results/RS-20260727c-arm-identifiability.md, wiki/findings/claims/CL-20260727-arm-identifiability.md, wiki/findings/results/RS-20260726e-framework-coverage.md, wiki/goodness-senses.md, wiki/decisions/resolved/D-20260724-04-pair-relative-sense-weights.md, PROJECT.md

The traceability inventory — what a framework release could say today, and on what

What this is. ARM-framework's first step (S035), extended at S041: for every candidate operational recommendation the project's evidence could support, the evidence, its class, the language pairs it is evidenced on, and whether it would survive into a release. Built at S035, over 17 result pages, 8 anchors, 10 source pages, 2 theory pages, 1 essay, 1 conjecture and 1 open question. It compiles nothing new; it sorts what exists.

Why it could be built while the release is blocked. Charter A7 gates a release on Tier D. It does not gate taking stock. T5 · Framework had never supplied a principal unit in 34 sessions, and ARM-framework recorded the reason as "deferred was read as nothing to do here."

The one-line answer. The project can supply a vocabulary, a set of diagnostic questions, and a taxonomy of options. It cannot supply a single rule that decides a case. Of fourteen candidate recommendations, one has the shape do X rather than Y addressed to a translator, and it is the one Tier D's failure makes inadmissible — and as of S041 it is inadmissible twice over, on grounds calibration would not fix (§4).

S046 added the fourteenth, and the honest note is that the ratio did not merely fail to move — the pattern behind it is now visible. The new candidate is the fourth consecutive addition that is prescriptive about method rather than about translating. framework/control-arm-spec.md is a real deliverable and it is not framework content in the sense a release needs. What that says about the arm is in §3 item 2, and it bears on how the arm's remaining budget is spent.


1. The evidence classes

Sorting by class rather than by topic is the whole method here, because Tier D's failure does not damage the evidence base uniformly — it destroys one class of it and leaves the rest untouched.

class what it is standing after RS-20260726d-tierD-heldout
X1a externally anchored, second-read — a published translation read against its source, or documented reception/scholarship, with an independent check on the reading intact. Unaffected by jury calibration
X1b externally anchored, single-reader — same kind of object, never independently checked intact but unverified. The one verification run the project has done retracted one claim and corrected three
X2 machine-measured — a number computed from stored texts, independently recomputed intact. No judgment involved
X3 panel-scored — rests on jury verdicts INADMISSIBLE. Tier D NOT PASSED; config/models.md NOT CALIBRATED; no jury verdict carries evidential weight about any translation
X4 internal judgment only — the lead's reading, unanchored never sole support for a framework recommendation (charter §2.2, and wiki/findings/README.md: "No claim rests solely on internal-judgment-only grounds")

2. The inventory

# candidate recommendation class senses pairs evidenced on shape claim page
1 Declare the target register; "natural" is not a single corpus X1a naturalness, purpose-fit EN as target, any source procedural CL-20260726-register-declaration
2 Ask what the fluency cost, not whether it is too fluent X1a naturalness, cultural-mediation, affect JA→EN diagnostic CL-20260726-fluency-cost
3 The handling set for culture-bound items is eight-wide; choose knowingly X1a ×3, X1b ×2 cultural-mediation, style-correspondence, consistency JA→EN, RU→EN, EN→FR, JA→JA, OE→EN taxonomic CL-20260726-handling-set
4 Declare the pair; sense weights are pair-relative X1b synthesis all 3-cell design ratified rule (D-20260724-04) (a decision, not a claim)
5 Grammar-borne meaning transcodes into lexis and loses systematicity X1a style-correspondence RU→EN, JA→EN (EN-target only) descriptive CL-20260726-grammar-transcoding
6 The false friend does its damage inside a drift window X1b style-correspondence, accuracy OE→EN, JA→JA (diachronic) descriptive (mechanism) CL-20260726-drift-window
7 Published translators do not handle a forked class uniformly X1a ×3, X1b ×2 consistency, cultural-mediation 5 pairs descriptive, normatively unsettled CL-20260726-forked-class-nonuniformity
8 Archaism buys inheritability back — foreignisation is a capability, not only a stance X1b + X1a (3 Tier 2 sources in the original) style-correspondence, purpose-fit OE→EN descriptive (folded into #6's evidence; not separately claimed — see §4)
9 Long source periods are split in published practice X1a, n = 1 style-correspondence, naturalness JA→EN descriptive (refused — see §4)
10 Measure the lead's contamination before selecting material X2 — all prescriptive (method) CL-20260726-lead-centrality
11 Check a published pair for dependence before using it as a baseline X2 — LA→EN, DE→EN, RU→EN prescriptive (method) CL-20260726-baseline-dependence
12 One self-revision pass buys naturalness without moving accuracy X3 naturalness, accuracy, style-correspondence JA→EN prescriptive (translation) (refused — inadmissible, and now on two independent grounds; see §4)
13 Measure whether a comparison's two arms are identifiable without reading, with a sign count rather than a correlation X2 — all prescriptive (method) CL-20260727-arm-identifiability
14 Never pool a paired comparison across length-sign strata — and never read the stratified figure as an estimate below ~22 items per stratum X2 — all prescriptive (method) framework/control-arm-spec.md (a specification, not yet a claim page)

3. What the table says, in four numbers

(Updated 2026-07-27, S041: the table now holds thirteen candidates. §3's four numbers are restated below with the thirteenth included; the diagnosis they support is unchanged and has become sharper.)

A fifth number, added 2026-07-30 (S068) — how far these fourteen rows reach when a third language pair is put to them. RS-20260730i-candidate-reach: 45 PT→EN decision sites, every live rendering written down at the moment of decision, three raters given all fourteen rows with C12's text as written.

A sixth number, added 2026-07-31 (S073) — what the fourteen rows do to a translator who is handed them. RS-20260731e-option-census: three independent subjects enumerated live English renderings at 24 loci of a fresh Polish translation, under no preamble, a length-matched sham of platitudes, and these fourteen rows verbatim.

A seventh number, added 2026-08-01 (S078) — and it is a number about the instrument rather than about the rows. RS-20260801b-census-author, ARM-option-census step 2, descriptive only: that run's F1 fired.

  1. Thirteen candidates; one is prescriptive about translating; it is inadmissible. #12 is the only entry of the form do X rather than Y addressed to a translator. It rests on the project's only regime comparison (RS-20260724-selfrevise-first, S010, 25 sessions ago), which is panel-scored on an uncalibrated jury. Tier D's failure does not weaken it — it removes it. The framework's emptiness is not an oversight; it is the measured downstream cost of the calibration gate.
  2. The four prescriptive claims that survive are about running the project, not about translating. #10, #11, #13 and now #14 are X2 — machine-measured, verified, replicated. They are the strongest statements in the repository and not one of them would appear in a release addressed to a translator. S041 added the third and S046 the fourth, and the ratio did not move: every session that strengthens this inventory strengthens the methods column. Four in a row is no longer a coincidence, and the mechanism is not mysterious. The release is gated on Tier D; the only work this arm has that is not gated is method; so the arm produces method. ARM-framework can go on producing sound X2 method claims indefinitely without getting one step closer to a release, and its two remaining sessions should be spent by someone who knows that.
  3. Nine of thirteen are procedural, diagnostic, taxonomic or descriptive. They give a translator a vocabulary and a set of questions. Under charter §3's requirement that recommendations be "stated so a user (or an autonomous run) can follow them", "here are eight things translators do" is not a recommendation.
  4. Zero of thirteen are evidenced on French→English, the pair of S035's own translation, and zero are evidenced on the JA→EN pair of S041's beyond #2 and #5. Under D-20260724-04 every one of them would carry untested if applied to a pair outside the five in column 5, which is most pairs.

4. Refusals, and they carry information

Three candidates were refused claim status, and the reasons are the inventory's real output.

A fourth is worth naming as a near-miss: CJ-20260724-scaffolded-borrowing is the project's one falsifiable prescriptive bet, it is well-motivated, and both of its testable forms require a jury with sense-localised authority. It cannot be tested until Tier D passes. It stays a conjecture.

5. What would move this page

In order of how much they would change the table, not of how easy they are:

  1. Repair the two Tier D controls and pass Tier D (wiki/backlog.md, owed). This is the only thing that converts any X3 evidence into a recommendation, and it unblocks CJ-20260724-scaffolded-borrowing. Amended S041: it no longer unblocks #12 on its own — §4 records a second defect, in the temperature control, that a calibrated jury re-run on the same design would inherit unchanged.
  2. ~~A second regime comparison.~~ BUILT 2026-07-27 (S041), and unscored by design. R06 (lead single pass) was created and frozen, R04 frozen at v1.0 with the draft now required to be frozen separately, and T-kusamakura-vii-bath-R06-v1 / -R04-v1 filed as a pair that is paired by construction — the lead analogue of the R01/R02 property. Sōseki 草枕 ch. VII, JA→EN, 1,956 source characters, both logs frozen, contamination measured at a longest run of 8 against a null-control floor of 5. The scoring is deliberately not done: it waits on a repaired Tier D and on the length-matched control §4 now requires. The durable part — the prose and the frozen decision logs — exists and does not decay. The new item this created: a future scoring of that pair must be designed around its own arm-identifiability figure (#13), which is what S041 pre-registered before translating a word of it.
  3. Second-read the two intralingual anchors (A-yosano-yomogiu, A-beowulf-ingeld). They carry #6, the project's most quantitative close-reading result, at class X1b. The one verification run the project has ever done moved four claims.
  4. ARM-typology-logs — open-coding the frozen translator's logs. §3's finding is that the project's recommendations are not operational; the logs are the only evidence it holds about what translators actually decide, and wiki/program.md Slate G already says operational recommendations should come from documented decisions rather than plausibility. RS-20260726e-framework-coverage is a single instance of exactly that method and found 11 of 21 decisions uncovered.
  5. A pair with English as source (ARM-breadth, top C1 priority since S015). #5 rests on two instances that share a target language.

5b. Amendment 2026-07-28 (S051) — the table has now been measured from the other end, and ARM-framework closes on it

RS-20260728i-coverage-independent: two independent readers, 63 logged translation decisions from two language pairs, 126 classifications, DECIDES 0. The readers were given all fourteen rows including #12, with its text as written and no note that it is inadmissible, and neither used DECIDES once. §3 item 1's "one is prescriptive about translating; it is inadmissible" is now joined by a stronger statement that does not depend on admissibility at all: on the 63 decisions this project has on record, #12 is invoked zero times, so reviving it would change nothing at the decision level.

Which rows a real decision reaches. 6 of 14 — #1, #2, #3, #5, #6, #7 — plus #9, a refused candidate, once. Seven were never invoked: #4, #8, #10, #11, #12, #13, #14, four of them the X2 method rows §3 item 2 calls the strongest statements in the repository. #3 alone carries 33 of the 74 invocations, 45%.

And one figure on this page is corrected upward. RS-20260726e's 10-of-21 mapping was independently re-read (κ 0.905 and 0.715), and the single decision both readers agree it got wrong is A3: coverage there is 11 of 21, not 10. §6's citation of the coverage figures carries that correction.

The coverage proportion is not an estimate of anything and this page must stop treating it as one — RS-20260728i §3 withdraws the inference before the run, on the pre-run critic's finding that priming changes which decisions get recorded, not only which get noticed. What survives is the zero.

The closure statement ARM-framework owed is framework/closure.md.

5c. Amendment 2026-07-29 (S056) — a fifteenth candidate was written, tested, and REFUSED

C15 is not in the table above and must not be added. It was written for E-20260729d as the project's first prescriptive candidate at the grain of a single decision — four ordered tests for handling a culture-bound item, every clause traceable to a site in a second-read Tier 1 anchor, i.e. class X1a — frozen before any source text was opened, and it is exactly the shape §3 item 3 says the table has never held.

It was refused on a registered failure criterion, and the criterion is about reproducibility rather than about evidence. Two independent readers applying it to 23 frozen culture-bound item sites agree on the handling it prescribes at 14 of 23 (κ 0.452). The same two readers, applying a deliberately groundless rule of the same shape, agree at 22 of 23 (κ 0.933). RS-20260729d-decision-grain §1. Charter §3 requires a release's recommendations to be followable; this one is not followed the same way twice.

Two further findings belong on this page rather than only on the result.

  1. #3's eight-wide handling set does not survive conversion into a procedure. Across 46 reader prescriptions under C15, only four of the eight handlings were ever returned — RETAIN, EQUIVALENT, SCAFFOLD, CONVERT. CALQUE, GLOSS, SUBSTITUTE and OMIT are unreachable from any branch of the rule. The translator found this from the inside (T-patsyuk-R04-v1 D19, a calque taken outside the rule because no test can return one) and two readers reproduced it from the outside. A prescription derived from #3 is a strictly narrower object than #3, and the narrowing was invisible to its author.
  2. §5b's reliability figures were an artifact of §5b's own null. RS-20260728i reported reader-versus-reader agreement at κ 0.809 / 0.714 and read it as a sound instrument. With no DECIDES anywhere, that was agreement about INFORMS versus NONE. Add one candidate that can decide and the same two models on the same 63 entries fall to κ 0.066. The zero was stable because it was empty.

Two controls exist inside E-20260729d and neither is a candidate: C16 (the sham) and C17 (a positive control that decides by counting words). Any session that finds either in this table should strike it.

6. Standing

provisional: true (charter §2.4). This page is an inventory, not a release: framework/v0.1/ does not exist, and on the evidence sorted above it should not yet. Every count in §3 is recomputable from the table; the coverage figures are RS-20260726e-framework-coverage's and carry that page's stated limits.