Repository path: framework/traceability-inventory.md · rendered 2026-09-09
Page metadata (front matter)
| type | ledger |
|---|---|
| id | traceability-inventory |
| status | active |
| created | 2026-07-26 |
| updated | 2026-07-29 |
| senses | accuracy, naturalness, voice, style-correspondence, affect, literary-quality, cultural-mediation, purpose-fit, consistency |
| provisional | true |
| links | framework/README.md, wiki/findings/results/RS-20260729d-decision-grain.md, wiki/arms/ARM-decision-grain.md, framework/closure.md, wiki/findings/results/RS-20260728i-coverage-independent.md, wiki/arms/ARM-framework.md, wiki/findings/results/RS-20260727c-arm-identifiability.md, wiki/findings/claims/CL-20260727-arm-identifiability.md, wiki/findings/results/RS-20260726e-framework-coverage.md, wiki/goodness-senses.md, wiki/decisions/resolved/D-20260724-04-pair-relative-sense-weights.md, PROJECT.md |
The traceability inventory — what a framework release could say today, and on what
What this is. ARM-framework's first step (S035), extended at S041: for every candidate operational recommendation the project's evidence could support, the evidence, its class, the language pairs it is evidenced on, and whether it would survive into a release. Built at S035, over 17 result pages, 8 anchors, 10 source pages, 2 theory pages, 1 essay, 1 conjecture and 1 open question. It compiles nothing new; it sorts what exists.
Why it could be built while the release is blocked. Charter A7 gates a release on Tier D. It does not gate taking stock. T5 · Framework had never supplied a principal unit in 34 sessions, and ARM-framework recorded the reason as "deferred was read as nothing to do here."
The one-line answer. The project can supply a vocabulary, a set of diagnostic questions, and a taxonomy of options. It cannot supply a single rule that decides a case. Of fourteen candidate recommendations, one has the shape do X rather than Y addressed to a translator, and it is the one Tier D's failure makes inadmissible — and as of S041 it is inadmissible twice over, on grounds calibration would not fix (§4).
S046 added the fourteenth, and the honest note is that the ratio did not merely fail to move — the pattern behind it is now visible. The new candidate is the fourth consecutive addition that is prescriptive about method rather than about translating. framework/control-arm-spec.md is a real deliverable and it is not framework content in the sense a release needs. What that says about the arm is in §3 item 2, and it bears on how the arm's remaining budget is spent.
1. The evidence classes
Sorting by class rather than by topic is the whole method here, because Tier D's failure does not damage the evidence base uniformly — it destroys one class of it and leaves the rest untouched.
| class | what it is | standing after RS-20260726d-tierD-heldout |
|---|---|---|
| X1a | externally anchored, second-read — a published translation read against its source, or documented reception/scholarship, with an independent check on the reading | intact. Unaffected by jury calibration |
| X1b | externally anchored, single-reader — same kind of object, never independently checked | intact but unverified. The one verification run the project has done retracted one claim and corrected three |
| X2 | machine-measured — a number computed from stored texts, independently recomputed | intact. No judgment involved |
| X3 | panel-scored — rests on jury verdicts | INADMISSIBLE. Tier D NOT PASSED; config/models.md NOT CALIBRATED; no jury verdict carries evidential weight about any translation |
| X4 | internal judgment only — the lead's reading, unanchored | never sole support for a framework recommendation (charter §2.2, and wiki/findings/README.md: "No claim rests solely on internal-judgment-only grounds") |
2. The inventory
| # | candidate recommendation | class | senses | pairs evidenced on | shape | claim page |
|---|---|---|---|---|---|---|
| 1 | Declare the target register; "natural" is not a single corpus | X1a | naturalness, purpose-fit | EN as target, any source | procedural | CL-20260726-register-declaration |
| 2 | Ask what the fluency cost, not whether it is too fluent | X1a | naturalness, cultural-mediation, affect | JA→EN | diagnostic | CL-20260726-fluency-cost |
| 3 | The handling set for culture-bound items is eight-wide; choose knowingly | X1a ×3, X1b ×2 | cultural-mediation, style-correspondence, consistency | JA→EN, RU→EN, EN→FR, JA→JA, OE→EN | taxonomic | CL-20260726-handling-set |
| 4 | Declare the pair; sense weights are pair-relative | X1b synthesis | all | 3-cell design | ratified rule (D-20260724-04) |
(a decision, not a claim) |
| 5 | Grammar-borne meaning transcodes into lexis and loses systematicity | X1a | style-correspondence | RU→EN, JA→EN (EN-target only) | descriptive | CL-20260726-grammar-transcoding |
| 6 | The false friend does its damage inside a drift window | X1b | style-correspondence, accuracy | OE→EN, JA→JA (diachronic) | descriptive (mechanism) | CL-20260726-drift-window |
| 7 | Published translators do not handle a forked class uniformly | X1a ×3, X1b ×2 | consistency, cultural-mediation | 5 pairs | descriptive, normatively unsettled | CL-20260726-forked-class-nonuniformity |
| 8 | Archaism buys inheritability back — foreignisation is a capability, not only a stance | X1b + X1a (3 Tier 2 sources in the original) | style-correspondence, purpose-fit | OE→EN | descriptive | (folded into #6's evidence; not separately claimed — see §4) |
| 9 | Long source periods are split in published practice | X1a, n = 1 | style-correspondence, naturalness | JA→EN | descriptive | (refused — see §4) |
| 10 | Measure the lead's contamination before selecting material | X2 | — | all | prescriptive (method) | CL-20260726-lead-centrality |
| 11 | Check a published pair for dependence before using it as a baseline | X2 | — | LA→EN, DE→EN, RU→EN | prescriptive (method) | CL-20260726-baseline-dependence |
| 12 | One self-revision pass buys naturalness without moving accuracy |
X3 | naturalness, accuracy, style-correspondence | JA→EN | prescriptive (translation) | (refused — inadmissible, and now on two independent grounds; see §4) |
| 13 | Measure whether a comparison's two arms are identifiable without reading, with a sign count rather than a correlation | X2 | — | all | prescriptive (method) | CL-20260727-arm-identifiability |
| 14 | Never pool a paired comparison across length-sign strata — and never read the stratified figure as an estimate below ~22 items per stratum | X2 | — | all | prescriptive (method) | framework/control-arm-spec.md (a specification, not yet a claim page) |
3. What the table says, in four numbers
(Updated 2026-07-27, S041: the table now holds thirteen candidates. §3's four numbers are restated below with the thirteenth included; the diagnosis they support is unchanged and has become sharper.)
A fifth number, added 2026-07-30 (S068) — how far these fourteen rows reach when a third language pair is put to them. RS-20260730i-candidate-reach: 45 PT→EN decision sites, every live rendering written down at the moment of decision, three raters given all fourteen rows with C12's text as written.
- Ten of the fourteen bear on nothing. Only C1 (31 bearing pairs), C3 (9), C2 (2) and C5 (1) reach any decision. C12 bears nowhere, and no method candidate — C10, C11, C13, C14 — bears anywhere, which is §3 item 2 measured a second time on a pair none of them was evidenced against.
- A bearing candidate excludes 0.14 of a live rendering. Not zero, and not one. Five exclusions in 45 sites, every one C1 or C3, every one ruling out exactly one option of four.
DECIDESis 0 of 43 — and 2 of 13 on the same sites with the option list halved. So the inventory's headline zero is a fact about these rows and a fact about how long a translator's option list is;framework/closure.md§1.4 now carries the correction, and the statistic to quote isk, not a rate.- The "does this row bear?" judgment does not reproduce (0.21–0.51 rater agreement on positive cells). Any coverage proportion computed from this table must be quoted with its rater pool.
A sixth number, added 2026-07-31 (S073) — what the fourteen rows do to a translator who is handed them. RS-20260731e-option-census: three independent subjects enumerated live English renderings at 24 loci of a fresh Polish translation, under no preamble, a length-matched sham of platitudes, and these fourteen rows verbatim.
- Handing a translator this table does not change how many renderings they have in play. FRAMEWORK − NONE = +0.292 / −0.667 / +0.417, mean +0.014, against a same-day byte-identical noise floor of 0.375. The registered prediction failed on its floor clause.
- A groundless fourteen-line preamble moved the count further, in two of three seats, and moved it down (−0.375 / −0.875 / +0.083). Whatever a page of advice does to a subject, it is not being done by the content of these rows. This is §5c's C15/C16 finding reproduced on a different instrument and at the level of the whole table rather than of one rule.
- §3's zero is not a lead artifact. No subject in any unprimed or framework-primed cell produced a site with ≤ 2 live renderings — nine of nine cells at 0.000, exactly as the lead's own census reads. The zero survives its first test against non-lead censuses, which is the direction
ARM-candidate-reachtold the successor to be least ready for. - But the option sets themselves do not reproduce: three-way Jaccard 0.170, and each subject recovers 0.22–0.26 of the lead's list.
kis computed over a set three competent subjects do not agree on, which is a different exposure from the one S068 closed.ARM-option-censusstep 2.
A seventh number, added 2026-08-01 (S078) — and it is a number about the instrument rather than about the rows. RS-20260801b-census-author, ARM-option-census step 2, descriptive only: that run's F1 fired.
- Twelve of the fourteen bear on nothing, and one of the two survivors is inert. Over 136 site-codings across six censuses of two Polish works, by majority: C1 bears at 49 sites, C3 at 31, every other candidate nowhere. C1 is the only candidate that ever excludes a rendering — 36 times — while C3 bears 31 times and excludes nothing at all, ever. The three independent measurements now run 7 → 10 → 12, and the newest adds that this evidence base in practice is one register instruction that does work, and one taxonomy that gets named and then does none.
- Whether
ksurvives a change of census author is NOT ESTABLISHED, in either direction, and the arm that asked closedresolvedon that. The run's control block failed to reproduce and its byte-identical noise floor was the size of the effect. §3's reach figures are still one translator's census — which is exactly whatARM-candidate-reachsaid when it closed, and it remains true after two arms and three sessions of trying to change it. framework/closure.md§1.4 now bounds how 0.14 may be quoted: with its evidence base named, never as a constant.
- Thirteen candidates; one is prescriptive about translating; it is inadmissible. #12 is the only entry of the form do X rather than Y addressed to a translator. It rests on the project's only regime comparison (
RS-20260724-selfrevise-first, S010, 25 sessions ago), which is panel-scored on an uncalibrated jury. Tier D's failure does not weaken it — it removes it. The framework's emptiness is not an oversight; it is the measured downstream cost of the calibration gate. - The four prescriptive claims that survive are about running the project, not about translating. #10, #11, #13 and now #14 are X2 — machine-measured, verified, replicated. They are the strongest statements in the repository and not one of them would appear in a release addressed to a translator. S041 added the third and S046 the fourth, and the ratio did not move: every session that strengthens this inventory strengthens the methods column. Four in a row is no longer a coincidence, and the mechanism is not mysterious. The release is gated on Tier D; the only work this arm has that is not gated is method; so the arm produces method.
ARM-frameworkcan go on producing sound X2 method claims indefinitely without getting one step closer to a release, and its two remaining sessions should be spent by someone who knows that. - Nine of thirteen are procedural, diagnostic, taxonomic or descriptive. They give a translator a vocabulary and a set of questions. Under charter §3's requirement that recommendations be "stated so a user (or an autonomous run) can follow them", "here are eight things translators do" is not a recommendation.
- Zero of thirteen are evidenced on French→English, the pair of S035's own translation, and zero are evidenced on the JA→EN pair of S041's beyond #2 and #5. Under
D-20260724-04every one of them would carryuntestedif applied to a pair outside the five in column 5, which is most pairs.
4. Refusals, and they carry information
Three candidates were refused claim status, and the reasons are the inventory's real output.
- #12 self-revision — refused as inadmissible. Not weak: unusable. The distinction matters for what a future session should do about it. Re-running it needs a calibrated jury, so it is downstream of the two rebuilt Tier D controls in
wiki/backlog.md, not of more translating. - Amended 2026-07-27 (S041),
RS-20260727c-arm-identifiability. The sentence above is wrong in one direction and right in another, and the correction changes what a re-run has to fix. #12 now carries a second defect that passing Tier D would not discharge: its temperature control does not support the inference drawn from it. On the TEMP arm the jury preferred the longer text on all five senses (gaps 0.06–0.27, corroborated by S010's own M5 flags on three of them), while temperature moved length randomly and largely (±151 words,A = 0.500exactly). So the pooled ≈0.53 thatRS-20260724-selfrevise-firstcites as "lowering temperature alone moves naturalness only to 0.53" is the average of two opposite-signed length-driven halves and would sit near 0.5 whatever temperature does to prose. A calibrated jury re-run on this design inherits the same uninformative control. The repair is neither calibration nor more translating: it is a control arm matched on length, or analysis within length-sign strata. - Amended 2026-07-28 (S046),
RS-20260728c-length-matching. This amendment closes that prescription rather than extending it: both repairs were priced and neither delivers what S010 wanted. Matching is achievable exactly — 421 words brought to 407 at six sites — but the matched text classifies as 0 REVERT / 4 NEW / 2 DEPART against a 0.0665 coincidence floor, so it is a third text, not a matched arm; a design that matches acquires a third condition rather than a repaired control. Stratifying fails a 0.30 interval-width criterion on 3 of 5 senses at 6 items and 5 of 5 at 10, and the computed requirement is 22–30 items per stratum against S010's 6. So #12 is not repairable by re-analysis at its own n — it needs roughly four to five times the pairs, which is a cost statement rather than a method statement, and it is the first time this page has been able to put a number on what a re-run would take. One finding cuts the other way and is recorded because the lead predicted it wrong: the jury's length response is graded, not sign-only (voicep = 0.011,literary-qualityp = 0.002 over 12 items), which is precisely the condition under which a tolerance-matched control would work. The option S041 wrote off as a fallback is the one the evidence leans toward, and it is unsettled at n = 12. - What was not found, reported because the session was built to find it. The MAIN arms are read-free identifiable (
A = 0.833; the revision came back shorter in 10 of 12 pairs), but the pre-registered rule that would have downgraded #12'snaturalnessresult to confounded did not fire: the best length rule reproduces the jurors' naturalness choices 0.635 of the time against the 0.760 it must match, and the two pairs where the revision got longer were preferred more. The pre-committed minority-sign test is powerless at n = 2. The 0.76 is neither shown to be a length artifact nor shown not to be, and both halves of that sentence are load-bearing. - #9 sentence-splitting — refused as under-evidenced. One published translator, one pair, one text. It is exactly the kind of plausible-sounding process advice a framework accretes without noticing, and the only thing standing between it and a release is that the count was written down.
- #8 archaism-as-capability — refused as not separable. Its Tier 2 side (Schleiermacher's Biegsamkeit, Yan Fu's register argument, Futabatei's 筆力) is strong and was read in the original; its Tier 1 side is one un-second-read anchor.
TH-20260725-capability-conditionsalready ruled that capability is a parameter of the space, not a row in the typology, and the same reasoning applies here: it conditions when a recommendation is available rather than making one.
A fourth is worth naming as a near-miss: CJ-20260724-scaffolded-borrowing is the project's one falsifiable prescriptive bet, it is well-motivated, and both of its testable forms require a jury with sense-localised authority. It cannot be tested until Tier D passes. It stays a conjecture.
5. What would move this page
In order of how much they would change the table, not of how easy they are:
- Repair the two Tier D controls and pass Tier D (
wiki/backlog.md,owed). This is the only thing that converts any X3 evidence into a recommendation, and it unblocksCJ-20260724-scaffolded-borrowing. Amended S041: it no longer unblocks #12 on its own — §4 records a second defect, in the temperature control, that a calibrated jury re-run on the same design would inherit unchanged. - ~~A second regime comparison.~~ BUILT 2026-07-27 (S041), and unscored by design.
R06(lead single pass) was created and frozen,R04frozen at v1.0 with the draft now required to be frozen separately, andT-kusamakura-vii-bath-R06-v1/-R04-v1filed as a pair that is paired by construction — the lead analogue of the R01/R02 property. Sōseki 草枕 ch. VII, JA→EN, 1,956 source characters, both logs frozen, contamination measured at a longest run of 8 against a null-control floor of 5. The scoring is deliberately not done: it waits on a repaired Tier D and on the length-matched control §4 now requires. The durable part — the prose and the frozen decision logs — exists and does not decay. The new item this created: a future scoring of that pair must be designed around its own arm-identifiability figure (#13), which is what S041 pre-registered before translating a word of it. - Second-read the two intralingual anchors (
A-yosano-yomogiu,A-beowulf-ingeld). They carry #6, the project's most quantitative close-reading result, at class X1b. The one verification run the project has ever done moved four claims. ARM-typology-logs— open-coding the frozen translator's logs. §3's finding is that the project's recommendations are not operational; the logs are the only evidence it holds about what translators actually decide, andwiki/program.mdSlate G already says operational recommendations should come from documented decisions rather than plausibility.RS-20260726e-framework-coverageis a single instance of exactly that method and found 11 of 21 decisions uncovered.- A pair with English as source (
ARM-breadth, top C1 priority since S015). #5 rests on two instances that share a target language.
5b. Amendment 2026-07-28 (S051) — the table has now been measured from the other end, and ARM-framework closes on it
RS-20260728i-coverage-independent: two independent readers, 63 logged translation decisions from two language pairs, 126 classifications, DECIDES 0. The readers were given all fourteen rows including #12, with its text as written and no note that it is inadmissible, and neither used DECIDES once. §3 item 1's "one is prescriptive about translating; it is inadmissible" is now joined by a stronger statement that does not depend on admissibility at all: on the 63 decisions this project has on record, #12 is invoked zero times, so reviving it would change nothing at the decision level.
Which rows a real decision reaches. 6 of 14 — #1, #2, #3, #5, #6, #7 — plus #9, a refused candidate, once. Seven were never invoked: #4, #8, #10, #11, #12, #13, #14, four of them the X2 method rows §3 item 2 calls the strongest statements in the repository. #3 alone carries 33 of the 74 invocations, 45%.
And one figure on this page is corrected upward. RS-20260726e's 10-of-21 mapping was independently re-read (κ 0.905 and 0.715), and the single decision both readers agree it got wrong is A3: coverage there is 11 of 21, not 10. §6's citation of the coverage figures carries that correction.
The coverage proportion is not an estimate of anything and this page must stop treating it as one — RS-20260728i §3 withdraws the inference before the run, on the pre-run critic's finding that priming changes which decisions get recorded, not only which get noticed. What survives is the zero.
The closure statement ARM-framework owed is framework/closure.md.
5c. Amendment 2026-07-29 (S056) — a fifteenth candidate was written, tested, and REFUSED
C15 is not in the table above and must not be added. It was written for E-20260729d as the project's first prescriptive candidate at the grain of a single decision — four ordered tests for handling a culture-bound item, every clause traceable to a site in a second-read Tier 1 anchor, i.e. class X1a — frozen before any source text was opened, and it is exactly the shape §3 item 3 says the table has never held.
It was refused on a registered failure criterion, and the criterion is about reproducibility rather than about evidence. Two independent readers applying it to 23 frozen culture-bound item sites agree on the handling it prescribes at 14 of 23 (κ 0.452). The same two readers, applying a deliberately groundless rule of the same shape, agree at 22 of 23 (κ 0.933). RS-20260729d-decision-grain §1. Charter §3 requires a release's recommendations to be followable; this one is not followed the same way twice.
Two further findings belong on this page rather than only on the result.
#3's eight-wide handling set does not survive conversion into a procedure. Across 46 reader prescriptions underC15, only four of the eight handlings were ever returned —RETAIN,EQUIVALENT,SCAFFOLD,CONVERT.CALQUE,GLOSS,SUBSTITUTEandOMITare unreachable from any branch of the rule. The translator found this from the inside (T-patsyuk-R04-v1D19, a calque taken outside the rule because no test can return one) and two readers reproduced it from the outside. A prescription derived from#3is a strictly narrower object than#3, and the narrowing was invisible to its author.- §5b's reliability figures were an artifact of §5b's own null.
RS-20260728ireported reader-versus-reader agreement at κ 0.809 / 0.714 and read it as a sound instrument. With noDECIDESanywhere, that was agreement aboutINFORMSversusNONE. Add one candidate that can decide and the same two models on the same 63 entries fall to κ 0.066. The zero was stable because it was empty.
Two controls exist inside E-20260729d and neither is a candidate: C16 (the sham) and C17 (a positive control that decides by counting words). Any session that finds either in this table should strike it.
6. Standing
provisional: true (charter §2.4). This page is an inventory, not a release: framework/v0.1/ does not exist, and on the evidence sorted above it should not yet. Every count in §3 is recomputable from the table; the coverage figures are RS-20260726e-framework-coverage's and carry that page's stated limits.