Repository path: framework/README.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | framework-readme |
| status | active |
| created | 2026-07-23 |
| updated | 2026-09-04 |
The framework
The synthesis object (charter §3): a practical, grounded framework for producing and evaluating literary translations, released in explicit versions (v0.1/, v0.2/, …).
framework/v0.3/ EXISTS as of 2026-09-04 (S244), on Tom's direction (PROJECT.md §11): the handbook organised by translation problem, one entry per family, each carrying guidance for a human translator and the same guidance as pipeline steps with the human entry points named. It consolidates v0.1 and v0.2 rather than restating them — read framework/v0.3/README.md for the index, the template and the procedure; framework/v0.3/entries/HB-change-of-language.md is the finished exemplar. v0.2 is frozen as the record at v0.2.52: no new numbered section is added to it; every finding after S243 goes into a v0.3 entry. Written by W1 sessions (wiki/plan.md).
framework/v0.2/ EXISTS as of 2026-08-08 (S133) — the second release, and it is a subtraction: prediction 1 is retired after five attempts in five language pairs, and the statement S1 that replaces it says the loss R1 exists to repair is, on this evidence, uncommon. R1's text is unchanged and no recommendation is added or withdrawn. v0.2 does not restate v0.1 and does not supersede it — read framework/v0.1/README.md first, then framework/v0.2/README.md for what changed. Evidence: RS-20260808b-discordance-fails.
framework/v0.1/ EXISTS as of 2026-08-02 (S091) — the first release, carrying one operational recommendation (R1, displaced marking) and an explicit statement of what it is not evidenced for. The gate was re-read rather than waived: charter A7 requires Tier D to have run and one disciplined regime comparison to have results, and both had (S086, S089). Tier D's failure removes evidence class X3 — anything resting on a jury's judgment of quality — not the release. See framework/v0.1/README.md; the arithmetic of what X3's loss costs is framework/closure.md, which v0.1 does not supersede.
Each release must contain:
- Operational recommendations — what to do, stated so a user (or an autonomous run) can follow them, with human involvement as a designed, parameterized dimension.
- Traceability — every recommendation cites the workshop or poetics evidence supporting it, or is marked
untested. - Predictions — what applying this release to fresh text will yield; later sessions score them, including the losses.
- Stress test — the release applied to a text outside the canon, with the evaluated result.
- Changelog — for every revision, the evidence that forced it.
- Pair declaration + cross-pair
untestedmarking (ratifiedD-20260724-04, 2026-07-25) — each release must declare the source→target pair(s) on which each recommendation is directly evidenced, and markuntestedany recommendation applied to a pair for which it has no direct workshop or precedent evidence (charter §2.1 at pair granularity). A two-axis pair profile (structural × cultural-referential distance, perTH-20260724) is permitted and encouraged as declared, explicitly-provisional descriptive metadata, but a release is not required to weight its recommendations from that profile — that promotion is gated on the revisit trigger in the resolved decision page.
First release gate (restated 2026-07-25, charter A7): not before Tier D detection calibration (charter §5) has run and at least one disciplined regime comparison has results. Tier P does not gate the release; recommendations that depend on ranking near-peer translations carry provisional without it.
Taking stock is not gated, and the stock has been taken (S035). framework/traceability-inventory.md sorts every candidate operational recommendation the project's evidence could support, by evidence class, sense and evidenced pair. The finding a release has to live with: one of twelve candidates is prescriptive about translating, and Tier D's failure makes it inadmissible. The rest are a vocabulary, a taxonomy of options and a set of diagnostic questions. Measured against a real translation task, they covered 11 of 21 decisions and decided none (RS-20260726e-framework-coverage, corrected upward on two independent readers at RS-20260728i §2 and stale here until S056). A fifteenth candidate, prescriptive at the grain of a single decision, was written and refused at S056 — not for want of evidence but because two readers do not apply it the same way (RS-20260729d-decision-grain).
The gate is unmet on BOTH limbs, and saying "blocked on Tier D" was half the picture (2026-07-28, S051, framework/closure.md). Tier D has run and NOT PASSED; and the second condition above — at least one disciplined regime comparison with results — has never been satisfied either. The project's only scored comparison (S010) is panel-scored on an uncalibrated jury with a control whose inference is withdrawn; the two paired R06/R04 comparisons that exist (kusamakura-vii-bath, S041; takasebune, S051) were built and deliberately not scored, because scoring them needs the same jury. The two limbs are not independent: passing Tier D is what would convert two already-built pairs into the missing evidence.
ARM-framework closed resolved at S051 on the second of its two declared endings, with the closure statement at framework/closure.md. Its sharpest number: on 126 independent classifications of 63 real translation decisions, the project's fourteen candidates decided zero (RS-20260728i-coverage-independent), and 6 of the 14 were reached by any decision at all.
Tom deferred the first release on 2026-07-25 — "the framework work release can wait." Nothing is scheduled. Releases are earned by evidence, and the evidence base is being widened first (charter §1, A1; wiki/program.md Slates E–G).