Repository path: workshop/experiments/E-20260729e-revision-pass/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260729e-revision-pass |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | — |
| internal-judgment-only | true |
| links | wiki/arms/ARM-revision.md, workshop/regimes/R04-lead-close.md, workshop/regimes/R06-lead-single-pass.md, config/models.md, config/budget.md, wiki/method-notes.md |
E-20260729e — what the second pass does, and whether the log knows
Frozen before the held-out translation was begun and before any API call was dispatched.
ARM-revision step 1. Amendments made after the independent pre-run critic pass are recorded in
§10 with the finding that caused each; nothing else in this file changed after the freeze commit.
1. Question
Between a frozen R06 draft and the R04 revision of it, what changes, and how much of
what changes appears in the translator's log?
2. Why it is answerable now and was not before
R04 v1.0 §Procedure 2a (S041) made every R04 run produce a paired R06 output. Eight pairs
now exist and no session has ever read them as a corpus — corrected per amendment A2, which
struck this sentence's original "none has been read": this session read the artifacts to build
the extraction rules and the diff, and computed their edit rates, before freezing this file. What
it had not done, and what the design's hygiene actually rests on, is any M/E scoring, any
log-matching or any interpretive reading of the edits. Three carry an explicit revision log — a contemporaneous
account, written by a lead that could not anticipate this measurement, of what the second pass
did. That makes the log checkable against the text for the first time.
3. Materials
The eight existing pairs, frozen in earlier sessions, all status: frozen:
| work | source language | draft ¶ / words | revision ¶ / words | revision log? |
|---|---|---|---|---|
alfred-preface |
Old English | 4 / 938 | 4 / 944 | no (log is not pass-separated) |
bargamot |
Russian | 17 / 673 | 17 / 688 | no |
bettelweib-locarno |
German | 3 / 407 | 3 / 421 | no |
kusamakura-vii-bath |
Japanese | 12 / 1,210 | 12 / 1,206 | yes — D25–D36 |
saigo-no-ikku |
Japanese | 31 / 1,216 | 31 / 1,219 | no |
takasebune |
Japanese | 9 / 895 | 9 / 905 | yes — D28–D42 |
wang-liulang |
Chinese (classical) | 4 / 1,875 | 4 / 1,879 | yes — D10–D22 |
yingyi-jiejixing |
Chinese (modern) | 7 / 1,532 | 7 / 1,561 | no |
Paragraph counts match pair-for-pair in all eight, which is the extraction's own check.
The held-out ninth pair, translated after this file is frozen: Machado de Assis,
«O enfermeiro» (Várias Histórias, 1896), Portuguese → English — the project's first
Portuguese. Source: pt.wikisource Página:Várias histórias.djvu/175–183, single witness,
pagequality level 1 (not proofread), which is declared as a limitation on the artifact.
Comparator for the contamination gate: Isaac Goldberg, Brazilian Tales (Four Seas, 1921),
Gutenberg #21040, located by heading offset only, no prose displayed (method note abm).
4. The mechanical edit list
extract.py pulls the translated prose out of each artifact under a per-work rule written by
hand after inspecting every file; build_edits.py diffs the paired paragraphs token-wise.
An edit is a maximal run of non-matching tokens, where two non-matching runs separated by GAP = 2 or fewer identical tokens are merged into one. GAP is declared here and the GAP = 0 ("raw") count is reported alongside every merged count, so the merge rule's effect is visible rather than assumed. Context: CTX = 12 tokens either side.
Frozen result over the eight pairs: 192 edits (292 raw) at the freeze commit a012e6f, and
191 edits (289 raw) after amendment A7 repaired the tokenizer. Both numbers are on the
record; the first is in git history, the second is what the reader pass runs on. punctuation-only
edits: 8 of 191.
5. Procedure
- Freeze this design and the 192-edit list. (Done at the freeze commit.)
- Independent pre-run critic pass — one non-lead, non-rater model. Findings accepted or declined in writing before any other call.
- Translate the held-out pair. R06 draft written straight through and frozen as its own
commit; contamination gate run on unit A before unit B is drafted (the standing selection
gate,
CLAUDE.md); then the R04 self-revision, with a revision log written under R04 §3. The lead's log is frozen before the diff over its own pair is computed. - Match logs to edits. For each pair carrying a revision log, each logged decision is matched to zero or more mechanical edits by the verbatim strings the log itself quotes. The matching rule is frozen in §6.
- Reader pass. Three non-lead models score every edit on two graded axes (§7), blind to whether the edit is logged, blind to the work, and blind to this design.
- Analysis and independent verification —
verify.pyrecomputes every reported number from the stored raw bodies and imports nothing from the analysis script.
6. The log-matching rule, frozen
This section as frozen was wrong, and amendment A5 replaces it. It read: a decision matches if its quoted draft form is a substring of the edit's draft span with context, or its quoted revised form is a substring of the edit's revision span with context. Two defects — the critic found the second, the implementation found the first.
- Containment runs the other way. A log quotes a whole phrase ("moisten the spring in
secret" → "moisten the springtime in secret") and the mechanical span is the minimal
difference inside it (
spring→springtime). The rule as written matched 2 of 12 decisions in a log that itemises ten changes. oradmits a declared non-change. A term the revision kept still appears in the draft text, so a one-sided test scores it as a logged change.
The rule that runs, in match_logs.py: a decision matches an edit when quoted material
identifies both sides of it — each side either aligned with a quote by containment in either
direction (both strings ≥ 2 tokens), or, where the changed span is a single token, grown outward
with its own unchanged context until it reaches 2 tokens and found inside a quote, or exactly
equal to a single-token quote. Multi-edit matches: every matched edit is logged, the decision
counted once.
A decision with no match is declared-non-change if its own text says the draft was kept
(kept, unchanged, not changed, stands, untouched, left alone, no change), and
unmatched-claim otherwise. An edit matched by no decision is unlogged.
Every quoted form is taken from the log verbatim; the lead does not paraphrase to make a match. Any decision requiring a judgment call to match is listed by id in the results and counted separately, and gate F4 fires on their number.
7. The two graded axes
Categorical labels are used here only where a graded axis cannot replace them, on this
project's own measurement: RS-20260729b-graded-drift (S054) moved three-rater agreement from
α = 0.51 to 0.78 / 0.89 by splitting one categorical label into two graded axes on the same
items and the same raters, and backlog note (bdm) asks whether the project's other
categorical schemes carry the same defect. This design answers that question by not building
one.
Each reader scores each edit 0–100 on:
- M — meaning. How much does the revision change what the passage says? 0 = it asserts exactly the same thing; 100 = it asserts something materially different.
- E — English. How much does the revision change how the English reads — rhythm, register, idiom — setting aside what it says? 0 = indistinguishable in reading; 100 = a completely different-sounding piece of English.
Readers see the draft span, the revision span and the surrounding English context. They do not see the source text. That is a deliberate limitation and it bounds what M means: M measures change in what the English asserts, not correctness against the source. No reader here can say whether an edit repaired a mistranslation; the claim this design can support is about what moved, not about what was fixed. Stated in advance so it is not softened afterwards.
8. Registered gates — every one fires before any count is reported
- F1 — reader reliability. Krippendorff's α (interval) across the three readers must be ≥ 0.60 on each axis. Below it, no reader-based count from that axis is reportable as an estimate. (S054's graded axes reached 0.78 and 0.89; S056's categorical pass fell to κ 0.066 once a label became reachable.)
- F2 — range use (note bdq, S056). Each reader must span ≥ 40 points on each axis across the corpus, and ≥ 10% of edits must score ≥ 50 on M in at least one reader. An axis on which nothing fires measures the scheme's silence, not the corpus.
- F3 — the positive control, and the session must not rest on reading a number the instrument
may be unable to produce. Sixteen synthetic items are shuffled into the reader batch and are
not identified as controls: 8 meaning-controls (a negation removed, a number changed, a
named agent swapped, a quantifier reversed) built to score high on M and low on E, and 8
surface-controls (contraction, British/US spelling, comma placement, an adverb moved) built
to score low on M and moderate on E. Criterion: mean M(meaning-controls) − mean M(surface-
controls) ≥ 30, and mean E(surface-controls) ≥ 20. If the controls do not separate, the axes
are not measuring what they are named for, no M/E comparison is reported at all, and the
arm's own §Done-when clause about closing
retiredbecomes live. - F4 — matching auditability. If more than 20% of logged decisions need a judgment call to match an edit, the log-vs-diff primary is reported as bounded by lead judgment rather than as a measurement.
9. Predictions, registered
Written before the ninth pair existed and before any call.
- P1 — the log records the substantive edits. Mean M of
loggededits exceeds mean M ofunloggededits by ≥ 15 points. If P1 fails, the translator's log is not a size-ordered sample of what the pass did. - ~~P2~~ WITHDRAWN as a prediction by A1 and computed as an observation, because its truth value was available from frozen materials at registration time. Observed, on the three retrospective pairs: unlogged fractions 0.000, 0.000 and 0.364 — 29 of 33 mechanical edits are accounted for by the log, and two of the three logs account for every one. The prediction survives only for the held-out pair, where it is genuinely unknown.
- P3 — knowing changes it. The held-out pair, whose log is written by a lead that knows the
count is coming, has a lower
unloggedfraction than the mean of the three retrospective pairs. This is confounded by construction and the confound is the point: a null says the gap is structural rather than a matter of attention. - P4 — the second pass is mostly not about meaning. Mean E exceeds mean M across the whole corpus by ≥ 20 points.
- ~~P5 — the drift question~~ WITHDRAWN by A4 and scheduled as
ARM-revisionstep 2. Its denominator was never stated, so its threshold could not fail informatively, and the comparator texts are not stored in this repository by charter §7.5. It is a unit of work, not a sub-part of one.
Not a prediction, and recorded so it is not presented as one: the per-pair edit rate is already known to the design, because the edit list was built before this file was written. It ranges from 0.6 to 5.9 edits per 100 draft words — a factor of ten. That is an observation made before any prediction and it is reported as such.
10. Panel bindings, cost, and the pre-run critic
Roles per config/models.md; slugs are provenance, logged in the run records.
| role | panel seat | why |
|---|---|---|
| pre-run critic | P4 | a subject in nothing here; P1/P3/P5 are the readers, and a reader may not critique the instrument it is about to be |
| readers ×3 | P1, P3, P5 | the RS-20260729b-graded-drift configuration, which is the only graded-axis rater set this project has a reliability figure for |
Pre-flight estimate, built from max_tokens and not from an assumed output length (note
abc): critic 1 call at max_tokens 16,000 (note bdl — P4 failed at 6,000 in S044 and
returned cleanly at 16,000 in S055) ≈ $0.29 worst case; readers 6 calls (3 readers × 2 axes,
separate calls per axis on RS-20260729b's own finding that one call for two axes bleeds them)
at max_tokens 12,000 ≈ $0.83 worst case. Worst case ≈ $1.12; every output-dominated run
in this ledger has landed at 15–34% of worst case. Today's headroom at the freeze: $3.909597.
Free, and never ledgered: the translation, both contamination gates, the extraction, the diff, the log matching, the drift measurement and every analysis.
10b. Amendments after the pre-run critic pass
moonshotai/kimi-k3 (P4), one call, stop, in 4,287 / out 3,633, provider Fireworks,
$0.101034. Verdict NEEDS-AMENDMENT, seven findings — four BLOCKING, three MANDATORY. All seven
accepted; one sub-clause of finding 5 declined in writing. Note (rr), fifteenth consecutive
session.
- A1 (finding 1, BLOCKING) — P2 was not a prediction. "The unlogged fraction is > 0 in each pair carrying a revision log" was computable from frozen materials at registration time. P2 is withdrawn as a prediction and computed as an observation (§9); the prediction survives only for the held-out pair, where it is genuinely unknown.
- A2 (finding 2, BLOCKING) — a false sentence about this design's own hygiene. §2 said the eight pairs had never been read; §4 says the extraction rules were written after inspecting every file. The accurate statement, and the one that now stands: the artifacts were read to build the extraction rules and the diff, and their edit rates were computed, before this file was written; no M/E scoring, no log-matching and no interpretive reading of the edits had occurred. The blindness this design actually claims is the readers' and the held-out pair's, not the lead's.
- A3 (finding 3, BLOCKING) — F3 could pass on a single undifferentiated axis. The gate tested only that M separates and that E fires at all, so three readers scoring one "how big is the change" dimension would pass it and P4 would then be reported on axes never shown to be distinct. F3 gains the critic's two criteria verbatim, unaltered: mean E(meaning-controls) ≤ 30, and mean E(surface-controls) − mean E(meaning-controls) ≥ 15.
- A4 (finding 4, BLOCKING) — P5's denominator was never stated, so the gate could not fail
informatively. P5 is WITHDRAWN from this design and scheduled as
ARM-revisionstep 2. The comparator texts are correctly not stored in this repository (charter §7.5 forbids storing a copyrighted text whole), so running P5 means re-fetching and re-aligning six comparators — a unit of work, not a sub-part of one. The backlog row it discharges (T1, S050, age 7) is therefore discharged by scheduling into an arm, not by a rushed n = 2. (The critic's secondary worry, that method noteabmforbids the measurement, does not hold:tools/dependence_check.pycomputes a longest common run without displaying comparator prose, which is how every such figure in this project has been produced.) - A5 (finding 5, MANDATORY) — the matching rule was wrong in the direction the critic named and
in one it did not. Accepted: match against the changed span, not the context window; freeze a
multi-edit rule (every matched edit
logged, the decision counted once). Declined in writing: the clause putting the declared-non-change keyword test first. Run first on these logs it classifies 6 of 12, 6 of 15 and 6 of 13 decisions as non-changes, becausekept,keepsandholdsoccur inside descriptions of real changes ("The revision keeps the interiority and gives up the possessive"); it is therefore applied only to decisions that matched nothing. And a defect the critic did not find, which the implementation did: containment runs the other way — a log quotes a whole phrase and the mechanical span is the minimal difference inside it — so the rule as frozen matched 2 of 12 decisions in a log that plainly itemises ten changes. The final rule is inmatch_logs.py's docstring and requires both sides of an edit to be identified by quoted material, with short spans grown by their own unchanged context. This rule was developed against the three retrospective logs and is therefore tuned to them. The held-out pair is where it is applied untuned, and that is now one of the reasons the held-out pair exists. - A6 (finding 6, MANDATORY) — P3 is re-registered as asymmetric. Only the null licenses a conclusion (the gap is structural, not a matter of attention). A lower unlogged fraction on the ninth pair is uninterpretable by construction — foreknowledge, a first-ever source language, an unproofread witness and n = 1 variance are not separable — and will be reported as uninterpretable, never as a confirmation.
- A7 (finding 7, MANDATORY) — the tokenizer shattered words, and the corpus is re-frozen.
\w+|[^\w\s]split intra-word hyphens and apostrophes; the critic saw the signature in the sample it was shown. Taking the deterministic of the two fixes it offered, words are now kept whole and each edit carries a mechanicalkind. Effect, reported rather than assumed: 191 edits (289 raw), against 192 (292) before — one edit, and the artifact rate is small.punctuation-onlyedits are 8 of 191 (4.2%), under the 25% threshold registered with this amendment, so primaries are reported on the full set with the lexical-only set as a check. Such edits are not dropped: a?becoming a.is a revision decision.
What the amendments cost this design, stated plainly: one of its five predictions is withdrawn (A4), one is downgraded to an observation (A1), one is made one-directional (A6), and its central matching rule had to be rebuilt (A5). Two predictions survive as registered — P1 and P4.
11. What this design cannot establish
- Nothing about quality. Tier D has not passed. No sentence here says a revision improved anything, and the axes are named for change, not for repair.
- Nothing about correctness against the source, because readers do not see the source (§7).
- Nothing general about translators. Every edit in the corpus was made by the lead, and
CL-20260726-lead-centralitycarries that limit. The backlog row "A translator's log that is not the lead's" (T1/T5, S051) is the thing that would fix it and this design does not. - Nothing about pairs whose logs are not pass-separated. Five of the eight artifacts write one log across both passes, so their edits can be counted but not matched. They enter the M/E analysis and are excluded from the log-vs-diff primary, which is a smaller denominator than the corpus and is reported as such.