Repository path: workshop/experiments/E-20260730c-revision-close/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260730c-revision-close |
| status | frozen |
| created | 2026-07-30 |
| updated | 2026-07-30 |
| senses | accuracy, naturalness, style-correspondence |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-revision.md, workshop/experiments/E-20260729e-revision-pass/design.md, wiki/findings/results/RS-20260729e-revision-pass.md, wiki/findings/results/RS-20260730-grain-clause.md, workshop/experiments/E-20260730b-c16-redraw/design.md, workshop/translations/niewola-tatarska/R04-v1/translation.md, config/models.md, config/budget.md, wiki/method-notes.md |
E-20260730c — close ARM-revision: rebuild the control that failed, and find out whether the instrument holds still long enough for the rebuild to mean anything
Frozen before the independent pre-run critic pass and before any reader call. The translation
limb was frozen earlier and separately (T-niewola-tatarska-R06-v1 at 14c8989,
T-niewola-tatarska-R04-v1 and its revision log at 73ba25c, contamination figures at 1bf2831).
Amendments made after the critic pass are recorded in §11 with the finding that caused each; nothing
else in this file changes after the freeze commit.
1. What this is
ARM-revision step 2, the arm's last, three items in the order the arm names them:
- (a) rebuild the F3 surface controls and re-run the reader pass, so that P4 — computed at 22.48 at S057 and withheld by F3's registered consequence — becomes reportable or is shown not to be;
- (b) the drift question absorbed from
wiki/backlog.md(T1, opened S050) by amendment A4: does the R04 revision's longest common run with a published comparator exceed the R06 draft's? - (c) report per goodness sense, or record that the axes cannot reach one.
And one thing the arm did not ask for, which the previous session made unavoidable.
RS-20260730-grain-clause (S061) found that a reader instrument of the same family returned
κ 0.452 → 0.737/0.808 on a byte-identical request one day apart. S057's F3 diagnosis and this
session's rebuild are separated by exactly that gap. So the rebuild carries its own cross-day
control, and the control is a gate on attribution rather than a curiosity: see F5.
2. The wire between the limbs, in one sentence
The Polish pair was translated as translation, in session, and it is the first prospective instance of item (b)'s statistic — draft frozen, revision frozen, comparator never opened, run measured afterwards — where all five other pairs in that measurement are retrospective.
3. Materials
The reader batch is S057's batch with FIVE items substituted in place. build_items.py copies
S057's frozen materials/items.json (221 items: 205 real edits over nine R06/R04 pairs, 8 meaning
controls, 8 surface controls), preserves every key and every slot, and replaces the four text
fields of five of the eight surface controls. No reshuffle, no reseed. 216 of 221 items are
byte-identical at byte-identical positions — asserted by build_items.py rather than described —
and that is what makes F5 measurable. (As frozen this said eight substituted and 213 identical; the
redesign of §5 reduced the change. The lower number is what ran.)
Item (b)'s corpus, and n is stated here before any threshold is set — which is the failure A4 withdrew P5 for. Mechanical inclusion rule: a pair enters if (i) its R06 and R04 prose are extractable and (ii) the artifact already records a comparator reachable at a Project Gutenberg ebook id. That rule gives n = 6:
| pair | source language | comparator | Gutenberg |
|---|---|---|---|
bargamot |
Russian | (as recorded on the artifact) | #49598 |
enfermeiro |
Portuguese | Goldberg 1921 | #21040 |
kiseru |
Japanese | Shaw 1930 | #78105 |
kusamakura-vii-bath |
Japanese | Takahashi 1927 | #73131 |
wang-liulang |
Chinese (classical) | Giles | #43628 |
niewola-tatarska |
Polish | Curtin 1898 | #36583 |
Excluded, each with the reason, because a silently truncated corpus reads as a complete one:
bettelweib-locarno — its artifact states no free comparator exists; alfred-preface and
saigo-no-ikku — contamination UNMEASURED, no comparator recorded; takasebune,
yingyi-jiejixing, senilia-clean — comparators exist and are reachable but not at a Gutenberg
id, so they fail clause (ii) and are excluded for uniformity of retrieval rather than for
unavailability; malory-worship — intralingual, and #1252 is its source, not a comparator
translation. Six of thirteen candidate pairs enter. That is the number A4 predicted and it is
arrived at by rule, not by choosing.
4. What is not re-run, stated so no borrowed number sneaks in
The log-versus-diff primary of RS-20260729e §3 (42 of 47) is untouched and is not recomputed here.
The edit list is not rebuilt. The Polish pair is not added to the reader batch — adding it would
break the byte-identity F5 depends on — so it contributes to item (b) and to nothing else.
5. The surface controls — WITHDRAWN AND REBUILT BLIND after the critic pass
§5 as frozen described eight surface controls the lead had built by hand, under a criterion
("match the corpus on mechanical kind") derived from S057's failure scores. The pre-run critic
returned BLOCKING finding 2 against exactly that, and it is right: controls built by an agent that
has seen the failed threshold are selection-confounded by construction, and no amendment repairs
that after the fact. The eight were WITHDRAWN UNSCORED. They remain in git history at 913f5f7.
Full record: critic/dispositions.md.
What runs instead. The surface band is split into two declared bands at the same eight slots, with no reshuffle and no change in batch size:
- band L (n = 5), the corpus-kind band, BUILT BLIND.
build_blind.pyaskedqwen/qwen3.7-max—config/models.md's probed-but-not-selected first reserve, and not a reader, not the readers' declared reserve, and not the critic — for five edits that substitute different English words for the original words while leaving the proposition identical, matched to the corpus's span statistics. It was not told that a gate exists, what the axes are, any threshold, or any S057 number. The registered selection rule, fixed in the script before the call, was the first five well-formed items in the order returned; that is what is used. Slots 56, 96, 108, 158, 174. - band N (n = 3), the non-corpus-kind negative band, KEPT BYTE-IDENTICAL from S057 — two punctuation-only edits and one pure word-order rearrangement. Slots 48, 85, 88.
F3's surface mean is computed on band L only. Band N supplies the within-run contrast gate F7.
The diagnosis, now measured rather than asserted (critic finding 3). §5 as frozen said the rebuild would "match the corpus on mechanical KIND"; the critic showed that claim was false of the lead's own materials. The replacement criterion was computed before the blind build:
Does the revision span introduce a word type absent from the draft span? 184 of the 205 real edits do — 89.8%. Four of S057's eight surface controls do, and all four are two contractions and two misspellings — the same word in another form. Not one of S057's eight surface controls substitutes a different English word for a word. That is the operation nine tenths of the corpus consists of, and its total absence from the control set is a sharper statement of note (bdu) than "too weak".
Span, realised blind (build_items.py), against the target the builder was given:
| draft-side span, median / mean | revision-side, median / mean | |
|---|---|---|
| the 205 real edits | 2.0 / 2.717 | 2.0 / 3.195 |
| S057's surface controls | 2.5 / 4.500 | 2.0 / 4.125 |
| the lead's withdrawn controls | 3.0 / 3.750 | 3.0 / 3.750 |
| band L, built blind | 2.0 / 2.800 | 2.0 / 3.200 |
Matched to within 0.083 and 0.005 of a word, by a builder that could not see a threshold.
The meaning controls are byte-identical to S057's and are not touched. With band N kept, 216 of the 221 items are byte-identical at byte-identical slots.
The residual confound, declared: the builder's prompt was written by a lead that had seen S057's scores. The builder's ignorance is of the scores and thresholds, not of the framing.
6. Procedure
- Freeze this file,
build_items.pyandmaterials/items.json. (This commit.) - Independent pre-run critic pass — P4, the only seat that is neither a reader nor the declared reserve. Findings accepted or declined in writing before any reader call.
- Reader pass: 3 seats × 2 axes = 6 calls, prompts and parameters unchanged from
E-20260729e/run_readers.pyexcept the eight substituted items. - Item (b): re-fetch the six comparators, locate each by heading offset, measure
tools/dependence_check.pyon (R06 vs comparator) and (R04 vs comparator). No comparator prose is displayed at any point (method noteabm). - Analysis, then independent verification recomputing every reported number from the stored raw bodies.
7. The axes and the prompt
Unchanged from E-20260729e §7, including the limitation stated there: readers do not see the
source, so M measures change in what the English asserts, not correctness against the source.
8. Registered gates
- F1 — reader reliability. Krippendorff's α (interval) across the three readers ≥ 0.60 on each axis. (S057: 0.822 / 0.855.)
- F2 — range use. Each reader spans ≥ 40 points on each axis; ≥ 10% of real edits score ≥ 50 on M in at least one reader. (S057: the span criterion passed, the 10% clause FIRED at 0.020 / 0.054 / 0.000. It is expected to fire again; it is registered unchanged so that the comparison is like-for-like, and firing again is a reproduction datum rather than a surprise.)
- F3 — the positive control, four criteria, unchanged from
E-20260729e§8 as amended by A3. M(meaning) − M(band L) ≥ 30 · E(meaning) ≤ 30 · E(band L) ≥ 20 · E(band L) − E(meaning) ≥ 15. The last two are the two that failed. Registered consequence, unchanged: if F3 does not pass, no M/E comparison is reported at all and P4 stays withheld for a second session. - F4 — not applicable; no log matching is re-run.
- F5 — THE ATTRIBUTION GATE, new, tightened by critic finding 1 (BLOCKING).
Over the 216 items that are byte-identical at byte-identical slots (205 real edits, 8 meaning
controls, band N's 3), for each of the six (seat, axis) cells compare today's scores with S057's:
mean absolute difference, Pearson r, the change in cell mean, and the per-item Δ
distribution (finding 5).
Criterion: mean |Δ| ≤ 3 points AND Pearson r ≥ 0.90, in every one of the six cells.
As frozen this read ≤ 10 and ≥ 0.80. The critic's arithmetic: E(surface) must move 12.87 → ≥ 20,
a gap of 7.13, so a 10-point tolerance let the gate pass in precisely the case it existed to
catch.
And attribution is a comparison, not a threshold (A1): any movement in E(surface) that does
not exceed the observed cross-day mean |Δ| on the 216 identical items is reported as
UNATTRIBUTABLE, regardless of F5's verdict.
If F5 fails, the session may not claim that rebuilding the controls changed F3, item (a) is
unanswerable with the instruments this project has, and
ARM-revision'sretiredending becomes live on that ground rather than on F3's. The failure is confounded two ways and the design claims neither over the other (A5): cross-day drift, OR within-batch context effects from the five substituted items. - F7 — the kind contrast, new, from critic finding 7. mean E(band L) − E(band N) ≥ 10. If the non-corpus-kind band scores as high as the corpus-kind band, the substitution hypothesis is refuted and any F3 pass licenses only the weaker claim "any surface perturbation now scores ≥ 20".
- F6 — the diagnosis check. Note (bdu) says a control weaker than the material cannot calibrate it, and S057 measured the corpus's median real-edit E at 27.33. That figure is recomputed today from the re-run. If the corpus's own median E has moved by more than 10 points, note (bdu)'s supporting number is itself day-dependent and the note is reported as narrowed.
9. Predictions, registered
- Q1 — the blind-built band L passes F3's E side. mean E(band L) ≥ 20 and E(band L) − E(meaning) ≥ 15. n = 5, and the n is part of the prediction. Failure means the E axis does not respond to meaning-preserving lexical substitution at the corpus's own kind and magnitude, which would be a finding about the axis and not about the controls.
- Q2 — F5 holds in all six cells. Stated with its prior: S061 found a same-family instrument
unstable across one day on a 9-label categorical task with 23 items. This is a 0–100 graded
task over 221 items, and
RS-20260729b/RS-20260729eboth found graded axes far more reliable between readers than categorical ones. The prediction is that graded axes are also more stable across days. A failure is the more consequential outcome and would say that S057's α figures, this session's F3 verdict, and every future cross-session comparison on this instrument are all drawn from a moving instrument. - Q3 — P4. mean E − mean M over the 205 real edits ≥ 20. S057 computed 22.48 and could not report it. Reportable only if F3 passes — and, per critic finding 6 (A6), if F5 fails, P4 is reported as a within-session descriptive statistic ONLY, S057's 22.48 stays withheld permanently, and item (a) is recorded as unanswerable. The two gates were jointly inconsistent as frozen: item (a)'s stated purpose was to make 22.48 reportable, while F5 could declare its cross-session meaning unclaimable.
- Q4 — item (b): a null. The revision does not systematically move the text toward the
published comparator. The test and its power are fixed here, before the numbers: a two-sided
sign test over the non-tied pairs of {run(R04) > run(R06)}. With n = 6, only 6 of 6 in one
direction reaches p ≤ 0.05 (exact p = 0.031); 5 of 6 gives p = 0.219. So this test can detect
only a near-deterministic effect, and every other outcome is reported as a null with the count
printed. That is stated in advance because A4 withdrew P5 precisely for having a threshold that
could not fail informatively. Secondary, continuous, and equally underpowered at n = 6: the
per-pair change in shared 7-gram count, draft → revision, reported with its sign pattern and no
significance claim.
Prior in the same direction:
RS-20260729h§8 measured a different statistic (n = 5 shared-type rate) over sixsenilia-cleanpairs and found away on 5 of 6, toward on 0 of 6. - Q5 — item (c): the axes cannot reach a goodness sense, and the reason is structural. Both axes
measure change; every sense in
wiki/goodness-senses.mdis evaluative. Projecting a change measure onto an evaluative sense requires a quality judgment, and Tier D has not passed. The prediction is that (c) is discharged by a written impossibility plus the descriptive breakdown that is available — per-pair and per-source-language E and M — not by a per-sense number. Registered as a prediction so that producing a per-sense number anyway would count as refuting it.
10. Panel bindings and pre-flight
| role | seat | why |
|---|---|---|
| pre-run critic | P4 moonshotai/kimi-k3 |
neither a reader nor the declared reserve; the S053 role-collision fix, tenth session running |
| readers ×3 | P1, P3, P5 | S057's seats exactly — F5 requires the same seats or it measures nothing |
| surface-control builder | qwen/qwen3.7-max |
added by amendment A2. Probed-but-not-selected in config/models.md; not a reader, not the readers' reserve, not the critic. A materials-construction role, not a judging role |
| declared reserve | P2 | as S057, declared here and not chosen later (notes (b), (bdl)) |
Pre-flight, built from max_tokens and not from an assumed output length (note abc), at the
corrected P1 price:
- actuals for the two calls already made, added at the amendment commit: critic $0.0876726
(in 5,398 / out 4,771,
Together, 122 s); blind builder $0.0406038 (in 822 / out 8,902,Alibaba, 166 s,max_tokens4,000 raised to accommodate hidden reasoning). The estimate below is the one frozen before either. - critic 1 call at
max_tokens16,000: out 16,000 × $15.00/M = $0.240, in ≈10,000 × $3.00/M = $0.030 → $0.270 - readers 6 calls at
max_tokens12,000, prompt ≈15,000 tokens: P1 2 × (0.090 + 0.019) = $0.218; P3 2 × (0.072 + 0.030) = $0.204; P5 priced at the worst plausible provider ($1.65/$3.30 per M, the measured S022 case, not list) 2 × (0.040 + 0.025) = $0.130 → $0.552 - reserves, if two fire: 2 × ($0.090 + $0.023) = $0.226
Worst case ≈ $1.05. Every output-dominated run in this ledger has landed at 15–34% of worst case. Headroom at this freeze: $4.3973500561 (today's $5.00 less S061's $0.5664084814 and this session's gate at $0.0362414625).
Free, and never ledgered: the Polish translation and both its logs, all six contamination measurements, the item rebuild, every analysis and the verifier.
11. Amendments after the pre-run critic pass
moonshotai/kimi-k3 (P4), one call, stop, in 5,398 / out 4,771, provider Together,
$0.0876726. Verdict NEEDS-REDESIGN — two BLOCKING, four MANDATORY, one ADVISORY. All seven
accepted, none declined. Note (rr), twentieth consecutive session, and the first
NEEDS-REDESIGN this project has taken since S021.
Amendments A1–A7, each with the finding that caused it and what it cost, are in
critic/dispositions.md. In summary: A1 tightened F5 from ≤10/≥0.80 to ≤3/≥0.90 and made
attribution a comparison against the measured drift rather than a threshold; A2 withdrew the
lead's eight hand-built surface controls unscored and replaced them with five built blind by a
seat that saw no score; A3 replaced the "match on mechanical kind" criterion with the measured
new-word-type criterion (89.8% of real edits, 0 of 8 S057 surface controls); A4 matched span
blind to within 0.083 of a word; A5 made F5's failure two-way confounded in writing;
A6 registered that a failed F5 makes P4 within-session-only and item (a) unanswerable;
A7 added band N and gate F7.
A design that loses its own control set to its own critic is the mechanism working. Nothing else in this file changed, and no reader call had been dispatched when these were applied.
12. What this design cannot establish
- Nothing about quality. Tier D has not passed; the axes are named for change, not repair.
- Nothing about correctness against the source, because readers do not see the source.
- Nothing general about translators. Every edit in the corpus was made by the lead
(
CL-20260726-lead-centrality), and item (b)'s six pairs are six passages by one translator. - Nothing that separates cross-day drift from batch-context effects, if F5 fails (A5).
- Nothing about the direction of the drift statistic beyond a near-deterministic effect, at n = 6 (Q4).