Repository path: workshop/experiments/E-20260805e-naturalness-wording/design.md · rendered 2026-09-09
Page metadata (front matter)
Frozen design (v2) — does the failing number belong to the wording or to the prose?
ARM-tierP step 3, the arm's declared closing condition, and the discharge of wiki/backlog.md's
only owed row — condition 4 of D-20260802-13's ratifying vote: re-run Tier D's failed
drop(naturalness) specificity test under the revised wording before any revised-sense verdict is
called calibrated.
v1 was frozen, then put through the independent pre-run critic pass (critic.md, seat x-ai/grok-4.5,
verdict NEEDS-REDESIGN, sixteen findings, eight BLOCKING). This is the amended version: all sixteen
accepted, fourteen in full and two in part with the refused remedy reasoned. A run may proceed only
against this version, and nothing had been dispatched when the amendments were made.
No senses: field: this page designs an evaluation of the instrument, not of a translation, and
asserts no evaluative claim about any translation. The lead's two new references are scored by the
panel, blind and unattributed; the lead judges nothing (charter §5).
0. What this run can and cannot be — read first
- Tier D is NOT PASSED and nothing here changes that. This run re-measures one cell of one condition. It does not dispatch the sham, the held-out arm or the prior positive control, and a specificity condition is one of three.
- It is a re-measurement, not a new instrument. Payloads, jurors, orderings, unit definition,
scoring and the six senses are carried from
E-20260801f; the OLD arm's payloads are byte-identical to that run's stage 4. - The manipulated variable is the whole
naturalnesscriterion string, and it changes in three places at once. §3.2 lists them. No sentence this run produces may attribute an outcome to the struck escape clause, to the target-only frame, or to the named register anchor separately. Everything is attributed to the §3.2 REVISED string as a whole (three inseparable edits). - So the obligation is discharged as a BUNDLE TEST and not otherwise. Condition 4 of
D-20260802-13's vote asks for the test under the revised wording; that is what runs here. A clause-specific discharge — whether the escape clause in particular drives the number — is out of scope and needs the single-edit design recorded in §10. This is the first BLOCKING finding of the critic pass and it is accepted without hedging. - A null is the more interesting outcome and is registered as such (§6, P4p). If the number does not move, the S034/S086 figure is a fact about how O4-damaged prose reads, not about how the question was worded — and that is a claim about translation, not about the apparatus.
1. The subject-rule sentence, written before the unit was designed
What does this unit teach about translating literature or evaluating translations? — It tests whether a reader's sense that an English translation has gone wrong as English can be held apart from the accuracy errors that made it go wrong, and whether that inseparability is a property of the damaged prose or of the words the question was asked in. Two prior runs found that eight accuracy errors in a 300–440-word passage stop reading as eight accuracy errors and start reading as prose that has failed.
The unit is nonetheless method work under wiki/tracks.md's subject rule, and takes the rule's
stated exception rather than pretending otherwise: a named deliverable is blocked. ARM-tierP
cannot write the Tier P statement RS-20260804i §12 makes available while an obligation stands that
forbids calling any revised-sense verdict calibrated or comparing it with a Tier D result. It is the
arm's last step and its declared closing condition.
2. Question
Does drop(naturalness) at the heavy 8-site accuracy dose — the single number on which Tier D
failed at S034 (1.11) and at S086 (1.12) — survive the naturalness criterion string that
D-20260802-13 installed; and does it survive on materials outside the Russian corpus both prior
runs shared?
3. Materials
3.1 Six items, two provenances
| item | source | reference | words | provenance |
|---|---|---|---|---|
T8-GA |
Turgenev, «Свидание» | Garnett | 299 | carried byte-identical from E-20260801f (S086) |
T8-HB |
Turgenev, «Свидание» | Hapgood | 341 | carried byte-identical |
T8-K |
Korolenko, «Сон Макара» | lead | 394 | carried byte-identical |
T8-L |
Korolenko, «Лес шумит» | lead | 397 | carried byte-identical |
N8-A |
Machado de Assis, «O enfermeiro» ¶20–23 | lead, this session | 394 | new |
N8-B |
Machado de Assis, «O enfermeiro» ¶43–46 | lead, this session | 394 | new |
Each item's second text is its reference plus O4 × 8, the accuracy operator reused verbatim from
E-20260725-tierD-ladder §3. materials/build.py records a SHA-256 of every carried string and
re-derives nothing. The two new edit tables are materials/perturbations.md: eight sites each, two of
each of O4's four types, sites numbered in textual order with the order asserted by the build.
The two new items are not independent of each other. Same story, same translator, same hour. They are two passages, not two works, and no figure treats them as independent draws.
3.2 The two criterion strings, verbatim
OLD — exactly the string tools/run_tierD.py put to the jury at S034 and S086:
Reads as fluent, idiomatic English prose WITHIN ITS OWN EVIDENT REGISTER AND PERIOD; nothing rings as translationese unless the source rings strange in the same place. Judge only whether the English is idiomatic on its own terms: do NOT penalise a text for being of an older or a more recent idiom than the other, and do not reward whichever is closer to present-day usage. Threshold competence, not distinction.
REVISED — condensed from wiki/goodness-senses.md §naturalness as it now stands, by the same
practice that produced the OLD string from the entry as it then stood:
Reads as fluent, idiomatic English prose, scored ON THE TARGET TEXT ALONE. This criterion measures DISTANCE FROM TARGET-LANGUAGE NORM — not whether a departure is justified, faithful, purposeful, or good overall. A low score is not by itself a verdict that the translation is bad; it records markedness against the register named here. THE REGISTER ANCHOR FOR THIS ITEM IS:
<anchor>. Judge only whether the English is idiomatic for that register: do NOT penalise a text for being of an older or a more recent idiom than the other, and do not reward whichever is closer to present-day usage. Threshold competence, not distinction.
Three edits, inseparable, and §0.3 binds every sentence about them:
- The source-conditional escape clause is gone — nothing rings as translationese unless the
source rings strange in the same place — struck by
D-20260802-13and replaced on the list by a separate sense,perceived-source-carriage. - A target-only frame is added — scored on the target text alone, distance from target-language norm, not whether a departure is justified, faithful, purposeful, or good overall.
- A register anchor is named rather than left to the juror's discovery of its own evident register and period. The critic's finding 2 is right that this is a second treatment and not a clarification: it can suppress register-uncertainty penalties, pull scores toward the named band, or inflate drops where the label mismatches the prose. The remedy is not to hide it but to name it as edit 3, to take the label out of the lead's hands (§3.3), and to report the effect split by anchor group (§6.2).
Held constant on purpose: the do-not-penalise-idiom-age instruction and threshold competence,
not distinction are carried into REVISED unchanged, so the contrast is not also a contrast between
an instructed item and a bare one. The other five sense definitions are byte-identical across arms,
and analysis/verify.py asserts that the two payloads differ in exactly one line.
3.3 The register anchor is assigned blind, by a third party — stage 0
The v1 rule assigned the anchor from the reference's provenance. The critic's finding 3 is that a date of production is not a register, and it is accepted with the critic's own preferred remedy.
Stage 0, dispatched before anything else. A seat that is neither a juror (P1, P2, P5) nor the
critic (P3) — P4 moonshotai/kimi-k3 — is shown the six undamaged reference texts,
unlabelled, in a fixed order, told nothing about the source, translator, date, experiment or
hypothesis, and asked to assign each exactly one of the three anchors wiki/goodness-senses.md
permits: period-idiomatic, contemporary-vernacular, unmarked/literary-contemporary.
Whatever it returns is used, verbatim, in the REVISED string for that item. It is recorded in
materials/anchors.json with the seat's one-line reason for each, and the raw body is preserved. A
returned label outside the permitted three aborts the run. The lead makes no register judgment
anywhere in this design.
Seat amendment A3. The seat was to be P4 moonshotai/kimi-k3; it returned finish_reason:
length with 7,997 reasoning tokens and empty content at cap 8,000, $0.129576 (preserved as
runs/stage0-anchors.attempt1-P4.json; note (bhf), tenth firing). Raising a $15/M-output cap after a
seat has spent its whole allowance on hidden reasoning is not a fix, and this is a labelling task, so
the seat was changed to mistralai/mistral-medium-3-5 — also probed-but-not-selected, also
neither a juror nor the critic, and not a hidden-reasoning seat. It answered on first dispatch for
$0.006381.
What stage 0 returned, before any juror was paid:
| item | reference | assigned anchor |
|---|---|---|
T8-GA |
Garnett, 1890s | unmarked/literary-contemporary |
T8-HB |
Hapgood, 1903 | unmarked/literary-contemporary |
T8-K |
lead, 2026 | unmarked/literary-contemporary |
T8-L |
lead, 2026 | unmarked/literary-contemporary |
N8-A |
lead, 2026 | unmarked/literary-contemporary |
N8-B |
lead, 2026 | unmarked/literary-contemporary |
Three consequences, all of them recorded here rather than discovered later.
- The anchor is a constant across all six items, so the anchor×item and anchor×cell confounds the critic's findings 3 and 10 identified do not arise in this run. Edit 3 of §3.2 is still an edit — the REVISED string names a register where OLD asked the juror to find one — but it names the same register everywhere.
- §6.2's anchor split is vacuous and is reported as vacuous. The provenance split
{
T8-GA,T8-HB} against {T8-K,T8-L} still runs. - A blind reader shown Garnett's 1890s English and Hapgood's 1903 English does not call either of
them period-idiomatic, on the seat's own reasons ("polished and literary without period or
colloquial markers", "refined and literary, lacking period or vernacular traits"). That is one
seat on one payload and licenses nothing; it is recorded because the design would otherwise have
assigned those two items
period-idiomaticby the provenance rule the critic struck, and the difference between the two rules is not hypothetical.
3.4 Contamination on the new references, measured before the design was written
tools/dependence_check.py, each new reference against Isaac Goldberg's complete 1921 English of the
same story (Brazilian Tales, PG #21040, 4,220 words), reporting the longest run and the shared
7-gram count as CLAUDE.md requires:
| 7-grams | 12-grams | 15-grams | longest run | |
|---|---|---|---|---|
N8-A |
21 | 8 | 5 | 19 tokens |
N8-B |
6 | 0 | 0 | 11 tokens |
The critic's finding 7 is that "the comparison is a text against a damaged copy of itself" blocks
the independence objection and not the measurement objection: overlap can still move a number, above
all if an edit site sits inside or beside the shared run. The check it asks for was run before
dispatch. The 19-token run — "him back to life; it was too late — the aneurism had burst, and the
colonel was dead. I went" — occupies characters 1332–1424 of N8-A's reference. No edit site
intersects it. The nearest is site 7 (I drew back in terror, characters 1228–1249), which ends
83 characters before the run begins; the next nearest is site 8, 419 characters after it.
N8-A and N8-B are reported separately throughout, so that the item carrying 21 shared 7-grams and
the item carrying 6 can be compared rather than pooled silently.
4. Arms, jury, and dispatch
4.1 Twenty-four payloads, 72 calls, plus stage 0
| stage | items | criterion string | payloads | calls | units |
|---|---|---|---|---|---|
| 0 | 6 references, undamaged | — | — | 1 | — |
| 1a | 4 carried | OLD | 8 | 24 | 12 |
| 1b | 4 carried | REVISED | 8 | 24 | 12 |
| 2a | 2 new | OLD | 4 | 12 | 6 |
| 2b | 2 new | REVISED | 4 | 12 | 6 |
A payload is one item in one ordering; a call is one dispatch of a payload to one juror. Order swap is mandatory: every item is dispatched twice with the slots swapped.
Stage order is 1a → 1b → 2a → 2b, so that what truncates first is what can be lost.
The seat probe is the longest payload in the run, not a synthetic one — critic finding 11, and
note (bja): a probe bounds a seat on the payload shape it probed and on no other. T8-L under
REVISED, order 0 — measured at 8,572 characters ≈ 2,143 tokens, the longest of the run's 24
payloads — is dispatched to all three jurors first; its three bodies are real data. If any seat returns an empty body or finish_reason: length there,
the run stops before anything else is dispatched.
Partial-stage rule (critic finding 11). The obligation cell completes only if all 16 carried payloads score. If any is missing, no sentence from §6.1 may be written; the run reports what it has as an incomplete re-measurement.
4.2 Jury
P1, P2, P5, resolved from config/models.md at run time and logged as provenance — the same
three seats as S034 and S086, because the OLD arm is a retest and a retest with a different jury is
not one. P3 and P4 are excluded from the jury exactly as in those runs, which is also what frees them
for stage 0 and the critic pass; this is a power limitation, not a finding about those models.
Six senses, scored 1–7 for each text, plus a forced overall preference, no ties. Strict JSON. Judgment is not parallelized (charter §6): strictly sequential dispatch.
4.3 Unit and cell statistics
The unit is (juror × item), with the two orderings averaged within it. A unit is +1 if the
reference is preferred in both orderings, −1 if the second text is, 0 if split. Under a null
of independent coin-flip preferences P(+1) = P(−1) = 0.25, P(0) = 0.5. Identical to E-20260801f
§6.3.
drop(sense) for a unit is score_ref(sense) − score_var(sense), averaged over the two orderings. A
cell's drop(sense) is the mean over its units. Cells are 12 units (carried) and 6 units
(new).
All exact probabilities are computed by enumeration in analysis/rules.py, which imports nothing
from tools/.
5. Failure criteria, registered before dispatch
| id | criterion | if it fires |
|---|---|---|
| FC1 | Detection must survive the criterion-string change. In each cell the targeted firing rule must fire: ≥ 8 of 12 at +1 and none −1 (carried, exact null P 0.000594) or ≥ 5 of 6 at +1 and none −1 (new, exact null P 0.003174). FC1 is an OVERALL-PREFERENCE rule, not a naturalness-targeted or accuracy-targeted one — a construct mismatch inherited from E-20260801f and named here rather than re-imported silently (critic finding 14). It is therefore accompanied by a reported check: drop(accuracy) > 0 in the same cell |
drop(naturalness) in a cell where FC1 did not fire is uninterpretable and is reported without a verdict. Every row of §6.1 is gated on FC1 firing in BOTH carried cells (critic finding 8). If FC1 fires but the accuracy check does not, the construct mismatch is flagged on the result page |
| FC2 | Spillover, bounded proportionally to this run's own accuracy drop (critic finding 9 — the v1 bound reused the naturalness pass bar with no argument): |drop(accuracy)ᴼᴸᴰ − drop(accuracy)ᴿᴱⱽ| ≤ 0.25 × drop(accuracy)ᴼᴸᴰ on the carried cell. The absolute difference is reported too |
the payload difference is not confined to the naturalness line; the contrast is confounded and is reported as such |
| FC3a | Retest, reported and not a gate (critic findings 4 and 13): |drop_natᴼᴸᴰ, this run − 1.12|, the OLD carried cell against RS-20260802-tierD-verdict §2. This is a statistic, not a threshold. v1 used it as a resolvability gate, which the critic showed is a random throttle: a lucky retest near 1.12 would license almost any wording effect and an unlucky one would suppress a large one |
reported in every case; licenses nothing on its own |
| FC3b | Resolvability, defined entirely within-run. The effect is resolvable iff the permutation test of §6 P2p fires and |effect| ≥ δ_min = 0.25 (frozen here; the same scale as P4p, so the two are complementary) | if it does not hold, nothing is licensed about the criterion string; both cell means are reported |
| FC4 | Ceiling. If the reference's naturalness score is 7 in ≥ 50% of a cell's ratings, drops in that cell are floor-limited |
the cell's drop(naturalness) is reported as a lower bound, not as a value |
FC1–FC4 are evaluated and reported whether or not they fire.
6. Predictions, registered before dispatch
- P1p (primary).
drop(naturalness), REVISED, carried cell, ≤ 0.75 — §6.6 condition 3 would have passed under the revised string. - P2p (the paired test). Per-unit δ =
drop_natᴼᴸᴰ −drop_natᴿᴱⱽ over the 12 carried units. Exact permutation test: under the null that the OLD/REVISED label is exchangeable within a unit, enumerate all 2¹² = 4,096 sign assignments of the observed δ and compute the two-sided P for |δ̄|. P2p fires iff P ≤ 0.05 and δ̄ > 0. Magnitudes are kept and zero-δ units contribute nothing under sign-flipping, which is why this replaces v1's ties-dropped sign test (critic finding 5 showed that test could fire only on near-unanimous patterns). The sign test is retained as a secondary statistic, stated correctly: one-sided exact binomial on positive against negative after dropping ties. - P3p (descriptive by default). The new cell's δ̄ and unit signs are reported. The words reproduces and extends beyond Russian may be used only if FC1 fires in both new cells and ≥ 5 of the 6 non-tied δ are positive (critic finding 12).
- P4p — the registered rival, and the outcome this design exists to be able to report. |effect| ≤ 0.25 on the carried cell: the number does not move. Then the S034/S086 figure is a property of how O4-damaged prose as implemented here reads — including O4's length-changing sites, which are a live alternate path to a naturalness penalty (critic finding 15) — and not of the §3.2 REVISED string.
6.1 The result → claim map, frozen
Every row is void unless FC1 fires in both carried cells and all 16 carried payloads scored. Otherwise the only licensed sentence is uninterpretable under FC1 or incomplete under the partial-stage rule.
| # | this run's OLDᶜ | REVISEDᶜ | FC3b resolvable | what may be written |
|---|---|---|---|---|
| 1 | > 0.75 | ≤ 0.75 | yes | On these stored scores, drop(naturalness) at threshold 0.75 would have decided the specificity condition FAILED at the 8-site dose under the OLD string and PASSED under the §3.2 REVISED string. Attribution is to the three-edit bundle only |
| 2 | > 0.75 | ≤ 0.75 | no | The threshold is crossed by an effect no larger than δ_min or not distinguishable from unit-level noise. Both means reported; nothing licensed about the criterion string |
| 3 | > 0.75 | > 0.75 | yes | The bundle moved the number and did not move it across the threshold. Both means and the effect reported |
| 4 | > 0.75 | > 0.75 | no | Both measured means exceed 0.75 and the OLD–REVISED gap is within this run's own unit-level noise. Not licensed: that the number was shown stable, that it "survives", or anything about the struck clause (critic finding 6) |
| 5 | ≤ 0.75 | either | either | Row 1's second clause is unavailable: this run's own OLD cell did not reproduce the failure, so no sentence may say the condition would have FAILED under the OLD string here. Report FC3a prominently; the run becomes a retest result about drop_nat's stability and says so |
| 6 | any | ≤ 0.75 and |effect| ≤ 0.25 | — | P1p and P4p both hold. The threshold is crossed with essentially no movement between the two strings — i.e. the pass, if any, is not the criterion string's doing. This row overrides rows 1–3 wherever it applies |
Charter §5.5 binds every sentence: on these stored scores, statistic S at threshold T would have decided D.
6.2 Pre-registered exploratory split (critic finding 10)
The wording effect is reported split by assigned anchor group and separately by provenance
{T8-GA, T8-HB} against {T8-K, T8-L}. Stage 0 assigned every item the same anchor, so the
anchor split is vacuous and is reported as vacuous; the provenance split runs. The primary stays
pooled, and §6.1 may not claim homogeneity across groups. This is descriptive: 12 units do not
support an interaction test.
7. Budget
Built from max_tokens, not from an assumed output length (note (abc)).
| juror | S034 measured max completion | max_tokens here |
S086's cap | headroom |
|---|---|---|---|---|
| P1 | 886 | 3,500 | 3,500 | 3.9× |
| P2 | 2,839 | 5,000 | 6,000 | 1.8× |
| P5 | 5,893 | 8,000 | 10,000 | 1.4× |
P2's and P5's caps are lower than S086's, for headroom, and this is declared rather than silent. A
cap cannot change the content of an accepted body; it can only make truncation likelier, and a
truncated body is never scored — it is retried at a raised cap as an amendment written into
config/budget.md before the call (note (bgt)). Worst-case input 2,600 tokens (measured longest
payload 2,143).
| juror | price in / out per M | worst case per call |
|---|---|---|
| P1 | $1.00 / $6.00 | $0.02360 |
| P2 | $1.50 / $7.50 | $0.04140 |
| P5 | $1.65 / $3.30 — worst plausible provider, not list | $0.03069 |
Per payload $0.09569.
| stage | payloads | calls | worst case, reserved |
|---|---|---|---|
| pre-run critic | — | 3 (2 dead) | spent: $0.15947 |
| 0 blind anchors | — | 2 (1 dead) | spent: $0.135957 |
| 1a + 1b carried | 16 | 48 | $1.531 |
| 2a + 2b new | 8 | 24 | $0.766 |
| retry allowance | 2 | 6 | $0.191 |
| total | 24 | 82 | $2.747 |
Against $2.9276 of headroom at session start. The reservation is per stage and is checked
before the stage is entered; a stage that does not fit is not entered and the deferral goes to
NEXT.md. Central estimate from S086's realised per-call mean on this exact task: $0.0054 × 72 ≈
$0.39.
8. Verification
analysis/verify.py recomputes every reported number from the stored raw bodies and asserts,
separately (critic finding 16):
- every carried string against the SHA-256 in
items.json; - every OLD payload against
tools.run_tierD.build_prompton the same item — the retest claim depends on this; - every REVISED payload differing from its OLD twin in exactly one line, that line beginning
- naturalness:, and containing the anchor stage 0 assigned to that item; - every stored body's
prompt_sha256against a rebuild of its payload; - the exact null probabilities by enumeration against closed form, and the permutation test by full enumeration;
- the per-request cost sum.
Then mutation tests: each writes a corrupted copy, asserts the bytes on disk changed, and asserts that a named reported number changes.
9. Known limits, stated before the run
- One cell of one condition. Tier D is not re-run; no verdict on Tier D follows.
- Three edits, one measurement (§3.2, §0.3–0.4). Bundle attribution only.
- The units are not independent. 12 units are 3 jurors × 4 items; a juror with fixed taste contributes correlated units, so exact null probabilities are lower bounds, as at S086.
- The new cell is 6 units on 2 non-independent passages and is anchored by whatever stage 0 assigns; it can support a direction, not a rate, and P3p is descriptive unless its rule fires.
- The retest is at temperature 0.2, not 0.
drop_natis not a fixed property of the payload, which is why FC3a reports and FC3b gates from within-run structure only. - The REVISED string has never been put to a jury on this task. If it behaves oddly, this run cannot distinguish the revision does something from this condensation of the revision does something; the condensation is printed verbatim in §3.2 so the distinction stays available.
- O4 as implemented changes length: the variants run +2.3% to +4.4% longer, and 5–6 of each 8 sites change the word count. Any inseparability claim is about this operator including its surface disruption, not about pure semantic error (critic finding 15).
10. What a clause-specific discharge would need, recorded so it is not lost
The cheapest non-confounded design, from the critic's finding 1: OLD against OLD-minus-the-escape-
clause only — same register phrasing, no anchor token, no target-only frame; one contrast, the same
payloads, 8 payloads / 24 calls on the carried cell. That would answer does the escape clause drive
drop(naturalness)? This run does not ask it.
A process consequence, moved out of §6.1 because it is not a licensed result (critic finding 6):
whatever this run returns, a Tier D design that reports a condition-3 number should state which
naturalness string it was measured under. The two strings now both exist in the repository.