Repository path: workshop/experiments/E-20260807d-marking-work/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260807d-marking-work |
| status | frozen |
| created | 2026-08-07 |
| updated | 2026-08-07 |
| senses | style-correspondence, voice, cultural-mediation, accuracy |
| provisional | true |
| internal-judgment-only | false |
| links | wiki/arms/ARM-marking-work.md, framework/v0.1/README.md, wiki/findings/results/RS-20260806e-published-loss.md, wiki/findings/results/RS-20260805h-content-or-marking.md, wiki/findings/results/RS-20260802e-displaced-marking.md, workshop/translations/saigo-no-ikku/R04-v2/translation.md, workshop/translations/saigo-no-ikku/R06-v2/translation.md, workshop/regimes/R04-lead-close.md, config/models.md, config/budget.md, wiki/base/consulted.md |
E-20260807d — where does relational marking do work the scene does not already do?
ARM-marking-work step 1 (T5). Session S128, 2026-08-07 UTC.
Frozen before any dispatch. Nothing below is edited after the first API call except by a numbered amendment, dated, with its reason and the finding that forced it.
1. The question, and why it is this one and not prediction 1 again
framework/v0.1 §3 prediction 1 — on fresh Class A sites in a new pair, R1 recovers a marking at
more than half — has been withheld four times, at S101, S111, S116 and S121–S122, and the
obstacle has moved three times: the ruler, then the census, then the admission condition. The fourth
run's own reading (RS-20260806e §7, written into the release's §3) is that the prediction is
mis-specified rather than unlucky:
A version of prediction 1 that could be discharged has to define its admission on the source's marking and the target's device, not on a reader's recovery.
The reason is now understood exactly. An admission gate defined as the published English fails to
convey the relation selects for sites whose content is relation-neutral, because the sites where
the content states the standing are exactly the sites the gate throws out. RS-20260806e measured
the consequence: King 1914 carries no English device of address at 9 of 9 Kleist sites and is
judged to convey the relation at 8 of 9. The device was gone; the relation was not; the scene had
already said it.
That result contains its own successor question, and framework/v0.1 §8 Q-b now carries it in these
words:
not do translators mark these relations, but where in a text does the marking do work that the scene does not already do?
The Kleist run answered at one site of nine — A2, where the speaker's content works against
the deference his address claims — and one site is a hypothesis, not a finding. This run tests that
hypothesis.
The hypothesis, in one sentence
The marking does work the scene cannot do exactly where the propositional content is DISCORDANT with the standing the grammar claims, and buys nothing where content and standing are CONCORDANT.
If it holds, prediction 1 can be re-worded in a version — 1′ — whose admission condition is
source-side and checkable without a jury: the source marks; the content is discordant with the
marking; the target has a device. That is the deliverable ARM-marking-work exists to reach.
The subject-rule sentence, written before the design
What does this unit teach about translating literature or evaluating translations? — It measures where in a literary text a translator's decision to carry a grammatical mark of social standing changes what a reader can recover, and where it changes nothing; that is a craft rule about when to spend a device, and it is the difference between a framework recommendation a translator can act on and one that fires everywhere and therefore nowhere.
The unit is a principal unit on T5 under the subject rule's named exception: a named deliverable
is blocked. framework/v0.1's only recommendation has a first prediction that four runs could not
discharge, and the release itself says the fix belongs in a version. The arm says so in its first line.
2. The design in one paragraph
Take Ōgai's «Saigo no ikku» (1915), a courtroom story whose entire fabric is Japanese address
marking. Extract every quoted utterance of ten or more characters — 37, by a rule with no
judgment in it. Put all 37 to three seats shown only the Japanese, who say (a) whether the
utterance grammatically marks the speaker→addressee standing and (b) whether its propositional
content is CONCORDANT or DISCORDANT with that standing. Select, by hash, 8 sites per stratum
from those the seats admit. For each, build a minimal pair out of the lead's filed close
rendering: MARKED, the filed span, and BARE, the same span with the address devices deleted and
every proposition held fixed. Grade both, blind to arm, by three seats who have never seen the
Japanese, against relation statements written by a seat shown no English at all. The primary
is an interaction: the MARKED − BARE recovery gap should be large at DISCORDANT sites and near
zero at CONCORDANT ones.
Why an interaction and not a main effect. A main effect is what prediction 1 has been asking for
and is exactly what ceiling on BARE destroys: if the scene conveys the relation anyway, every
absolute rate is high and the difference is small. An interaction is immune to that — it asks where
the difference is, not how big it is — and it fails cleanly if the moderator is not real.
The wire between the limbs, in one sentence. The translation limb renders the story's two official scenes whole and, in a log frozen before any of this existed, registers which utterances the translator believes carry the discordance; the study limb asks independent readers where the discordance actually is and whether the translator's devices bought anything there.
3. Materials
Source. 森鴎外 「最後の一句」 (1915), Aozora Bunko card 45244, 底本 「山椒大夫・高瀬舟」岩波文庫.
The passage is the story's second half whole and continuous — the gate scene, the magistrate's office,
and the white-sand examination — 72 paragraphs, 4,976 Japanese characters, stored at
materials/source-jp.txt and workshop/translations/saigo-no-ikku/span-source.txt.
Why this work. Ōgai names the phenomenon himself, in the passage's last paragraph: 「献身のうちに潜む反抗の鋒」 — the blade of defiance hidden within self-devotion — and says it went into every official in the room. That the phenomenon is present in this text is the author's claim, not the lead's. Which utterances carry it is what is being measured.
Translations, both frozen before this design was written.
T-saigo-no-ikku-R06-v2 (single-pass draft, 2,573 words) and T-saigo-no-ikku-R04-v2 (close
revision, 2,606 words), with the whole translator's log, including §L9's registered prediction.
Contamination. Declared none on T-saigo-no-ikku-R04-v2 §Contamination. The direct measurement
is unreachable — exactly one English translation of this story exists (Dilworth & Rimer 1977, in
copyright, lending-only) — and the artifact says so rather than asserting independence. A same-author
sibling cell measures the lead's Ōgai against Taketomo Torao's public-domain 1918 «Takase Bune»,
length-matched at 525 tokens: 0 shared 7-grams, 0 twelve-grams, longest run 6, against
160 / 27 / 12 / run 24 for an independent published pair. Run before any locus was selected
and after both logs were frozen. The design does not turn on it: both graded arms are the lead's
own English, matched proposition for proposition, so a remembered phrasing sits in both.
The utterance inventory, materials/utterances.json: every quoted span of ≥10 non-whitespace
Japanese characters, in text order, no other filter. 37 of them. The lead applied no relevance
criterion; whether an utterance marks anything is Stage A's question, not the lead's.
4. The frozen BARE construction rule
BARE is MARKED with the devices deleted and nothing else changed. The device lexicon is frozen
here and is taken from the translator's log §L5, which enumerated every device the revision added
before this design existed:
sir·your honour/his honour·if it please your honour·was pleased to/be pleased to/are pleased to·gracious/graciously·the likes of you·I should judge·enquire above(→put the matter to him) ·creature(→girl)
Deletion may adjust punctuation and nothing else. No word is added. No proposition is added or
removed. Where deletion leaves an ungrammatical fragment, the minimal repair is applied and
recorded in materials/pairs.json with the before and after, and that site is flagged for the
parity screen.
A site where MARKED and BARE come out identical is not gradable and is not graded. Its count
is reported as G2, the number of source-marked sites at which a close literary rendering carried
no device at all — which is framework/v0.1 §8 Q-b's own question asked of this translation.
5. Stages, seats, and what each seat may see
Seat roles are bound so that no grading seat has ever seen the Japanese, and no seat grades
against a relation statement it wrote. Slugs from config/models.md, logged as provenance.
| stage | role | seats | sees |
|---|---|---|---|
| 0 | pre-run critic | nvidia/nemotron-3-ultra-550b-a55b (non-panel) |
the design and the Japanese |
| A | source-side classification | A1 P1, A2 P2, A3 P5 |
the Japanese only |
| B | selection | (no call — deterministic) | — |
| C1 | relation statements | P2 |
the Japanese only, no English of any kind |
| C2 | leak screen on the statements | P1 |
the statements and the Japanese |
| D | propositional parity screen | P1, P5 |
the two English spans only |
| E | grading | G1 P3, G2 P4, G3 qwen/qwen3.7-max |
one relation statement + one English span |
| F | recognition probe, after grading | P3, P4 |
the English spans only |
qwen/qwen3.7-max is the declared panel reserve (config/models.md) and is used here in a reader
role, not a juror role — nothing in this run is a judgment of quality, and Tier D is NOT PASSED. It is
used because the panel has five members, three of them are needed on the source side, and a grading
seat that has read the Japanese would recognise the utterance it is grading.
Stage A — the classification, and the exact question put
Each seat receives all 37 utterances with their containing paragraph and a one-line staging note (who speaks, to whom, in what setting) in Japanese and English metadata only — never an English rendering of the utterance. For each it returns:
marks: does the Japanese grammatically mark the standing of speaker to addressee?yes/nodirection:up/down/levelconcordance:CONCORDANTif what the speaker actually says is consistent with the standing the grammar claims;DISCORDANTif the content works against it — a deferential form carrying refusal, accusation, correction, or a claim the speaker's position does not entitle them to; and the mirror for down-address.reason: one clause.
Admission G1: an utterance is admitted iff ≥2 of 3 seats return marks: yes.
Label: the majority concordance among admitting seats. No majority → excluded.
Stage B — selection, deterministic
Among admitted sites that are gradable (MARKED ≠ BARE, §4), order within each stratum by
sha256("E-20260807d|" + site_id).hexdigest() and take the first 8. Fewer than 8 available in a
stratum: take all of them.
Stage C — the relation statements
P2, shown the Japanese utterance, its paragraph, and the staging note, and no English rendering,
writes for each selected site a one-or-two-sentence statement of the relation between speaker and
addressee as the utterance's grammar presents it. P1 then screens each statement, shown the
statement and the Japanese: does the statement restate the utterance's propositional content? A
statement that does is a leak — it would let a grader score content rather than marking — and its site
is dropped from the primary and reported.
Stage D — the propositional parity screen
P1 and P5, shown the two English spans unlabelled and in randomised order, answer: do these
state the same facts? does either assert anything the other does not — about size, age, rank, period,
or anything else? This screen exists because of two prior findings that are the release's own:
RS-20260805h §3 (the adjective predicates where the suffix modulates) and RS-20260806e §7 /
framework/v0.1 §4 (a device can assert a register the no-added-fact clause does not reach). The
screen therefore asks about register by name, not only about facts.
Stage E — grading
Items, per admitted site:
| arm | what it is |
|---|---|
MARKED |
the filed R04 span |
BARE |
the same span, devices deleted (§4) |
POSITIVE |
an explicit English statement of the relation — instrument ceiling |
WRONG |
an explicit English statement of the reversed relation — instrument floor |
Plus REPEAT: at four sites, MARKED appears a second time under a different item id, in a
different block. Items are shuffled by sha256("E-20260807d|items|" + seat_id), split into two
blocks per seat, and carry opaque ids. Each item: the relation statement and one English span, and
the question "Does this English convey the relation stated? YES or NO", plus one clause of reason.
Temperature 0.0.
6. The primary, and every prediction, registered
Let r(arm, s) = the fraction of the three grading seats returning YES for arm at site s, and
gap(s) = r(MARKED, s) − r(BARE, s).
P1 — the interaction (the primary). All three conditions must hold:
mean gap over DISCORDANT−mean gap over CONCORDANT≥ 0.25;- the exact permutation test over all relabelings of sites between the two strata, holding the stratum sizes fixed, gives P ≤ 0.05;
- the instrument controls are clean (
F2,F3below silent).
P2 — the ceiling that killed prediction 1, reproduced under proposition control.
r(BARE) pooled over CONCORDANT sites ≥ 0.70. This predicts that RS-20260806e's 8-of-9 is not
an artifact of published translations being unmatched, and that the scene really does carry the
relation where content agrees.
P3 — the translator's frozen list. The log §L9 names 8 utterances as DISCORDANT before any seat
was called. Prediction: ≥ 6 of those 8 carry a DISCORDANT majority in Stage A. Scored by exact
hypergeometric against the realised base rate; reported with its P whatever it is.
P4 — registered against the lead's own expectation, so it can embarrass him. r(MARKED) at
DISCORDANT sites will be below 0.90 — the marking helps and does not fully carry. If it comes back
at ceiling, the lead over-stated how hard this is.
G2 — descriptive, no threshold. The number of admitted, source-marked sites at which the close
rendering carried no device at all, reported as a fraction of admitted sites.
7. Failure criteria, declared before dispatch
| # | fires when | consequence |
|---|---|---|
F1 |
fewer than 5 gradable sites in either stratum after Stage A | P1 withheld — no power |
F2 |
pooled r(POSITIVE) < 0.80 or pooled r(WRONG) > 0.20 |
P1 withheld — the instrument is not discriminating |
F3 |
REPEAT pairs disagree at > 1 of 4 |
P1 withheld — precision unestablished |
F4 |
> 2 of the selected sites fail the Stage D parity screen at either seat | P1 withheld, figures reported descriptively |
F5 |
Stage A reaches no concordance majority at > 4 of the selected sites, or Fleiss κ over the three seats ≤ 0.0 |
P1 withheld — the moderator is not a property of the text |
F6 |
mean device count per site differs between strata by > 1.0 | reported as a confound beside P1, not withheld — device density is a real covariate and hiding it would be worse than naming it |
F7 |
both Stage F seats name the work and the author from the English alone | reported, does NOT withhold — declared before dispatch. The graders never see the Japanese, the relation statements are source-derived, and MARKED/BARE are equally recognisable, so recognition cannot produce the interaction. It is measured because S123 and S125 established that this project reports it. |
8. What this run cannot establish, written before it runs
- Nothing about quality. Tier D is NOT PASSED. No seat ranks anything; four seats read for a
stated relation and two compare facts. Every sentence of the result is
provisional. - Nothing about human readers.
framework/v0.1§2's standing clause applies unchanged. - One work, one pair, one translator. JA→EN, one hand, one story. A moderator that holds here is a hypothesis elsewhere.
MARKEDis one translator's device set, not the space of available devices. A site where this rendering's device bought nothing is not a site where no device could.- The strata are not randomised, they are observed. Sites are what they are; the design measures an interaction across a classification, not an experiment on an assigned treatment.
9. Spend, pre-flight
Today's ledger (UTC 2026-08-07) stands at $1.377379055 of $5.00 across S125–S127, leaving
$3.622620945. Worst case is built from max_tokens and the worst plausible provider price, per
note (abc).
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| 0 critic | 1 | 16,000 | $0.15 |
| A classification | 3 | 6,000 | $0.13 |
| C1 statements + C2 screen | 2 | 6,000 / 4,000 | $0.09 |
| D parity | 2 | 4,000 | $0.05 |
| E grading | 6 | 6,000 | $0.42 |
| F recognition | 2 | 2,500 | $0.06 |
| subtotal | 16 | $0.90 | |
| re-dispatch allowance (note (bhf) has fired sixteen times) | $0.20 | ||
| declared | $1.10 |
Lead translation is $0 and is never ledgered (charter §3, A4). The two renderings, the whole-story reading, the contamination cell and all analysis cost nothing.
10. Verification
verify.py recomputes every reported number from the raw bodies in runs/, importing nothing
from analyse.py, and recomputes the permutation null by exhaustive enumeration. It runs at
least three mutation tests — perturbing a stored judgement, a stratum label, and a control arm —
and each must change the outcome it is supposed to change. Raw bodies are written before anything is
computed from them and discarded bodies are preserved in runs/discarded/ and never overwritten
(note (bhd)).
11. Amendments, after the pre-run critic and before any other dispatch
Pre-run critic pass 1, nvidia/nemotron-3-ultra-550b-a55b (non-panel), 12,941 characters,
finish_reason: stop, $0.0206148. Verdict NEEDS-REDESIGN, 5 BLOCKING and 7 ADVISORY.
Raw body at runs/critic.txt. All twelve findings are accepted. Two of the critic's proposed
remedies are overruled and replaced, with the reason given. No API call other than the critic's had
been made when these were written.
A1 — the staging notes are removed from Stage A entirely (BLOCKING 1, accepted whole)
The critic is right and the finding would have voided the primary. Stage A's prompt supplied a one-line English note naming the speaker and the addressee — "the gatekeeper to Ichi, a girl of sixteen" — which states the institutional standing the seat was supposed to read off the Japanese grammar. The concordance judgment would then have been content-against-supplied-standing, i.e. a property of the lead's English metadata.
Stage A now shows the Japanese utterance and its containing paragraph and nothing else. The seat
additionally returns speaker and addressee as inferred from the Japanese, which makes the
inference visible and checkable instead of supplied. materials/staging.json is retained in the
repository but is not read by Stage A; it is used only in Stage C's prompt construction, where
the same objection does not arise because a relation statement is supposed to describe the standing.
Correction to that last clause, made in the same pass: Stage C is exactly where a supplied standing
would do the most damage, since the statement is the yardstick. Stage C's prompt also drops the
staging note. staging.json is now read by nothing and is kept only as a record of what was
withdrawn.
A2 — the seats are rebound so no grader has read the Japanese and no seat scores its own statement (BLOCKING 3, accepted whole)
The critic is right: P2 was to classify concordance at Stage A and then write the relation
statements, so the yardstick would have encoded the label the primary is measured against.
| stage | before | after |
|---|---|---|
| A classification | P1, P2, P5 | unchanged |
| C1 relation statements | P2 | P3 — sees the Japanese, did not classify |
| C2 leak screen | P1 | P1 — unchanged, with the residual named below |
| D parity | P1, P5 | unchanged |
| E grading | P3, P4, qwen |
P4, qwen/qwen3.7-max, z-ai/glm-5.2 — none has seen the Japanese or written a statement |
| F recognition | P3, P4 | P4, qwen/qwen3.7-max |
The residual, named rather than hidden. P1 classified at Stage A and also runs the leak screen.
The leak screen asks one binary question — does this English sentence give away what the speaker is
saying? — and P1 never grades and never sees MARKED or BARE. The panel has five members and
this design needs five disjoint source-side roles; this is the one reuse, it is on the least
label-sensitive role, and it is a stated limit of the run.
Two of the three graders are non-panel (qwen/qwen3.7-max, the declared reserve; z-ai/glm-5.2,
probed 2026-07-23 and held in reserve). This is a departure and it is deliberate: a grader who has
read the Japanese would recognise the utterance it is grading, which is a worse defect than model
heterogeneity in a reader role. Nothing here is a jury and Tier D is NOT PASSED. Per-seat YES
rates are reported and P1 is re-run leaving each seat out in turn (A7).
A3 — BARE is built by a frozen substitution table with no repairs (BLOCKING 4, accepted; one remedy overruled)
The critic is right that "minimal repair, recorded afterwards" is an unbounded licence. The rule is
now a frozen, exact-string substitution table (materials/substitutions.json), applied
mechanically by build_pairs.py with no run-time judgment. A site whose substituted text is not
grammatical English is EXCLUDED and counted; nothing is repaired.
Overruled: the critic's remedy that the device lexicon be "defined by an independent Japanese linguist before the lead's translation, or drawn from a published grammar." The lexicon is not a theory of Japanese; it is the list of devices that are actually present in the English being tested, and it was enumerated in the translator's log §L5 before this design existed. No independent annotator can supply the devices of a text they did not write. Determinism is the available fix and it is the one applied.
G2 is reported as a property of the translation, not of the source, in those words.
A4 — the hypothesis is restated in the terms the run actually measures (BLOCKING 5, accepted; one remedy overruled)
The critic is right that §1's wording — work the scene does not already do — names a property of the Japanese scene while the design measures English readers recovering a Japanese-grammar relation from English. The registered hypothesis is now:
English address devices let readers who see no Japanese recover the speaker→addressee standing that Japanese grammar marks, exactly where the utterance's own propositional content is DISCORDANT with that standing, and not where it is CONCORDANT.
And "the scene" is defined: it means the utterance's own propositional content, which is what the graders see. The surrounding narrative is deliberately withheld from graders so that nothing but the utterance's content can carry the relation.
Overruled: the critic's alternative remedy, adding a condition in which graders see the translated surrounding context. It doubles the item count and answers a different question; the design says so rather than quietly not doing it. This is now a named limit in §8.
A5 — the utterance inventory drops its length filter (ADVISORY 8, accepted)
The ≥10-character floor was a lead judgment presented as none. The inventory is now every quoted
span in the copy-text, 56 of them, and marks: no at Stage A is what removes the ones that mark
nothing. Ids U01–U37 are unchanged so that the translator's log §L9, frozen before the critic
ran, still names the same utterances; the nineteen added spans carry S-ids.
A6 — the parity probes are concretised and "fail" is defined (ADVISORY 10, accepted)
Register is replaced by two concrete probes: does either passage imply that the speaker is socially
above, below, or level with the person addressed? and does either passage imply a historical period
or a degree of ceremony the other does not? A site FAILS the parity screen if either seat answers
same_facts: no, or answers anything but neither to the extra-assertion probe. F4 is scored on
that definition.
A7 — two sensitivity analyses are registered now (ADVISORY 6 and 7, accepted)
- Device density. Device counts per site are reported per stratum, and the primary is recomputed
on gap per device as a registered sensitivity.
F6is unchanged. - Leave-one-seat-out. The primary is recomputed three times, each with one grading seat removed.
If the direction of the interaction is not preserved in all three, that is reported beside
P1.
A8 — an underpowered fallback is pre-registered (ADVISORY 11, accepted)
If a stratum lands at 3 or 4 gradable sites, F1 still fires and P1 is still withheld; the
observed difference is reported with an exact permutation P and the words underpowered, not a
result. Below 3, nothing is reported but the counts.
A9 — a recognition sensitivity is pre-registered (ADVISORY 12, accepted)
If F7 fires, the primary is recomputed excluding items whose spans the recognisers named, and
both figures are reported. F7 still does not withhold, for the reason already given.
A10 — the contamination declaration gains the sentence the critic asked for (ADVISORY 9, accepted)
"Unreachable" is not "unread." Added to T-saigo-no-ikku-R04-v2 §Contamination: the lead declares
that it did not consult the Dilworth & Rimer rendering, or any other English rendering of this story,
in any form, at any point in this session — and that this is a declaration about the session, not a
claim about training data, which the sibling cell bounds and nothing can establish directly.
A11 — what P3 licenses is narrowed in advance (BLOCKING 2, accepted in substance)
The critic reads P3 as the lead predicting the lead. It is not: §L9 is scored against three seats
that see only the Japanese, and with A1 applied they are not shown the lead's reading of anything.
But the critic's underlying point stands, so the licence is written down before the number exists:
P3 measures agreement between one translator's frozen reading of his source and independent
readers of that source. It licenses nothing about translation, nothing about English, and nothing
about the primary, and it is reported outside P1's conditions.
A12 — P2's classification cap raised (note (bhq)), in-run
A2 (P2) returned finish_reason: length with 733 content characters and 5,758 of 5,948
completion tokens spent on hidden reasoning, at a cap the other two seats cleared on the identical
input. The cap was the variable, not the seat (note (bhq)); it is raised to 24,000 for this seat
only and the seat is kept. Dead body preserved at runs/discarded/classify-A2.attempt1.json.
A13 — A2 is split into two blocks, in-run
The re-dispatch also returned length, with 48 of 56 objects complete and well formed. The body
is kept as block 1 and only the eight unfinished items are re-dispatched as block 2 at the same
cap. Declared before the tail's content was looked at. A1 and A3 returned stop.
A14 — the leak screen fires at 9 of 12 and the primary is withheld a second time
Declared before Stage D or E was dispatched, and before any graded figure existed.
Stage C2 (P1) returns leak: yes at 9 of the 12 sites, including all four DISCORDANT ones.
Applying §5's rule as written — a statement that restates the content is a leak and its site is
dropped from the primary — leaves 3 CONCORDANT sites and 0 DISCORDANT. P1 is therefore
withheld on this ground as well as on F1, which had already fired at 4 DISCORDANT sites against a
floor of 5.
The rule is applied as written and is not weakened after seeing it fire. That is the move the verification discipline exists to stop.
And the reason it fired is a finding, not an accident. A statement of the relation between two
speakers in dialogue cannot be free of the speech act, because the relation is enacted through
the speech act: a supplicant addressing a high official entails petitioning. The E-20260802e
procedure was built on narrated relations, where a relation can be described without describing what
anybody said. On dialogue the procedure's yardstick and its object cannot be separated, and this
run is the measurement of that. It is the same shape as RS-20260806-same-man's finding that the
propositional parity making a paired rendering a controlled contrast is what readers answer same
person? from.
The graded limb is run anyway, and every figure it produces is DECLARED DESCRIPTIVE before
dispatch. What it can still measure, with leaky statements: whether removing the device changes
recovery at all, within site, against the same statement in both arms; whether the instrument
discriminates on this material (POSITIVE, WRONG, REPEAT); and the S06/S18 minimal case, in
which the entire difference between the two arms is the word sir. No claim of the form "the
marking does work at discordant sites" may be made from this run, and the result page says so.
A15 — amendment A6 made the parity screen unsatisfiable, and the failure is the manipulation check
Declared before Stage E was dispatched.
F4 fires at 12 of 12 pairs, and the reason is an error this design introduced when it accepted
the pre-run critic's ADVISORY 10. A6 replaced the vague word register with the concrete probe
does either passage imply that the speaker is socially above, below, or level with the person
addressed where the other does not? — and that probe IS the treatment. MARKED differs from
BARE in exactly one respect: it marks the speaker's standing. A screen asking whether the two
differ in standing must fail every pair by construction, and it did.
The critic's finding was right and the remedy was wrong, and the remedy was accepted without
checking it against the design's own manipulation. That is the lead's error and it is recorded here
rather than smoothed. F4 is applied as written: P1 is withheld a third time.
The same twelve bodies, read as what they actually are, are the strongest measurement in the run.
Two independent seats, shown the two English passages unlabelled and in randomised order and told
nothing about the study, identify the difference between them as the speaker's social position
relative to the person addressed, at 12 of 12 pairs. On the propositional probe alone — D2's
same_facts — the pairs are judged to state the same facts at 9 of 12, the three exceptions
(P02, P07, P09) being sites where the seat counted the deference itself as a fact. Both seats
name the difference as social standing or deference at 11 of 12 pairs. The stripping rule does exactly and
only what §4 claims it does, and that is a manipulation check passing at ceiling, not a screen
passing.
What follows for the graded limb, declared now: it is run, every figure is descriptive, and the question it can still answer is narrower and still worth the money — given a source-derived statement of the standing, do readers recover that standing better from the marked English than from the bare English, and by how much? A null there would bear on the procedure; a difference would price the device. Neither licenses the hypothesis in §1.
A19 — POST-RUN, and it is a defect of the design rather than of the run
Written at hand-off, after every figure existed. It is labelled post-run for that reason and it changes no number.
§3's contamination paragraph asked whether a published English rendering of this work was reachable
and measured a same-author sibling cell when none was. It never asked whether this project had
already translated the passage itself. It had: T-saigo-no-ikku-R04-v1, S027, 2026-07-27, §4
of this story entire — paragraphs [42]–[72] of this passage — with its own frozen log.
Measured now: the lead against itself, eleven days apart, 218 shared 7-grams, 74 twelve-grams, 35
fifteen-grams, longest run 24 tokens, against 160 / 27 / 12 / 24 for two independent published
translators (materials/selfoverlap.json). Note (bhb) says a lead re-rendering is not an
independent second opinion of anything, and this design's gate was not built to look for it.
Consequences, both recorded on RS-20260807d: the within-site contrast is untouched, because
MARKED and BARE are both drawn from this rendering; §L9's registered prediction is not
independent, because the S027 log analyses five of the eight utterances it names. The present
renderings are re-filed as R04-v2 / R06-v2 and S027's v1 artifacts are intact.