Repository path: workshop/experiments/E-20260804g-yardstick-repair/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260804g-yardstick-repair |
| status | frozen |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| senses | style-correspondence, accuracy, voice, cultural-mediation, naturalness |
| provisional | true |
| links | wiki/arms/ARM-r1-warrant.md, framework/v0.1/README.md, wiki/findings/results/RS-20260802e-displaced-marking.md, wiki/findings/results/RS-20260804-yardstick.md, workshop/experiments/E-20260802e-displaced-marking/design.md, workshop/experiments/E-20260802e-displaced-marking/materials/sites.md, config/models.md, config/budget.md, wiki/goodness-senses.md |
E-20260804g — the same six sites, and somebody else says what the source conveys
ARM-r1-warrant step 1 (T5). Frozen before any seat is addressed and before any new English is
written. framework/v0.1 §8 Q-c names this run's shape in advance; §5 below follows it and declares
the one place it departs from it.
1. What is being asked
framework/v0.1 contains one recommendation, R1:
Where the source marks a relation or attitude by a grammatical form the target lacks, the absence of a same-category counterpart is not the absence of the marking. Render the site again under a brief that requires the marking to appear, and let it fall wherever the target does mark such things. Record the loss only if that second attempt fails.
Its machine-read warrant is one figure: at E-20260802e, three independent seats recovered the
relation from the FORCED rendering at 5 of 6 sites and from the FROZEN (filed) rendering
at 0 of 6. The seats scored each rendering against a written statement of the relation, and the
lead wrote that statement — the same lead who had written all five texts in play and knew which
was which.
S101 put the same procedure on FR→EN with the relation statements written by a seat that had been
shown no English, and the separation vanished: FROZEN 0.810, FORCED 0.738, an independent seat's
plain translation also 0.810 (RS-20260804-yardstick). That run differed from E-20260802e in
language, in sites and in the translator's state of knowledge as well, so it refutes nothing. This
run differs in one thing.
The question: does R1's separation survive a yardstick the lead did not write?
2. The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5)
What does this unit teach about translating literature? — Whether the one piece of advice this project gives a translator does what it claims: when a source marks something with a grammatical form English has no counterpart for, does re-rendering the site so the marking falls on a device English does have actually put the relation into the English — judged against a statement of what the source conveys written by somebody who never saw a translation of it.
The yardstick's authorship is the control. The reason it matters is itself a translation proposition and not an apparatus one: a translator who declares a loss should not also be the one who says what was lost. This unit is a principal unit rather than method work under the subject rule's named exception — a published figure in the project's only framework release may be false, and it is the only figure supporting its only recommendation.
3. Materials — what is held byte-identical, and what is not
Held byte-identical to E-20260802e/materials/renderings.json (sha256
58f0912bc27184c94180496a18674cbff5b8a9761032a6109876291a452f9cbd, copied to
materials/sites-frozen.json, checked by the verifier):
| field | what it is |
|---|---|
source |
the source passage at each of the six sites |
gloss |
the literal gloss |
before / after |
the surrounding context, identical across all conditions so only the span varies |
spans.FROZEN |
the filed rendering, verbatim from its translation page |
spans.FORCED |
the 2026-08-02 re-rendering under R1's brief |
spans.DECOY |
the 2026-08-02 length-matched unmarked control |
Replaced, and this is the variable:
| field | 2026-08-02 | here |
|---|---|---|
relation |
written by the lead | written by A1, a seat shown the source, the gloss and the construction under study and no English rendering of any kind |
spans.POSITIVE |
an explicit gloss of the lead's relation | an explicit gloss of A1's relation, written by the lead from it — framework/v0.1 §8 Q-c names the ordering as the mistake E-20260804 paid F1 for |
The six sites, unchanged: S1 Akutagawa 「煙管」 JA→EN (stacked humbling auxiliaries) · S2 Turgenev
«Роза» RU→EN (formal вы between lovers) · S3 Garshin «Сигнал» RU→EN (dative of experience) ·
S4 Korolenko «Сон Макара» RU→EN (affectionate diminutives) · S5 Pu Songling 《王六郎》 LZH→EN (humble
1sg 僕) · S6 Bécquer «El rayo de luna» ES→EN (grammatical gender). Their provenance and the frozen
log claim at each are E-20260802e/materials/sites.md, unchanged and not re-derived here.
Contamination. Carried forward verbatim from E-20260802e §3 with its reasoning, because the
material has not changed: no published comparator is opened and none is needed — the claim under
test is existential (does a rendering that carries the relation exist), and a rendering that
reproduces a published translator's device is still an existence proof. Per CLAUDE.md's standing
rule the absence of a reachable comparator is declared on the artifact rather than passed over. Note
(bhb) cuts toward the null here and is the design's friend: the lead matches itself at up to 37
contiguous tokens across sessions, so a new rendering that merely reproduces an old one marks
nothing and counts against R1.
4. Stage A — the yardstick, written by somebody who has seen no English
A1 = P5 deepseek/deepseek-v4-pro, a seat used in no other role in this run. It receives, per
site: the source passage, the literal gloss, and a plain-language identification of the construction
under study ("the verb ending in …", "the pronoun the speakers use for each other"). It receives
no English rendering of any kind — not FROZEN, not FORCED, not the 2026-08-02 relation
statement, not the log.
It returns, per site, one sentence stating what the source passage conveys to a reader beyond what
a plain rendering of its words would say — the same object E-20260802e/materials/sites.md calls
the relation.
A2 = qwen/qwen3.7-max, the panel's documented first reserve (config/models.md), receives the
byte-identical prompt independently. A2's statements are not graded. They exist for two reasons,
both declared now:
- A stability measurement. Is "what this passage conveys" a stable object across two
independent readers, or does each see a different thing? Coded by the lead against the frozen
three-way rule in §7 (
SAME/OVERLAPPING/DIFFERENTreferent) and committed before any grading seat is addressed — the commit order is the guarantee, and the verifier checks it. - A declared fallback. P5 has returned
finish_reason: lengthon four occasions in the last two sessions. If A1 fails both attempts, A2's statements become the yardstick and the run says so on its face. That rule is written here, before dispatch, so it cannot be chosen after seeing anything.
The leak screen (mandatory, mechanical). A relation statement that names the source's
grammatical device tells the graders the answer. Every A1 statement is passed through a frozen
banned-lexicon check (analysis/screen.py, committed with this design): pronoun, plural, singular,
formal, informal, familiar, diminutive, suffix, prefix, dative, accusative, nominative, case,
declension, gender, feminine, masculine, honorific, auxiliary, humble, particle, inflection,
conjugation, grammar, grammatical, verb form, tense, mood, subjunctive, agreement. A statement that
fires is re-requested once, with the screen's rule quoted and nothing else changed; a statement
that fires twice puts its site outside the denominator, named. The screen is run and its result
committed before the graders are addressed.
5. The one declared departure from "unchanged in every other respect"
framework/v0.1 §8 Q-c asks for the six sites "unchanged in every other respect." This design adds
a fifth arm, and the reason is a hole the repair would otherwise leave open.
REAIM — the lead renders each site again, in session, from the source and A1's statement, under R1's own brief: the relation must appear, and it may fall on any device English marks such things with — a pronoun, a verb, an address noun, a courtesy formula, an adjective, the argument structure of the clause. No metalinguistic commentary; that is what POSITIVE is for, and a REAIM rendering that states the relation rather than carrying it is a failed REAIM.
Why the arm is necessary. FORCED was written in 2026-08-02 to repair what the translator's log
said was lost. RS-20260804-yardstick §4 found that at 3 of 8 sites the independent reader's
statement of what the passage conveys is not the thing the log recorded as lost. Where that
happens here, FORCED is aimed at a target the graders are not scoring against, and a low FORCED
score would be uninterpretable — it would not distinguish "R1's move cannot reach this relation"
from "the 2026-08-02 rendering was aimed elsewhere." REAIM is the arm that separates those.
What it costs, stated rather than hidden. The graders see five renderings per site where
E-20260802e's saw four. The primary comparison (§7 R2/R3) is stated on FROZEN, FORCED and DECOY,
whose bytes did not move; the arm count did, and DECOY is the control that detects it — if DECOY
rises above its 2026-08-02 floor by more than F1's bar, the grading task itself has changed and the
run is descriptive only.
REAIM is written by the lead knowing the statement the graders will score against. That is a real advantage over FROZEN, which was written to no test at all — and it is the same advantage FORCED had in 2026-08-02. REAIM is therefore never reported as evidence that R1 beats a filed translation; it is reported as evidence about whether the relation A1 names is reachable in English at all by ordinary narrative means.
6. Stage B — grading
Identical in form to E-20260802e §4, and the runner is that experiment's, with the condition list
extended by one:
- Three non-Anthropic seats — P1
openai/gpt-5.6-terra, P2google/gemini-3.6-flash, P3x-ai/grok-4.5. The same three as 2026-08-02. - Each receives the source, the gloss, A1's relation statement, and the five renderings unlabelled, and answers for each: would an English-speaking reader with no access to the source come away with the RELATION — YES or NO, plus one line of reason.
- Presentation order per site is fixed by
sha256(site id | seat id | ordering), never chosen by the lead. Two orderings per seat, the second the reverse of the first — six bodies. - A seat counts as recovering the relation only if it answers YES in both of its orderings.
Fixed here, before the data, exactly as at
E-20260802eandE-20260804. - A site counts as recovered under a condition when ≥ 2 of 3 seats recover it.
7. Predictions — registered, six, before any call
Every one is stated on the same scale as E-20260802e so the two runs are directly comparable.
| # | prediction | 2026-08-02 value |
|---|---|---|
| R1p (primary) | FORCED recovered at ≥ 4 of 6 sites | 5 of 6 |
| R2p (the one Q-c doubts) | FROZEN recovered at ≤ 1 of 6 sites | 0 of 6 |
| R3p (the separation itself) | (sites FORCED recovered) − (sites FROZEN recovered) ≥ 4 | +5 |
| R4p | DECOY recovered at ≤ 1 of 6 sites | 0 of 6 |
| R5p | POSITIVE recovered at 6 of 6 sites | 6 of 6 |
| R6p | REAIM recovered at ≥ 4 of 6 sites | — (new arm) |
The lead's stated expectation, recorded so it can be wrong: R2p and R3p FAIL. On the S101 precedent I expect FROZEN to rise to 3 or more of 6 once somebody else says what the passage conveys, and the separation to shrink below 4. I expect R1p to hold, S6 to fail for FORCED as it did before, and R6p to hold at 5 or 6 — because REAIM is written to A1's own words. If R2p and R3p hold, R1's warrant survives the hardest test available to this project and I was wrong.
The stability coding rule, frozen now (§4, A2): for each site, the lead codes A1's statement against A2's as SAME (the two name the same thing about the passage), OVERLAPPING (each names something the other's wording contains), or DIFFERENT (a reader given one and a reader given the other would look for different things in the English). Coded and committed before any grading body is dispatched. This is a measurement, not a gate.
Also measured, and it is free: for each site, whether A1's statement names the thing the frozen
translator's log recorded as lost. Same three-way coding, same commit, same rule. E-20260804 found
3 of 8 divergent; this is the second measurement of that quantity.
8. Failure criteria — pre-committed
- F1 — the grading instrument failed. POSITIVE recovered at fewer than 5 of 6 sites, or
DECOY recovered at more than 1 site. → the run is descriptive only; nothing enters, amends or
is withdrawn from
framework/v0.1§2 on it, whatever R1p–R3p say. - F2 — the yardstick leaked. A relation statement that fires the §4 screen twice → that site is outside every denominator, and every figure on this page is stated over the surviving sites with the count named.
- F3 — the yardstick failed to arrive. If A1 fails both attempts, A2 governs (§4) and the run says so; if both fail, no grading seat is addressed and the run reports a dispatch failure and nothing else.
- F4 — seats. Fewer than three grading seats returning both orderings → the primary is
descriptive only and the shortfall is named. A
finish_reason: lengthbody is a seat failure and never a partial answer (note (b)). - A null is a result, and here it is the expensive one. If R2p and R3p fail,
framework/v0.1§2's traceability paragraph is qualified in place: R1's machine-read warrant was measured against a yardstick its own author wrote and does not reproduce when somebody else writes it, and the release then carries two human-anchored sites and no machine-read separation. That sentence goes into the release and the changelog, and R1's text is not thereby refuted — it is left standing with less under it, which is the honest reading.
9. What this design cannot show
- Nothing about quality. Tier D is NOT PASSED; evidence class X3 is inadmissible. This run
measures availability — whether a relation is recoverable from a piece of English.
RS-20260802e§4.1 already measured that the recovered marking costs about two-thirds of a naturalness point. - Nothing about human readers. Three language models are not a survey (charter §4). A YES licenses three independent readers recovered the relation from this English. Tom is never an experimental subject (charter §9).
- Nothing from a sample. Six sites are a census of the Class A population this project has identified, not a draw from anything, so no sampling inference is available and none is drawn.
- Nothing that separates language from yardstick across the two runs. This design fixes the
sites and varies the yardstick;
E-20260804varied language, sites and yardstick together. A disagreement between this run and 2026-08-02 is attributable to the yardstick within these six sites, and to nothing wider. - One leak this design does not close. The
glossis held byte-identical and it is lead-written; A1 reads the passage through it. A gloss that already hints at the relation would seed the statement. Holding it fixed is what makes the comparison clean, and the leak is a limit, not a control.
10. Verification
analysis/verify.py, importing nothing from tools/: asserts materials/sites-frozen.json equal by
sha256 to E-20260802e/materials/renderings.json; asserts every FROZEN, FORCED and DECOY span used
in a dispatched prompt is byte-identical to that file; asserts every source string verbatim against
its stored source file; recomputes the hash-derived rotations; recomputes all six predictions and all
four failure criteria from the raw stored bodies; recomputes the leak screen; asserts the stability
coding was committed before the first grading dispatch (by git commit order); and re-sums the billed
cost from the stored bodies. At least three mutation tests, each asserting the bytes on disk
actually changed (note (bgu)) and restoring them after (note (bhd)).
11. Budget
Declared worst case $0.90, built from max_tokens and never from expected output (note (abc)),
with the S022 routing margin.
| stage | calls | max_tokens |
list rate out | worst case |
|---|---|---|---|---|
pre-run critic (P4 moonshotai/kimi-k3) |
1 | 12,000 | $15.00/M | $0.18 |
| stage A — A1 (P5) and A2 (qwen), 1 each, 1 retry reserve each | 2 (+2) | 12,000 | $0.87 / $4.425 | $0.13 |
| stage B — 3 seats × 2 orderings | 6 | 12,000 | $6.00–$7.50/M | $0.49 |
| retry reserve | — | — | — | $0.10 |
Today's UTC headroom before this run: $3.331026740 of the $5.00 cap, five sessions already run. Lead translation is free and is never ledgered (charter §3, A4): REAIM, POSITIVE, the coding and all verification are lead work at no API cost.
A pricing correction read from the API today, recorded as a gate and not as a finding.
openai/gpt-5.6-terra returns $1.00 / $6.00 per M, against the $1.25 / $7.50 config/models.md
has carried since S061. The revisit trigger does not fire — the move is downward on the frontier
seat, which is the conservative direction for every estimate built from it — and the table is
corrected in place. x-ai/grok-4.5 reads back $2.00 / $6.00, unchanged.
12. Amendment A0 — the critic seat, mid-run and declared
P4 moonshotai/kimi-k3 was dispatched as pre-run critic and did not return a critique. Attempt 1
returned a non-JSON body — 12,221 bytes of keep-alive whitespace, note (bgc) — and was retried
free. Attempt 2 returned finish_reason: length with 12,000 completion tokens, all of them
reasoning, and zero characters of content, billed at $0.2128698 (provider Together, 240s).
Under note (b) a length body is a seat failure and never a partial answer.
Note (bhf) — measure a seat's appetite before depending on it — fires for the fifth recorded time, on the same slug, and for the second time on capacity rather than price.
The seat moves to z-ai/glm-5.2 (config/models.md, probed 2026-07-23, not selected; a
reserve), max_tokens 24,000, used in no other role in this run. It returned in 125s for
$0.03770906 — one sixth of what the failure cost — with ten findings, two BLOCKING, and
VERDICT: NEEDS-REDESIGN.
One contamination this session declares rather than hides. While diagnosing the P4 failure the lead read the tail of the rejected body's hidden reasoning, which contained partly-formed criticism of the leak screen. Nothing from it is adopted here as a finding, and it is not counted as a critic pass; but the lead had seen it before writing the amendments below, and one of them (A4) addresses a defect that reasoning had also reached. It is recorded because a rejected body that was read is not the same as a rejected body.
13. Amendments A1–A8, from the pre-run critic (critic.md)
VERDICT: NEEDS-REDESIGN. Ten findings, two BLOCKING. All ten accepted. Applied before any
subject seat was addressed. §§4–11 above are superseded where they conflict with what follows.
A1 (F1, BLOCKING) — REAIM leaves the primary grading task
The critic's finding: a fifth arm that is a near-guaranteed YES anchors the grader and can make them stricter on the subtler arms, inflating the separation the run exists to measure — and DECOY is a coarse post-hoc gate that declares the run descriptive after the damage is done rather than preventing it. It is right, and the fix it proposes (drop REAIM) is right about the primary.
REAIM is removed from stage B and becomes its own stage.
- Stage B grades {FROZEN, FORCED, DECOY, POSITIVE} — four arms, the same four as 2026-08-02, in the same template with the same hash rule. The primary is now byte-identical to the original in arm count and arm composition as well as in span content.
- Stage C grades {FROZEN, REAIM, DECOY, POSITIVE} — four arms, FORCED swapped for REAIM, same seats, same template, same hash rule, two orderings. Dispatched after stage B.
Stage C is not a replication of stage B: the same three models see FROZEN, DECOY and POSITIVE twice under the same yardstick, so agreement between the stages is a consistency figure and never an independent second measurement. It is reported as such.
A2 (F2, BLOCKING) — the yardstick seats get no English at all
The critic's finding, with its own examples, is the most serious thing it found: the frozen glosses already state the relation. S4's says little-dear-earth and little-wretched-yurts; S5's says your-servant; S6's says daughters rather than sons and feminine; S2's says you[plural] twice; S1's says graciously permitted to be allowed. An A1 handed those has not written an independent yardstick — it has paraphrased the lead's.
Worse, and the design missed it: at S2 the literal gloss and the FROZEN rendering are the same English sentence. Giving A1 that gloss would have shown A1 the filed rendering while the design's own §4 promised it would see none.
So A1 and A2 receive no English rendering or gloss of any kind. Their whole input is:
- the source passage, verbatim;
- a staging note — bare narrative fact about who is speaking to whom and what is happening, with no characterisation of the utterance ("A river ghost is speaking to his friend, a fisherman");
- a bare pointer: the source tokens under study, quoted and nothing else —
拝見させて頂きたい,вы / Посмотрите,страшно Семёну,землицы / юртёнки,仆 / 僕,hijas. The §4 pointers that described the construction in words ("the word the speaker uses to refer to himself") are struck: that one names the relation.
This is stronger independence than E-20260804 had, where the yardstick seat was given the
gloss. The graders' materials do not move: they still receive the frozen gloss byte-identical.
The risk this buys, named rather than argued away: a seat reading Literary Chinese or Japanese with no crib may misread the passage. The panel's language competence was screened at S015 (≥5/6 in Russian, French and Japanese; P5 specifically missed a Chinese polysemy on that screen). A2's independent statements are the check: where A1 and A2 describe the same passage differently, the run says so and does not treat A1's statement as established.
A3 (F4) — a lose condition, pre-committed
The critic is right that the lead's registered expectation of a null converts every outcome into a win. An inconclusive band is therefore fixed now, before any data. Let d = (sites where FORCED is recovered) − (sites where FROZEN is recovered); 2026-08-02's d was +5.
| d | reading | consequence for framework/v0.1 |
|---|---|---|
| ≥ 4 | the separation reproduces under an independent yardstick | §2 records the reproduction |
| ≤ 1 | it does not reproduce | §2's machine-read paragraph is qualified in place |
| 2 or 3 | INCONCLUSIVE | nothing is amended, in either direction |
In the inconclusive band this session buys no decision and says so on its face, in the result,
in framework/v0.1's changelog as a non-entry, and in the arm's log. That is the losing outcome and
it is now reachable.
A4 (F3, F5) — two gates the design could not fire
- F1 gains a third leg (F1c), the too-easy floor. If all four stage-B arms are recovered at ≥ 5 of 6 sites, the task discriminates nothing and the run is descriptive only, whatever DECOY did. The critic's F5 is right that POSITIVE — written by the lead from A1's own sentence — can always be made explicit enough to pass, so F1's POSITIVE leg alone was close to unfireable.
- The leak screen is retained and demoted. The critic's F3 is right that a word-list tests vocabulary and not content: "the speakers use the distant language of strangers with each other" passes the screen and still hands over the answer. The screen stays as a floor, not a guarantee, and the run reports the statements verbatim so a reader can judge the leak themselves. No lead-coded predictability rubric is added, because F8 is right that a lead-coded check on the lead's own conflict is not a check.
A5 (F6) — R6p gains a consequence
If REAIM is recovered at fewer than 4 of 6 sites, R1's text is qualified and not only its
warrant. R1 asserts that a rendering carrying the relation exists at such sites; REAIM is the lead
attempting exactly that with the target handed to him explicitly and no metalinguistic escape. A
failure there is evidence about the recommendation itself, and framework/v0.1 §2 records it.
A6 (F8, F9) — the auxiliary codings are demoted, and the confound they cannot answer is declared
- The A1/A2 stability coding and the A1-vs-log divergence coding are lead-coded, and the critic is
right that a lead-coded measurement carries the conflict this run exists to repair. They are
therefore
internal-judgment-onlydiagnostics, still committed before grading so the ordering is verifiable, and no conclusion rests on either. The critic's alternative — a third seat coding them — is not bought: it would put a fresh uncalibrated instrument under a figure that is not load-bearing. - F9 is the finding this design cannot fully answer and it is declared as a limit, not solved. A null could mean the yardstick's authorship matters, or it could mean A1 simply writes vaguer statements than the lead did, so that graders say YES to more renderings. The DECOY gate is the partial answer — statements vague enough to pass FROZEN should also pass a rendering that marks nothing, and R4p/F1 catch that — but the band in which A1's statements are vague enough for FROZEN and precise enough to exclude DECOY is not excluded by anything here. A mechanical specificity comparison (word count and content-word count, A1's statements against the lead's 2026-08-02 ones) is reported and is a proxy, not a control.
A7 (F7) — what the two-ordering rule actually buys
Ordering 1 is the hash-derived order and ordering 0 its reverse, so the both-orderings rule detects
sequence sensitivity only. A seat systematically disposed to answer YES passes it. This is
inherited from E-20260802e unchanged and deliberately, because changing the template would destroy
the comparison the run exists to make. It is a limit of both runs equally and is stated in both.
A8 (F10) — what the amendment to the release may and may not say
The critic is right that a null from three models on six sites, with a yardstick whose independence this design improved but cannot certify, is thin ground for withdrawing anything. So the consequence is bounded in advance:
- The amendment states what was measured and under what conditions — the separation reported at
RS-20260802edid not reproduce on the same six sites when the relation statements were written by a seat shown no English — and the gloss-seeding, specificity and model-count limits alongside it. - It does not say R1 is refuted, and it does not withdraw R1's text; the two human-anchored sites (Dole 1896, Hertzberg 1886) are untouched by anything here.
- In the inconclusive band it says nothing at all (A3).
Budget, re-declared after the amendments
| stage | calls | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic — spent | 2 (1 billed failure + 1 accepted) | 12,000 / 24,000 | $0.2505789 actual |
| stage A — A1, A2, one retry reserve each | 2 (+2) | 12,000 | $0.15 |
| stage B — 3 seats × 2 orderings | 6 | 6,000 (P2: 16,000) | $0.50 |
| stage C — 3 seats × 2 orderings | 6 | 6,000 (P2: 16,000) | $0.50 |
| retry reserve | — | — | $0.15 |
Declared worst case $0.90 → $1.60, against $3.331026740 of UTC headroom at session open. The increase is the critic's: one billed seat failure, and one stage added on its BLOCKING finding.
14. Amendment A9 — what stage A actually returned, written before any grader was addressed
Both yardstick seats returned on the first dispatch, $0.0036418896 (A1, deepseek) and $0.0200364 (A2, qwen). Neither had any English and both read all six passages, including the Literary Chinese, without a crib — the misreading risk A2 was bought against did not materialise.
A9.1 — F2 fires. S1 leaves every denominator, and the run is on FIVE sites
A1's pass-1 statements fired the leak screen at S1 (humble), S2 (formal) and S4 (diminutive). The three were re-requested once, per the frozen rule, with the banned list quoted and nothing else changed. S2 and S4 came back clean. S1 fired again on the same word.
So S1 — Akutagawa, the stacked humbling auxiliaries — is outside every denominator on this page.
It is the pre-committed rule doing what it was written to do, and it costs the run a site where
2026-08-02 had FORCED YES and FROZEN NO. On the five surviving sites, E-20260802e's own figures
are FORCED 4 of 5, FROZEN 0 of 5, DECOY 0 of 5, POSITIVE 5 of 5, d = +4. Those are the numbers
this run is compared against, and they are recomputed from the stored 2026-08-02 bodies by the
verifier, not quoted from the result page.
A1 could not describe S1's relation without naming the machinery, twice. That is recorded as an observation and nothing is built on it.
A9.2 — the thresholds, restated for five sites, before any grading dispatch
Scaled from §7, keeping each threshold's meaning rather than its integer:
| # | on 6 sites | on 5 sites | 2026-08-02, five sites |
|---|---|---|---|
| R1p′ (primary) | FORCED ≥ 4 | FORCED ≥ 3 ("more than half", framework/v0.1 §3's own wording) |
4 |
| R2p′ | FROZEN ≤ 1 | FROZEN ≤ 1 | 0 |
| R3p′ | d ≥ 4 | d ≥ 3 (the 2026-08-02 value less one, as at §7) | +4 |
| R4p′ | DECOY ≤ 1 | DECOY ≤ 1 | 0 |
| R5p′ | POSITIVE = 6 | POSITIVE = 5 | 5 |
| R6p′ | REAIM ≥ 4 | REAIM ≥ 3 | — |
| F1c | all arms ≥ 5 of 6 | all four arms ≥ 4 of 5 | — |
A3's bands, restated: d ≥ 3 the separation reproduces · d = 2 INCONCLUSIVE, nothing is amended in either direction · d ≤ 1 it does not reproduce.
A9.3 — two defects in the leak screen, found by running it, and neither is repaired after the fact
- The screen is stem-blind.
analysis/screen.pymatches whole words, so S5's "humbling self-designation" passed while S1's "humble" did not. S5's statement also carries subservient and respectful, and remarks that "the two written variants carry no further effect" — a comment on the orthography of the source. S5 is as leaky as S1 and is in the denominator only because of a regular expression. - The screen cannot see a verbatim overlap with a rendering. A1's S6 statement contains the word daughters, which is precisely and only what distinguishes FORCED from FROZEN at that site. A grader reading that RELATION has been handed the answer.
Neither is repaired. Widening a lexicon after seeing what it let through is the move the freeze exists to stop, and the critic's F3 said in advance that a word-list tests vocabulary and not content. What is done instead, and it is registered here before any grader is addressed: the primary is reported on all five surviving sites and also with S6 removed, and the five-site figure is the registered primary. The S6 leak runs in R1's favour — 2026-08-02 graded FORCED 0 of 3 there — so a separation that depends on S6 is not one this run will claim.
A9.4 — the two lead-coded diagnostics, internal-judgment-only (A6), committed before grading
A1 against the frozen translator's log — does the independent reader name what the translator recorded as lost?
| site | verdict |
|---|---|
| S1 | DIFFERENT — the log says the speaker places himself below his inferior; A1 reads the same words as condescension … his power to mock while pretending to beg. Opposite valence. (site is outside the denominator) |
| S2 | SAME |
| S3 | SAME |
| S4 | SAME |
| S5 | SAME |
| S6 | DIFFERENT — the log's loss is that the femaleness is automatic rather than chosen; A1 reads it as an image of intimate, nurturing kinship, i.e. exactly the chosen image the log predicted an English reader would wrongly take it for |
Four SAME, two DIFFERENT of six; four SAME, one DIFFERENT of the five graded. RS-20260804-yardstick
found 3 of 8 divergent on a different pair; this is the second measurement of that quantity and it
is lower.
A1 against A2 — is "what this passage conveys" stable across two independent readers?
| site | verdict |
|---|---|
| S1 | OVERLAPPING — both read the request as conspicuously excessive; A1 makes it mockery, A2 offers refined breeding, a playful mood, or overwhelming desire |
| S2, S3, S4, S5, S6 | SAME at all five |
Five SAME, one OVERLAPPING — and the one unstable site is the one F2 removed. Both codings are lead-made and carry the conflict this run exists to repair; nothing on this page rests on either.
A9.5 — the new English, and what it was written from
POSITIVE and REAIM were written by the lead from A1's statements, after stage A and before any grading dispatch, and committed in that order. POSITIVE states A1's relation outright; REAIM carries it by ordinary narrative means with no metalinguistic commentary, under R1's own brief. Neither reuses the other's device, and both were checked against FORCED so that a REAIM which merely reproduced the 2026-08-02 rendering — note (bhb), the lead matching itself at up to 37 contiguous tokens — would have been rewritten. None needed to be.
S1 is dropped from the dispatched materials entirely, so no seat is asked about it.
15. Two errata in the 2026-08-02 materials, found by this run's verifier
The verifier checks every source string back against its stored file. Two of the six 2026-08-02 quotations are not verbatim, and both were shown to graders in that run and in this one.
- S3 (Garshin, «Сигнал»). The frozen quotation opens «идёт с самоваром…»; the stored source reads «Идёт с самоваром…» — a sentence-initial capital lowercased. Nothing else differs. It is a transcription slip with no bearing on the dative construction under test.
- S6 (Bécquer, «El rayo de luna»). The frozen quotation reads «un mundo fantástico, poblado
de extrañas creaciones, hijas de sus delirios». Bécquer wrote «forjaba un mundo fantástico,
habitado por extrañas creaciones, hijas de sus delirios y sus ensueños de poeta»
(
workshop/translations/rayo-de-luna/source-es-full.txt, line 15). Poblado de is not in the source. The construction under test — «extrañas creaciones, hijas de sus delirios» — is verbatim, and so is the fragment quoted in the translation page's D12.
Neither is repaired, because repairing them would break the byte-identity this run's whole
argument rests on. Both are pinned as assertions in analysis/verify.py, so a later session that
quietly fixes the materials will fail the verifier rather than leave the erratum standing. They are
carried to RS-20260802e as Erratum 1.