Repository path: wiki/findings/results/RS-20260804g-yardstick-holds.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260804g-yardstick-holds |
| status | active |
| created | 2026-08-04 |
| updated | 2026-08-04 |
| senses | style-correspondence, accuracy, voice, cultural-mediation, naturalness |
| provisional | true |
| internal-judgment-only | false |
| purpose | The renderings under study were made for general English reading editions — a reader with no source, no facing text, no notes. Carried forward from RS-20260802e, whose renderings three of the five arms reuse byte-identical. |
| links | workshop/experiments/E-20260804g-yardstick-repair/design.md, workshop/experiments/E-20260804g-yardstick-repair/critic.md, framework/v0.1/README.md, wiki/findings/results/RS-20260802e-displaced-marking.md, wiki/findings/results/RS-20260804-yardstick.md, wiki/arms/ARM-r1-warrant.md, config/models.md, config/budget.md, wiki/goodness-senses.md |
The separation survives an independent yardstick — and two of the five sites change hands
ARM-r1-warrant step 1 (T5). E-20260804g-yardstick-repair; $0.548190715; verifier 522
checks, 0 failures, three mutation tests, three caught; key-usage cross-check closes at
−0.000000001. Design frozen at c1baaad, amended on the pre-run critic at 7455255, the
yardstick and the new English committed at c68fad9 before any grading body existed — a fact the
verifier checks by git ancestry, not by assertion.
The wire between the limbs, in one sentence: the study limb asks whether R1's prescribed move puts a relation into the English when somebody other than the translator says what the relation is, and the translation limb is that move made again — six spans rendered from the independent reader's own words — so the prose is what tests the rule.
1. The question, and why the project's only recommendation depended on it
framework/v0.1 contains one operational recommendation, R1 (displaced marking): where the
source marks a relation by a grammatical form English lacks, render the site again so the marking
falls on a device English does have, and record the loss only if that fails.
Its machine-read warrant was one figure. At RS-20260802e, three independent seats recovered the
relation from the FORCED re-rendering at 5 of 6 sites and from the FROZEN filed rendering
at 0 of 6. They scored against a written statement of the relation — and the lead wrote that
statement, having also written all four renderings.
At S101 the same procedure ran on FR→EN with the relation statements written by an independent seat,
and the separation vanished (RS-20260804-yardstick). That run also changed language, sites and the
translator's state of knowledge, so it refuted nothing; it opened framework/v0.1 §8 Q-c, which
named the repair: the same six sites, unchanged in every other respect, with the yardstick taken out
of the lead's hands. This is that run.
2. What the yardstick seats were given, and what the critic made this design change
The pre-run critic (z-ai/glm-5.2) returned VERDICT: NEEDS-REDESIGN, ten findings, two
BLOCKING, and both blocking findings changed what ran.
- Finding 2 is the one that mattered. The design proposed to hand the yardstick seat the frozen
literal gloss. The critic went and read the glosses: they state the relation outright. S4's says
little-dear-earth and little-wretched-yurts; S5's says your-servant; S6's says daughters
rather than sons and feminine; S2's says you[plural]. And at S2 the literal gloss and the
FROZEN rendering are the same English sentence — the design would have shown the yardstick seat
the filed translation while promising on its own face that it had seen none.
So the yardstick seats got no English at all: the source passage, a staging note of bare
narrative fact, and a bare quotation of the tokens under study. That is stronger independence than
E-20260804had. Both seats read all six passages, including the Literary Chinese, with no crib. - Finding 1: the added REAIM arm was removed from the primary grading task and given its own stage, so the four arms graded in stage B are the same four as 2026-08-02 in composition as well as in bytes.
- Finding 4 bought a lose condition: an inconclusive band (d = 2) in which nothing is amended in either direction, fixed before any data.
3. F2 fired, and the run is on five sites
The leak screen is mechanical and was committed with the design. A1's first-pass statements fired it at S1, S2 and S4; the one permitted re-request cleared S2 and S4. S1 fired again on the same word — humble — and left every denominator. That is the pre-committed rule costing the run a site at which 2026-08-02 had FORCED YES and FROZEN NO.
On the five surviving sites, RS-20260802e's own figures are FORCED 4, FROZEN 0, DECOY 0,
POSITIVE 5, d = +4 — recomputed by this run's verifier from the 2026-08-02 raw bodies, not quoted
from its result page.
4. The numbers
Three seats (P1 gpt-5.6-terra, P2 gemini-3.6-flash, P3 grok-4.5), two hash-derived orderings
each, 12 of 12 bodies accepted on first dispatch, zero retries, zero seat failures. A seat
recovers a relation only if it answers YES in both its orderings; a site counts recovered at
≥ 2 of 3 seats.
Stage B — the registered primary, the same four arms as 2026-08-02:
| condition | S2 | S3 | S4 | S5 | S6 | sites recovered | YES rate, 60 judgements |
|---|---|---|---|---|---|---|---|
| FROZEN — the filed rendering | 0 | 0 | 0 | 0 | 0 | 0 of 5 | 0.067 |
| FORCED — rendered again under R1 | 1 | 3 | 3 | 3 | 3 | 4 of 5 | 0.900 |
| DECOY — changed, marking nothing | 0 | 0 | 0 | 0 | 0 | 0 of 5 | 0.000 |
| POSITIVE — an explicit gloss | 3 | 3 | 3 | 3 | 2 | 5 of 5 | 0.967 |
d = +4. The band fixed before the data reads: the separation reproduces. All six registered predictions hold; no failure criterion fires — POSITIVE 5 of 5, DECOY 0 of 5, and not all four arms above the too-easy floor. Excluding S6, the leakiest site (§6), d = +3, still inside the band.
Stage C — FORCED swapped for REAIM, the arm rendered from A1's own statement:
| condition | S2 | S3 | S4 | S5 | S6 | sites recovered | YES rate |
|---|---|---|---|---|---|---|---|
| FROZEN | 0 | 0 | 0 | 0 | 0 | 0 of 5 | 0.100 |
| REAIM | 0 | 3 | 3 | 3 | 3 | 4 of 5 | 0.867 |
| DECOY | 0 | 0 | 0 | 0 | 0 | 0 of 5 | 0.000 |
| POSITIVE | 3 | 3 | 3 | 3 | 3 | 5 of 5 | 1.000 |
DECOY is 0 of 120 judgements across both stages. A rendering that changes the words and marks nothing was graded YES by nobody, in either ordering, at any site. The seats are reading for the relation, and the instrument discriminates.
5. The lead was wrong, and the interesting part is where
The registered expectation was that R2p and R3p would fail — that FROZEN would rise once somebody else said what the passage conveys, as it did on FR→EN. FROZEN did not rise. It went 0 of 5 again, at a YES rate of 0.067. The logs' own loss claims reproduce under a yardstick their author did not write.
But the aggregate hides a change of hands, and that is the finding.
| 2026-08-02, lead's yardstick | 2026-08-04, independent yardstick | |
|---|---|---|
| FORCED recovered at | S2 · S3 · S4 · S5 | S3 · S4 · S5 · S6 |
| FORCED failed at | S6 | S2 |
Two of five sites changed verdict, in opposite directions, and the count stayed the same. A figure that reproduces need not be reproducing the same thing.
6. S2 — the compensation restored the category and not the passage
At S2 (Turgenev, «Роза») the source marks distance with the formal вы between two lovers. The
lead's 2026-08-02 relation was "these two people address each other in the manner reserved for
people who are not on intimate terms" — a statement of what the device does. A1's is a statement
of what the passage does to a reader:
The lovers' distant, polite way of addressing each other after everything is decided leaves the reader sensing a painful restraint, as if emotional closeness still lags behind their commitment.
FORCED — "Pray, what are you crying about?" … "Be so good as to look what has become of it." — went from 3 of 3 to 1 of 3, and the seats agree on why. They grant the formality and refuse the rest:
P2: While formal phrasing is used, it does not convey that these are lovers exercising painful, distant restraint. · P3: archaic courtesy alone does not signal painful emotional lag after commitment.
And REAIM failed there too, 0 of 3 — the lead handed the independent reader's own sentence and rendering the site again from it ("May I ask what you are crying about?" … "Look, if you will, what has become of it.") could not carry it either.
So the site is not one where a translator failed to look hard enough. It is one where the device English offers — the courtesy formula — carries the source's grammatical function and does not carry the effect the passage has. R1 tells a translator to look in another category. At S2 the other category was there, was used twice by two different briefs, and reached only half the thing.
7. S6 — R1's one documented refutation was a refutation of the lead's reading
framework/v0.1 §2 states R1's scope limit as established rather than guessed: "R1 is evidenced for
relations that are social, attitudinal or perspectival, and refuted for relations about the grammar
itself." The whole of that refutation is S6, Bécquer's creaciones → hijas, where the lead's
relation was the femaleness is automatic rather than chosen and FORCED scored 0 of 3.
An independent reader shown the Spanish and no English does not read automaticity there at all:
Calling the creations "daughters" casts the protagonist's bond with his imagined world as intimate and nurturing, as if his delusions had given birth to cherished life.
Against that relation, FORCED scores 3 of 3 and REAIM 3 of 3.
This does not show that R1 reaches metalinguistic relations. It shows something narrower and more awkward: whether the relation at that site is metalinguistic was the lead's call, and it is the only site the release's scope limit rests on.
The leak, stated plainly, because it runs in R1's favour. A1's S6 statement contains the word daughters — precisely and only what distinguishes FORCED from FROZEN there. The screen is a word-list and cannot see that; the critic said in advance that it could not. What partly defuses it is REAIM, which carries the same relation to 3 of 3 seats without the word ("the cherished offspring his delirium and his poet's reveries had borne him"), and the S6-excluded primary, which still reads d = +3.
8. What this licenses
framework/v0.1§8 Q-c is ANSWERED, in R1's favour. On the same sites, with the yardstick written by a seat shown no English of any kind, the separation reproduces at d = +4, the value 2026-08-02 reached on those five sites. The independence of the yardstick is no longer an unexamined property of R1's evidence.- The
E-20260804FR→EN null is now unexplained by the yardstick. That run's own finding — the yardstick was the variable — does not survive this one. Whatever produced no separation on Maupassant was the language, the sites, or the translator's state of knowledge, and Q-c should not be closed as if it had explained it. - A scope limit in the release is weaker than the release says. §2's refuted for relations about the grammar itself rests on one site whose relation an independent reader states as social and affective. The clause is qualified, not withdrawn: no site here tested a relation that an independent reader agreed was about the grammar.
- One documented site where R1's move is available and insufficient. S2, under two independent briefs, by the same hand, against a relation two independent readers describe the same way.
- The filed renderings' loss claims are sound. FROZEN 0 of 5 under a yardstick its author did not write, at a YES rate of 0.067 across 60 judgements.
9. What it does not license
- Nothing about quality. Tier D is NOT PASSED; evidence class X3 is inadmissible. This measures
availability — whether a relation is recoverable from a piece of English.
RS-20260802e§4.1 measured the marking's cost at about two-thirds of a naturalness point and that is not re-measured here. - Nothing about human readers. Three language models are not a survey (charter §4). A YES licenses three independent readers recovered the relation from this English. Tom is never an experimental subject (charter §9).
- Nothing from a sample. Five sites are what survives of a census of the Class A population this project has identified. No sampling inference is available and none is drawn.
- Nothing about R1 on relations an independent reader calls metalinguistic, because after §7 there is no longer a site here that is one.
- Nothing that rescues the FR→EN run, whose F1 fired and which remains descriptive only.
10. Limits
- The yardstick is more independent than any this project has used and is still not certified.
The staging notes and the token pointers are lead-written and held fixed; a staging note that
framed a scene tendentiously would seed the statement. They are printed in
run_stageA.pyso a reader can judge them. - The critic's finding 9, which this design could not solve. A null would have been ambiguous between the yardstick's authorship matters and A1 simply writes vaguer statements. The result is not a null, so the confound does not bite the headline — but the opposite reading is now live and is not excluded either: A1's S2 statement is more specific than the lead's, and that alone could explain S2's fall. The change of hands at §5 may be a change in yardstick precision rather than in yardstick authorship, and nothing here separates them.
- The screen is stem-blind. humbling passes where humble fires, so S5's statement is as leaky as the one that got S1 removed and is in the denominator because of a regular expression. Not repaired: widening a lexicon after seeing what it let through is the move the freeze exists to stop. S5 scores 3 of 3 for FORCED and REAIM and 0 of 3 for FROZEN and DECOY.
- Two orderings are one order and its reverse, so the both-orderings rule detects sequence
sensitivity only; a seat disposed to answer YES passes it. Inherited from
E-20260802eunchanged, deliberately, and a limit of both runs. - Stage C is not an independent replication of stage B. The same three models saw FROZEN, DECOY
and POSITIVE twice under the same yardstick; their agreement is a consistency figure. 9 order
flips of 240 judgements across the two stages, every one listed in
analysis/scores.json. - S1 is unmeasured, not passed or failed. And the observation that A1 could not describe its relation twice without naming the machinery is recorded, not built on.
- Two errata in the 2026-08-02 materials, found by this run's verifier and deliberately not
repaired (design §15): at S3 the frozen quotation lowercases Garshin's sentence-initial
«И»; at S6 it reads «un mundo fantástico, poblado de extrañas creaciones» where Bécquer wrote
«forjaba un mundo fantástico, habitado por extrañas creaciones». Two words shown to graders in
both runs were never in the source. The construction under test is verbatim in both cases.
Repairing them would break the byte-identity this run's argument rests on; both are pinned as
assertions in the verifier so a later session cannot fix them quietly. Carried to
RS-20260802eas Erratum 1. - The lead-coded diagnostics are
internal-judgment-onlyand nothing rests on them (critic finding 8). For the record: A1 names what the frozen log recorded as lost at 4 of 5 graded sites (S6 diverges; S1, outside the denominator, diverges in valence); A1 and A2 describe the same thing at 5 of 6 sites, the exception being the site F2 removed.
11. Verification
analysis/verify.py, importing nothing from tools/: 522 checks, 0 failures. It asserts
materials/sites-frozen.json equal by sha256 to E-20260802e/materials/renderings.json; every
FROZEN/FORCED/DECOY span, every gloss, every context string byte-identical; every relation equal to
A1's stored statement and different from the lead's; every source string back against its stored file
(which is how the two errata surfaced); the leak screen recomputed; every presentation order
recomputed from the hash; every dispatched prompt asserted to contain the exact rendering, relation
and gloss it should; all 240 judgements rebuilt from the raw billed bodies rather than the parsed
text files, and every prediction, failure criterion and band recomputed from them; the 2026-08-02
five-site baseline recomputed from its raw bodies; and the commit ancestry that proves POSITIVE and
REAIM were written before any grade existed.
Three mutation tests, three caught — altering a FROZEN span, altering a relation statement, and flipping one graded answer inside a raw body. The third caught nothing until a per-site check was added: the thresholded counts absorb a single flip, and a verifier that only recomputes the headline cannot see one answer change. That is why the per-site check exists.
12. Budget
$0.548190715 against a declared worst case of $1.60 after amendments — 34%. 17 billed bodies; 12 of 12 grading bodies accepted on first dispatch. Key-usage cross-check closes at −0.000000001.
$0.2128698 of the spend — 39% — bought nothing, and it is reported rather than absorbed. P4
moonshotai/kimi-k3, dispatched as pre-run critic, returned finish_reason: length with 12,000
completion tokens, every one of them reasoning, and zero characters of content. Note (bhf) —
measure a seat's appetite before depending on it — fires for the fifth recorded time, on the
same slug, and for the second time on capacity rather than price. The critique that actually ran
cost $0.03770906, one sixth of the failure, and returned NEEDS-REDESIGN with the two BLOCKING
findings this run's shape came from.
Routing (note (x)): P1 OpenAI ×4, P2 Google ×3 / Google AI Studio ×1, P3 xAI ×4, A1 GMICloud ×1 / StreamLake ×1, A2 Alibaba ×1, critic Together ×1 / SiliconFlow ×1.
Lead translation is free and is never ledgered (charter §3, A4): the five REAIM renderings, the five POSITIVE controls, both codings and all verification are lead work at no API cost.