Repository path: wiki/findings/results/RS-20260814c-plain-yardstick.md · rendered 2026-09-09
Page metadata (front matter)
| type | result |
|---|---|
| id | RS-20260814c-plain-yardstick |
| status | open |
| created | 2026-08-14 |
| updated | 2026-08-14 |
| senses | affect |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260813f-affect-confound/design.md, wiki/arms/ARM-affect-confound.md, wiki/goodness-senses.md, wiki/decisions/resolved/D-20260804-16-affect-two-halves.md, wiki/decisions/resolved/D-20260813-17-affect-halves-discharge.md, wiki/findings/results/RS-20260813b-affect-yardstick.md, wiki/findings/results/RS-20260812f-affect-unprompted.md, workshop/translations/hoshi/R06-v1/translation.md, workshop/translations/hoshi/R08-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R08-resistancy.md, config/models.md, config/budget.md |
RS-20260814c — the plain description of an original cannot be written plainly, and the gate that checked it against the Japanese voided the run
ARM-affect-confound step 2, and the arm's last step. Design E-20260813f, frozen 2026-08-13
at 030c3b5, amended before dispatch on the pre-run critic's eight findings and again at 3b85eaf
on measurement (§12 A1–A6); the two arms frozen at 1822ac7 and 5f1fd0b before the design
existed. Everything here is internal-judgment-only and provisional: affect is untested,
Tier D is NOT PASSED, and no jury verdict in this project carries evidential weight.
261 bodies, 0 dead, $0.816705 including two discarded attempt sets. Verifier 1,594 checks, 0 failures; 9 mutation tests, 9 caught. Stage 5 — the 135 judging calls — was not dispatched, and §8 below says why.
1. The question, and the one-sentence answer
The design asked whether the H2 judgment — which of two renderings comes closer to doing to its
English reader what the original does to a reader of the original — moves when the prose style
of the description it is given moves. Two prior runs (RS-20260812f, Bulgarian;
RS-20260813b, Romanian) found that the plainer rendering wins H2 while the more marked
rendering wins H1, and both hang on the possibility that H2's seats were simply preferring
whichever English resembled the plainly-written description in front of them.
The run never reached that question, and the reason is the answer to a better one. Two of the gates the pre-run critic forced into the design fired at once:
The "plain" description is not plain — three seats call it ornate at 41 of 45 cells, unanimously on 11 of 15 segments — and the seat that reads Japanese calls it wrong about the source on 12 of 15 segments.
So the manipulation had no contrast to test: its baseline was already sitting at the marked pole it
was supposed to be moved away from. D-20260804-16 condition 3 does not discharge, and the
criterion that failed is criterion 2, clause 3 — the plain document must be measured plain —
with F2b (the source-fidelity gate) voiding the run before that.
What this is a question about, in the arm's own words: whether a reader who cannot read the source can be told which rendering comes closer to what the original does. That reader can only be told through a written description. This run is the first time this project measured the description itself, on either of the two properties the judgment needs it to have — being right about the original, and being stylistically neutral between the renderings — and it has neither.
2. What ran
15 segments of 国木田独歩 「星」 (1896), the piece whole, in two lead renderings frozen before the
design: T-hoshi-R06-v1 (single pass, source only, no rule set) and T-hoshi-R08-v1 (Venuti's ten
foreignizing rules). Seats per design §3: judges P1 openai/gpt-5.6-terra, P2
google/gemini-3.6-flash, P3 x-ai/grok-4.5; yardstick author and gate coder
qwen/qwen3.7-max; source-fidelity gate z-ai/glm-5.2, which wrote nothing and judges nothing.
| stage | what | calls |
|---|---|---|
YP |
the plain yardstick, written from the Japanese only | 15 |
YM |
the same content restyled ornate, from YP only, without the Japanese |
15 |
GJ |
does YP describe the Japanese? (critic finding 1) |
15 |
GA |
do the two arms state the same narrative facts? | 15 |
GY |
do the two documents say the same things? | 15 |
GD |
is each document PLAIN / ORNATE / SOURCEWARD? (criterion 2) |
96 |
GS |
which arm's prose is closer to the document's? | 90 |
| stages 1–4 | 261 | |
H1, H2P, H2M |
not dispatched — see §8 | 0 |
3. GD — the plain document is not plain, and neither arm reads as plain either
Criterion 2 has four clauses. Two fail.
| clause | requirement | observed | |
|---|---|---|---|
| 1 | the R08 sample is SOURCEWARD and the R06 sample is PLAIN |
R08 SOURCEWARD 3–0; R06 SOURCEWARD 3–0 |
FAILS |
| 2 | YM is ORNATE on ≥ 12 of 15 |
15 of 15, and 45 of 45 cells | holds |
| 3 | YP is PLAIN on ≥ 12 of 15 |
0 of 15 | FAILS |
| 4 | per-segment exclusion of any segment failing 2 or 3 | excludes all 15 | — |
Clause 3 is the finding. YP's instruction was explicit and not subtle: "In plain, ordinary
English — the register of a clear encyclopedia entry." What came back reads like this (segment
S06, the whole document is in runs/yp__S06.json):
Stars descend toward treetops while dew rises to rejoin the heavens, creating a reciprocal motion between celestial and terrestrial elements. … The resulting tone is serene yet charged with quiet intensity, immersing readers in a contemplative atmosphere where language itself enacts the harmony it depicts.
Per seat, on the 15 YP documents: P1 ORNATE 11, PLAIN 4; P2 ORNATE 15; P3 ORNATE
15. Only P1 ever dissents. On the 15 YM documents the verdict is ORNATE 45 of 45. The
manipulation's two poles are on the same side of a three-way choice, which is why GS moves on
only 5 of 15 segments (§4) and why nothing downstream could have measured what it was built to
measure.
Clause 1 is a second finding and a sharper one for translation. Asked which way a passage
departs from plain modern English, all three seats put both arms at SOURCEWARD, 9 cells of 9,
and their stated reasons name the same word:
P1onR06: "archaic and literal-sounding constructions such as "young man-star"…"P2onR06: "literal calques like "man-star"…"P3onR06: "literal compounds like "man-star"…"
R06 is the plain arm. It says man-star because 男星 has no English noun, and no rendering that
keeps the sky's two lovers distinct can avoid coining something. Where the source carries a word
English lacks, even the plainest rendering acquires the surface signature of a foreignising one,
and readers classify on that signature rather than on the prose around it. That is a claim about
reading translations, and it is the run's most transferable result — though it stands on one
pre-registered segment (S07) and 9 cells, and nothing more may be built on it than that.
4. GS — and the seats do not agree about the plain document
| document | matched to R06 |
matched to R08 |
unanimous |
|---|---|---|---|
YP (plain) |
6 | 9 | 3 of 15 |
YM (ornate) |
1 | 14 | 11 of 15 |
The style match moves between the two documents on 5 of 15 segments (S03, S07, S08, S10, S12),
against P3's clause (b) bar of 0.60. And the per-seat tables show why the YP row is close to
meaningless:
| seat | YP → R06 : R08 |
YM → R06 : R08 |
|---|---|---|
P1 |
2 : 13 | 1 : 14 |
P2 |
5 : 10 | 1 : 14 |
P3 |
13 : 2 | 3 : 12 |
P3 reads the plain document the opposite way from the other two, and the majority rule then
turns a 2–1 split into a run-level figure. On the ornate document all three converge. This is the
same shape RS-20260814b §5 recorded on its coding seats a few hours earlier — a rate that is an
artefact of which estimator you pick — and it recurs here on a different task and different seats.
5. P3, the manipulation check — FAILS, and F3 fires
Computed over segments with a GS match under both documents (§4 majority rule):
| clause | requirement | observed | |
|---|---|---|---|
| (a) | ≥ 8 segments matched under both | 15 | holds |
| (b) | the match differs between YP and YM on ≥ 0.60 |
5 / 15 = 0.333 | FAILS |
| (c) | YP matched to R06 on ≥ 0.80 |
6 / 15 = 0.400 | FAILS |
F3 fires and P2 is void. §5 of the design pre-committed, in writing and before the run, that
no override of F3 would be taken whatever the withheld comparison showed, and none is taken.
P2's denominator is in any case empty: every one of the 15 segments is excluded by F4 —
15 for YP_not_PLAIN, 12 for GJ WRONG, 10 because the style match did not move, 9 for GY
CHANGED. With n = 0 against a registered power floor of 11, P2 is WITHHELD as well as
void. There is no co-movement figure to report, licensed or unlicensed.
6. F2b — the gate the critic forced, and what the lead's own reading of it found
GJ returned WRONG on 12 of 15 segments against a bar of ≥ 5. F2b fires: the whole run is
void, P1 included, because H2P is built on YP.
The critic's finding 1 — nothing validated YP against the Japanese — was the sharpest thing that
pass produced, and this is what it bought. But a gate that fires at 80% deserves to be read rather
than counted, and the lead reads Japanese, so the twelve reasons were adjudicated by hand against
the source. They do not all stand.
| segment | GJ's complaint |
the lead's reading |
|---|---|---|
| S06 | YP calls the smoke "incense" |
GJ is right. 「ただ詩人が庭の煙のみ」 is the leaf-fire lit in S05 (「落ち葉つみたる一つへ火を移さしめて」). Turning a heap of burning dead leaves into incense is a real misreading, and it changes the scene |
| S02 | YP overstates the melancholy, ignoring 「楽しみて」 and 「誇る」 |
GJ is right. The paragraph ends the poet lived enjoying this garden morning and evening; the document makes it uniformly elegiac |
| S08 | YP renders 男星 as 天津乙女 |
GJ is wrong. 「天津乙女は…男星の肩に依れり」 — YP makes the maiden the subject and the male star the shoulder, exactly as the Japanese does |
| S10 | YP misidentifies the sleeping poet as the host |
GJ is wrong, and confused: the poet is the host and is asleep (「庭の主人に一語の礼なくて…詩人が寝顔を二人はしばし見とれぬ」) |
| S11 | YP makes the weeper one young woman, not 乙女の星 and 恋人たち |
GJ is wrong. 「乙女の星はこれを見て…涙うかべ」 — she weeps alone; the lovers act together only in the next clause |
| S12 | YP omits that the star is on her forehead |
a nitpick, not a content error |
| S15 | 「肩に垂るる黒髪」 is hair falling to the shoulders, not shoulder-length | a nitpick |
| S01 | the garden is "unbefitting the house in being broad", not "unusually spacious" | a nitpick |
| S05, S09, S13, S14 | mixed — one real omission (「青煙一抹」), one disputed gloss of 潯, two disagreements about tone rather than content | not adjudicated either way |
So two statements are both true and both go on the record. F2b fired on the registered
criterion and the run is void by it — no override, and the criterion was registered before anything
was dispatched. And the gate that voided it is not itself a reliable measure of how wrong the
yardstick is: of the five verdicts read closely, two are real errors and three are the gate
misreading either the Japanese or the description. What F2b's firing licenses is the yardstick
is not dependable, not the yardstick is wrong four times in five.
One error GJ did not catch, found by the same hand-read: on S10 「淡紅色の霞につつまれて乙女
の星先に立ち」 has the maiden-star wrapped in the pale-pink haze and going first; YP has the
smoke guiding her. A gate that misses that while firing on a forehead is a gate with a taste, not
a threshold.
7. GA, GY, and what the restyle did
GA — the arms are comparable. F1 does not fire. DIFFERENT on 2 of 15 against a bar of 3.
Both flags (S08, S09) are the same thing: 水色 rendered pale blue in R06 and water-colour in
R08. That is a difference of wording inside a colour term, which the prompt's first clause
excludes; the registered carve-out covers only a quoted-line flag on S11 or S15, so neither is
carved out and both are counted — and F1 still does not fire. On a stricter reading it fires
even less.
GY — the restyle moved content on 9 of 15. F2 fires. What it added is worth quoting, because
it is a fact about what "restyle this, change nothing" does:
- S01: the ornate version adds that the text is "an object of profound cultural and aesthetic significance demanding interpretive reverence";
- S02: it adds that the architecture renders the garden "an emblem of perpetual mutability perceived through refined aesthetic consciousness";
- S05: it adds "a meta-claim that every factual and atmospheric assertion remains inviolate beneath the elaborate stylistic superstructure" — the rewrite inserted a sentence asserting its own fidelity.
Asking a model to raise a description's register makes it add evaluative claims the description did not contain, and the claims it adds are appreciative. Anyone building a marked-prose control out of a restyle needs a content gate, and 9 of 15 is how often it bites.
8. Why stage 5 was not dispatched
F2b's registered consequence is the whole run is void, P1 included, because H2P is built
on YP. H1, H2P and H2M are the 135 calls that compute P1, and P1 is the only primary
they feed. Dispatching them would have produced numbers from a run already void by a criterion
frozen a day earlier, and the only use for those numbers would have been a comparison the design
does not license. They were not dispatched. That leaves ≈$0.40 unspent and leaves this run with
nothing to say about whether the H1/H2P separation replicates on a third language family —
stated plainly, because it was the second thing the run was for.
H1 alone is not a primary of this design, and running it on its own after seeing the gates
fail would have been an unregistered analysis chosen with the results in view. It was not run
either.
9. Verification, bodies and money
analysis/verify.py was written after the design was frozen, imports nothing from tools/ and
nothing from analyse.py, and reads analysis/results.json as a claim. 1,594 checks, 0
failures. It recomputes every cell choice against its stored body, every §4 majority from the
usable bodies, the P2 denominator from the raw record, the exact binomial critical values by
exhaustive enumeration, and the cost total over all attempts including discarded ones.
Nine mutation tests, nine caught. Three of the design's eight named mutations address H2P
/H2M cells or the P2 denominator, and this run has neither, so those three are recorded as not
exercised and replaced by the same mutation on a stage that exists; the substitutions are labelled
as such in the verifier's output. One of the original eight — a document swapped between YP and
YM — was MISSED on the first pass, because nothing tied a prompt's embedded document back to
the body that wrote it. That check was added (195 assertions of document provenance across GD,
GS and GJ) and the mutation is now caught.
Two defects in the frozen design, both found by the verifier and both recorded rather than patched away (design §12 A1, A2, and one more here):
- §10.3's assertion that no
GSprompt contains the words original, Japanese, source, effect, comparable or translation is unsatisfiable by construction. AGSprompt embeds aYPdocument, andYP's own §4 stage 1 instruction requires the word Japanese — "what the passage does to a reader of Japanese". Measured: 24 of 90GSprompts carry a banned word inside the document block, across 5 segments, the words being effect, japanese, original. The verifier asserts the ban on the prompt scaffolding, where it holds absolutely, and counts the document-block contamination rather than waving it through.GSis therefore not blind in the way the design claims: in a quarter of its cells the seat could see that the two passages are translations from Japanese. A third independent reasonP2could not have been trusted here. - §10.4's byte-identity claim across
H1,H2P,H2M, which cannot hold becauseH1asks a different question — scoped in A2 and moot, since stage 5 did not run.
Bodies. 261 dispatched, 0 dead, 266 attempts. F5 fires on no stage. Providers: OpenAI 62,
xAI 62, Alibaba 60, Google 45, Google AI Studio 17, CoreWeave 14, Phala 1.
Money. $0.816705 for the run, plus $0.023320 of transport probes, $0.840025 in all, against a
revised ceiling of $2.26 for stages 1–4. Key usage 107.830757299 → 108.692808434, delta
$0.862051, leaving $0.022027 unreconciled in the over-counting direction — the same lag this
session found at start-up on S181's closing snapshot ($0.265127 settling after that session closed).
Per-request costs are primary (CLAUDE.md); the residual is recorded and not chased.
$0.044710 of the run — 5.5% — bought nothing, two discarded attempt sets, both preserved under
runs/_discarded/ with their costs counted:
| discarded | why | cost |
|---|---|---|
yp__S01, 3 attempts at caps 1,600 / 3,200 / 6,400 |
note (bnk) on a third seat: reasoning: {"effort": "low"} sent and ignored, 4,662 reasoning tokens for 171 visible ones, then a body cut off mid-sentence with finish_reason: "stop" |
$0.043586 |
yp__S14, 1 attempt |
1,200 tokens of bare prose, no JSON anywhere, finish_reason: "stop" — a parse failure F6 covers and the runner's first usability test called usable |
$0.001124 |
10. Limits, at the size of the claims
- Model seats, not readers. Tier D is NOT PASSED. Every figure is a fact about how five seats
behave, and
GDandGSare themselves unanchored model judgments about style — design §9 limit 5, which this run's §4 table sharpens rather than resolves. - The
YP-is-ornate finding is about one yardstick author on one text.qwen/qwen3.7-max, reasoning disabled, on Meiji 擬古文, 15 documents. It says nothing about whether a human, or a different seat, could have written a plain one — and nothing about the yardsticks used inRS-20260812fandRS-20260813b, which were written by a different seat on different languages and have never been measured on either property. - The
F2bcount is the gate's, not the lead's. §6 gives the adjudication. Five of twelve were read closely; seven were not read at all, or not resolved. - The
SOURCEWARD-on-both finding rests on one segment and 9 cells, pre-registered as S07 before the run, but one segment nonetheless. - This run says nothing about the confound it was built to test. Not that the confound is absent, not that it is present. The instrument that would have measured it was never in a state to do so, and failing to build a control is not evidence about what the control would have shown.
- Nor does it say anything about whether the
H1/H2Pseparation replicates on Japonic. Stage 5 did not run.
11. What follows
D-20260804-16condition 3 does not discharge. The criterion that failed is criterion 2, clause 3, withF2bandF3firing before it.wiki/goodness-senses.md§affectis rewritten to say so, and usage rule 4 — never report a singleaffectfigure covering both halves — stands untouched, since it is not what condition 3 governs.ARM-affect-confoundclosesresolvedat 2 of 2, on its own terms: "an inconclusive third state is possible and is a failure to discharge, not a licence to try again inside this arm."- This is the second consecutive failure of this manipulation family, and the design
pre-registered what that means: "two failures make the manipulation itself the finding."
E-20260813bvoided on its own manipulation check;E-20260813fvoids on a source-fidelity gate and a directionality gate. A future design that wants to test the style confound must first solve the problem this run surfaced: producing a description of a foreign passage that three seats will call plain. Nothing in the two runs suggests asking for it works. - The two prior
affectresults keep their direction and acquire a caution.RS-20260812f§5 andRS-20260813b§5 rest on yardstick documents that were never checked for accuracy against the source and never checked for style, and this run is the only evidence anywhere in the project about what such a check finds. The caution is recorded onwiki/goodness-senses.md§affect; it is not a withdrawal, because nothing here measured those documents.