Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: wiki/findings/results/RS-20260814c-plain-yardstick.md · rendered 2026-09-09

Page metadata (front matter)
typeresult
idRS-20260814c-plain-yardstick
statusopen
created2026-08-14
updated2026-08-14
sensesaffect
internal-judgment-onlytrue
provisionaltrue
linksworkshop/experiments/E-20260813f-affect-confound/design.md, wiki/arms/ARM-affect-confound.md, wiki/goodness-senses.md, wiki/decisions/resolved/D-20260804-16-affect-two-halves.md, wiki/decisions/resolved/D-20260813-17-affect-halves-discharge.md, wiki/findings/results/RS-20260813b-affect-yardstick.md, wiki/findings/results/RS-20260812f-affect-unprompted.md, workshop/translations/hoshi/R06-v1/translation.md, workshop/translations/hoshi/R08-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R08-resistancy.md, config/models.md, config/budget.md

RS-20260814c — the plain description of an original cannot be written plainly, and the gate that checked it against the Japanese voided the run

ARM-affect-confound step 2, and the arm's last step. Design E-20260813f, frozen 2026-08-13 at 030c3b5, amended before dispatch on the pre-run critic's eight findings and again at 3b85eaf on measurement (§12 A1–A6); the two arms frozen at 1822ac7 and 5f1fd0b before the design existed. Everything here is internal-judgment-only and provisional: affect is untested, Tier D is NOT PASSED, and no jury verdict in this project carries evidential weight.

261 bodies, 0 dead, $0.816705 including two discarded attempt sets. Verifier 1,594 checks, 0 failures; 9 mutation tests, 9 caught. Stage 5 — the 135 judging calls — was not dispatched, and §8 below says why.

1. The question, and the one-sentence answer

The design asked whether the H2 judgment — which of two renderings comes closer to doing to its English reader what the original does to a reader of the original — moves when the prose style of the description it is given moves. Two prior runs (RS-20260812f, Bulgarian; RS-20260813b, Romanian) found that the plainer rendering wins H2 while the more marked rendering wins H1, and both hang on the possibility that H2's seats were simply preferring whichever English resembled the plainly-written description in front of them.

The run never reached that question, and the reason is the answer to a better one. Two of the gates the pre-run critic forced into the design fired at once:

The "plain" description is not plain — three seats call it ornate at 41 of 45 cells, unanimously on 11 of 15 segments — and the seat that reads Japanese calls it wrong about the source on 12 of 15 segments.

So the manipulation had no contrast to test: its baseline was already sitting at the marked pole it was supposed to be moved away from. D-20260804-16 condition 3 does not discharge, and the criterion that failed is criterion 2, clause 3 — the plain document must be measured plain — with F2b (the source-fidelity gate) voiding the run before that.

What this is a question about, in the arm's own words: whether a reader who cannot read the source can be told which rendering comes closer to what the original does. That reader can only be told through a written description. This run is the first time this project measured the description itself, on either of the two properties the judgment needs it to have — being right about the original, and being stylistically neutral between the renderings — and it has neither.

2. What ran

15 segments of 国木田独歩 「星」 (1896), the piece whole, in two lead renderings frozen before the design: T-hoshi-R06-v1 (single pass, source only, no rule set) and T-hoshi-R08-v1 (Venuti's ten foreignizing rules). Seats per design §3: judges P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5; yardstick author and gate coder qwen/qwen3.7-max; source-fidelity gate z-ai/glm-5.2, which wrote nothing and judges nothing.

stage what calls
YP the plain yardstick, written from the Japanese only 15
YM the same content restyled ornate, from YP only, without the Japanese 15
GJ does YP describe the Japanese? (critic finding 1) 15
GA do the two arms state the same narrative facts? 15
GY do the two documents say the same things? 15
GD is each document PLAIN / ORNATE / SOURCEWARD? (criterion 2) 96
GS which arm's prose is closer to the document's? 90
stages 1–4 261
H1, H2P, H2M not dispatched — see §8 0

3. GD — the plain document is not plain, and neither arm reads as plain either

Criterion 2 has four clauses. Two fail.

clause requirement observed
1 the R08 sample is SOURCEWARD and the R06 sample is PLAIN R08 SOURCEWARD 3–0; R06 SOURCEWARD 3–0 FAILS
2 YM is ORNATE on ≥ 12 of 15 15 of 15, and 45 of 45 cells holds
3 YP is PLAIN on ≥ 12 of 15 0 of 15 FAILS
4 per-segment exclusion of any segment failing 2 or 3 excludes all 15 —

Clause 3 is the finding. YP's instruction was explicit and not subtle: "In plain, ordinary English — the register of a clear encyclopedia entry." What came back reads like this (segment S06, the whole document is in runs/yp__S06.json):

Stars descend toward treetops while dew rises to rejoin the heavens, creating a reciprocal motion between celestial and terrestrial elements. … The resulting tone is serene yet charged with quiet intensity, immersing readers in a contemplative atmosphere where language itself enacts the harmony it depicts.

Per seat, on the 15 YP documents: P1 ORNATE 11, PLAIN 4; P2 ORNATE 15; P3 ORNATE 15. Only P1 ever dissents. On the 15 YM documents the verdict is ORNATE 45 of 45. The manipulation's two poles are on the same side of a three-way choice, which is why GS moves on only 5 of 15 segments (§4) and why nothing downstream could have measured what it was built to measure.

Clause 1 is a second finding and a sharper one for translation. Asked which way a passage departs from plain modern English, all three seats put both arms at SOURCEWARD, 9 cells of 9, and their stated reasons name the same word:

P1 on R06: "archaic and literal-sounding constructions such as "young man-star"…" P2 on R06: "literal calques like "man-star"…" P3 on R06: "literal compounds like "man-star"…"

R06 is the plain arm. It says man-star because 男星 has no English noun, and no rendering that keeps the sky's two lovers distinct can avoid coining something. Where the source carries a word English lacks, even the plainest rendering acquires the surface signature of a foreignising one, and readers classify on that signature rather than on the prose around it. That is a claim about reading translations, and it is the run's most transferable result — though it stands on one pre-registered segment (S07) and 9 cells, and nothing more may be built on it than that.

4. GS — and the seats do not agree about the plain document

document matched to R06 matched to R08 unanimous
YP (plain) 6 9 3 of 15
YM (ornate) 1 14 11 of 15

The style match moves between the two documents on 5 of 15 segments (S03, S07, S08, S10, S12), against P3's clause (b) bar of 0.60. And the per-seat tables show why the YP row is close to meaningless:

seat YP → R06 : R08 YM → R06 : R08
P1 2 : 13 1 : 14
P2 5 : 10 1 : 14
P3 13 : 2 3 : 12

P3 reads the plain document the opposite way from the other two, and the majority rule then turns a 2–1 split into a run-level figure. On the ornate document all three converge. This is the same shape RS-20260814b §5 recorded on its coding seats a few hours earlier — a rate that is an artefact of which estimator you pick — and it recurs here on a different task and different seats.

5. P3, the manipulation check — FAILS, and F3 fires

Computed over segments with a GS match under both documents (§4 majority rule):

clause requirement observed
(a) ≥ 8 segments matched under both 15 holds
(b) the match differs between YP and YM on ≥ 0.60 5 / 15 = 0.333 FAILS
(c) YP matched to R06 on ≥ 0.80 6 / 15 = 0.400 FAILS

F3 fires and P2 is void. §5 of the design pre-committed, in writing and before the run, that no override of F3 would be taken whatever the withheld comparison showed, and none is taken. P2's denominator is in any case empty: every one of the 15 segments is excluded by F4 — 15 for YP_not_PLAIN, 12 for GJ WRONG, 10 because the style match did not move, 9 for GY CHANGED. With n = 0 against a registered power floor of 11, P2 is WITHHELD as well as void. There is no co-movement figure to report, licensed or unlicensed.

6. F2b — the gate the critic forced, and what the lead's own reading of it found

GJ returned WRONG on 12 of 15 segments against a bar of ≥ 5. F2b fires: the whole run is void, P1 included, because H2P is built on YP.

The critic's finding 1 — nothing validated YP against the Japanese — was the sharpest thing that pass produced, and this is what it bought. But a gate that fires at 80% deserves to be read rather than counted, and the lead reads Japanese, so the twelve reasons were adjudicated by hand against the source. They do not all stand.

segment GJ's complaint the lead's reading
S06 YP calls the smoke "incense" GJ is right. 「ただ詩人が庭の煙のみ」 is the leaf-fire lit in S05 (「落ち葉つみたる一つへ火を移さしめて」). Turning a heap of burning dead leaves into incense is a real misreading, and it changes the scene
S02 YP overstates the melancholy, ignoring 「楽しみて」 and 「誇る」 GJ is right. The paragraph ends the poet lived enjoying this garden morning and evening; the document makes it uniformly elegiac
S08 YP renders 男星 as 天津乙女 GJ is wrong. 「天津乙女は…男星の肩に依れり」 — YP makes the maiden the subject and the male star the shoulder, exactly as the Japanese does
S10 YP misidentifies the sleeping poet as the host GJ is wrong, and confused: the poet is the host and is asleep (「庭の主人に一語の礼なくて…詩人が寝顔を二人はしばし見とれぬ」)
S11 YP makes the weeper one young woman, not 乙女の星 and 恋人たち GJ is wrong. 「乙女の星はこれを見て…涙うかべ」 — she weeps alone; the lovers act together only in the next clause
S12 YP omits that the star is on her forehead a nitpick, not a content error
S15 「肩に垂るる黒髪」 is hair falling to the shoulders, not shoulder-length a nitpick
S01 the garden is "unbefitting the house in being broad", not "unusually spacious" a nitpick
S05, S09, S13, S14 mixed — one real omission (「青煙一抹」), one disputed gloss of 潯, two disagreements about tone rather than content not adjudicated either way

So two statements are both true and both go on the record. F2b fired on the registered criterion and the run is void by it — no override, and the criterion was registered before anything was dispatched. And the gate that voided it is not itself a reliable measure of how wrong the yardstick is: of the five verdicts read closely, two are real errors and three are the gate misreading either the Japanese or the description. What F2b's firing licenses is the yardstick is not dependable, not the yardstick is wrong four times in five.

One error GJ did not catch, found by the same hand-read: on S10 「淡紅色の霞につつまれて乙女 の星先に立ち」 has the maiden-star wrapped in the pale-pink haze and going first; YP has the smoke guiding her. A gate that misses that while firing on a forehead is a gate with a taste, not a threshold.

7. GA, GY, and what the restyle did

GA — the arms are comparable. F1 does not fire. DIFFERENT on 2 of 15 against a bar of 3. Both flags (S08, S09) are the same thing: 水色 rendered pale blue in R06 and water-colour in R08. That is a difference of wording inside a colour term, which the prompt's first clause excludes; the registered carve-out covers only a quoted-line flag on S11 or S15, so neither is carved out and both are counted — and F1 still does not fire. On a stricter reading it fires even less.

GY — the restyle moved content on 9 of 15. F2 fires. What it added is worth quoting, because it is a fact about what "restyle this, change nothing" does:

Asking a model to raise a description's register makes it add evaluative claims the description did not contain, and the claims it adds are appreciative. Anyone building a marked-prose control out of a restyle needs a content gate, and 9 of 15 is how often it bites.

8. Why stage 5 was not dispatched

F2b's registered consequence is the whole run is void, P1 included, because H2P is built on YP. H1, H2P and H2M are the 135 calls that compute P1, and P1 is the only primary they feed. Dispatching them would have produced numbers from a run already void by a criterion frozen a day earlier, and the only use for those numbers would have been a comparison the design does not license. They were not dispatched. That leaves ≈$0.40 unspent and leaves this run with nothing to say about whether the H1/H2P separation replicates on a third language family — stated plainly, because it was the second thing the run was for.

H1 alone is not a primary of this design, and running it on its own after seeing the gates fail would have been an unregistered analysis chosen with the results in view. It was not run either.

9. Verification, bodies and money

analysis/verify.py was written after the design was frozen, imports nothing from tools/ and nothing from analyse.py, and reads analysis/results.json as a claim. 1,594 checks, 0 failures. It recomputes every cell choice against its stored body, every §4 majority from the usable bodies, the P2 denominator from the raw record, the exact binomial critical values by exhaustive enumeration, and the cost total over all attempts including discarded ones.

Nine mutation tests, nine caught. Three of the design's eight named mutations address H2P /H2M cells or the P2 denominator, and this run has neither, so those three are recorded as not exercised and replaced by the same mutation on a stage that exists; the substitutions are labelled as such in the verifier's output. One of the original eight — a document swapped between YP and YM — was MISSED on the first pass, because nothing tied a prompt's embedded document back to the body that wrote it. That check was added (195 assertions of document provenance across GD, GS and GJ) and the mutation is now caught.

Two defects in the frozen design, both found by the verifier and both recorded rather than patched away (design §12 A1, A2, and one more here):

  1. §10.3's assertion that no GS prompt contains the words original, Japanese, source, effect, comparable or translation is unsatisfiable by construction. A GS prompt embeds a YP document, and YP's own §4 stage 1 instruction requires the word Japanese — "what the passage does to a reader of Japanese". Measured: 24 of 90 GS prompts carry a banned word inside the document block, across 5 segments, the words being effect, japanese, original. The verifier asserts the ban on the prompt scaffolding, where it holds absolutely, and counts the document-block contamination rather than waving it through. GS is therefore not blind in the way the design claims: in a quarter of its cells the seat could see that the two passages are translations from Japanese. A third independent reason P2 could not have been trusted here.
  2. §10.4's byte-identity claim across H1, H2P, H2M, which cannot hold because H1 asks a different question — scoped in A2 and moot, since stage 5 did not run.

Bodies. 261 dispatched, 0 dead, 266 attempts. F5 fires on no stage. Providers: OpenAI 62, xAI 62, Alibaba 60, Google 45, Google AI Studio 17, CoreWeave 14, Phala 1.

Money. $0.816705 for the run, plus $0.023320 of transport probes, $0.840025 in all, against a revised ceiling of $2.26 for stages 1–4. Key usage 107.830757299 → 108.692808434, delta $0.862051, leaving $0.022027 unreconciled in the over-counting direction — the same lag this session found at start-up on S181's closing snapshot ($0.265127 settling after that session closed). Per-request costs are primary (CLAUDE.md); the residual is recorded and not chased.

$0.044710 of the run — 5.5% — bought nothing, two discarded attempt sets, both preserved under runs/_discarded/ with their costs counted:

discarded why cost
yp__S01, 3 attempts at caps 1,600 / 3,200 / 6,400 note (bnk) on a third seat: reasoning: {"effort": "low"} sent and ignored, 4,662 reasoning tokens for 171 visible ones, then a body cut off mid-sentence with finish_reason: "stop" $0.043586
yp__S14, 1 attempt 1,200 tokens of bare prose, no JSON anywhere, finish_reason: "stop" — a parse failure F6 covers and the runner's first usability test called usable $0.001124

10. Limits, at the size of the claims

  1. Model seats, not readers. Tier D is NOT PASSED. Every figure is a fact about how five seats behave, and GD and GS are themselves unanchored model judgments about style — design §9 limit 5, which this run's §4 table sharpens rather than resolves.
  2. The YP-is-ornate finding is about one yardstick author on one text. qwen/qwen3.7-max, reasoning disabled, on Meiji 擬古文, 15 documents. It says nothing about whether a human, or a different seat, could have written a plain one — and nothing about the yardsticks used in RS-20260812f and RS-20260813b, which were written by a different seat on different languages and have never been measured on either property.
  3. The F2b count is the gate's, not the lead's. §6 gives the adjudication. Five of twelve were read closely; seven were not read at all, or not resolved.
  4. The SOURCEWARD-on-both finding rests on one segment and 9 cells, pre-registered as S07 before the run, but one segment nonetheless.
  5. This run says nothing about the confound it was built to test. Not that the confound is absent, not that it is present. The instrument that would have measured it was never in a state to do so, and failing to build a control is not evidence about what the control would have shown.
  6. Nor does it say anything about whether the H1/H2P separation replicates on Japonic. Stage 5 did not run.

11. What follows

  1. D-20260804-16 condition 3 does not discharge. The criterion that failed is criterion 2, clause 3, with F2b and F3 firing before it. wiki/goodness-senses.md §affect is rewritten to say so, and usage rule 4 — never report a single affect figure covering both halves — stands untouched, since it is not what condition 3 governs.
  2. ARM-affect-confound closes resolved at 2 of 2, on its own terms: "an inconclusive third state is possible and is a failure to discharge, not a licence to try again inside this arm."
  3. This is the second consecutive failure of this manipulation family, and the design pre-registered what that means: "two failures make the manipulation itself the finding." E-20260813b voided on its own manipulation check; E-20260813f voids on a source-fidelity gate and a directionality gate. A future design that wants to test the style confound must first solve the problem this run surfaced: producing a description of a foreign passage that three seats will call plain. Nothing in the two runs suggests asking for it works.
  4. The two prior affect results keep their direction and acquire a caution. RS-20260812f §5 and RS-20260813b §5 rest on yardstick documents that were never checked for accuracy against the source and never checked for style, and this run is the only evidence anywhere in the project about what such a check finds. The caution is recorded on wiki/goodness-senses.md §affect; it is not a withdrawal, because nothing here measured those documents.