Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260814h-footing-price/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260814h-footing-price
statusfrozen
created2026-08-14
updated2026-08-14
sensesnaturalness
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-footing-price.md, framework/v0.2/README.md, workshop/regimes/R27-footing-max.md, workshop/regimes/R28-footing-selective.md, workshop/translations/genji-yomogiu/R06-v2/translation.md, workshop/translations/genji-yomogiu/R27-v1/translation.md, workshop/translations/genji-yomogiu/R28-v1/translation.md, wiki/findings/results/RS-20260814b-honorific-hands.md, wiki/goodness-senses.md, config/models.md

E-20260814h — what do three blind model judges get for the words a translator spends on a footing mark?

AMENDED 2026-08-14 on a pre-run adversarial critic pass (P3, NEEDS REDESIGN, 20 findings, 2 BLOCKING) taken BEFORE any grading call was dispatched. §10 lists every finding and what was done with it. The four largest consequences: (i) the permutation minima were wrong and the tests are now one-sided on registered directions, which is what makes the five-segment comparison reachable at all; (ii) the old primary P1 is demoted to a manipulation check — a judge that can read does not need to detect footing to notice ladyship; (iii) the new primary is NAT(D) − NAT(C), which neither the lexicon nor the construction guarantees; (iv) the decoy's ornament was re-registered because it was colloquial where C is courtly, and the title of this page no longer says "a reader".

ARM-footing-price step 1, study limb. The translation limb — T-genji-yomogiu-R28-v1, the middle rung of the ladder, with its full carried/declined log — was frozen and committed at 1bb7a66 before a line of this design was written.

The wire between the limbs, in one sentence: the third rendering asks whether the price framework/v0.2 §10.5 measured is avoidable — whether the footing can be bought at a fraction of the cost — and this run is what prices all three renderings by three judges that did not write them.

1. Question

framework/v0.2 §10.5 is a price list one hand wrote about its own sentences. It says that carrying a Japanese source's grammatical footing into English is available at 51 of 52 sites and costs +21.1% length and 78 marked deference tokens where the baseline has none. Nobody has ever been asked whether any of that is audible.

Do blind judges shown only the English register that anything was bought — is what they register a social relation between the people in the passage or merely a marked style — and is the naturalness the maximal rendering pays a price of deference or a price of bulk?

The third clause is the primary and the first two are checks, for the reason the pre-run critic gave and §6 records: a judge that can read does not have to perceive footing in order to notice the word ladyship.

The second half is the whole reason the run has four arms rather than three. RS-20260809h established that the project's perceived-source-carriage measure cannot separate them: CLUNKY — one of the lead's own translations with its words permuted, a strict word-multiset permutation — scored +1.4583 above the text it permutes, higher than five of six real ruled arms. A design that puts a heavily marked text against a plain one and finds the marked one "marks a social relation" has measured markedness and will have learnt nothing. The decoy arm is what makes the answer mean anything, and it is the arm this design would be worthless without.

2. Design in one line

Four renderings of the identical 1,227 characters — plain, selective, maximal, and an ornament-matched decoy carrying no deference at all — cut into six matched segments and put to three blind panel seats, one segment per request, source-blind, arm-blind, neighbour-blind, answering two 1–7 scales and one closed category.

3. Materials — the four arms

arm text words footing devices what it is
A T-genji-yomogiu-R06-v2 836 0 the hand not asked. R06: source-only, single pass
B T-genji-yomogiu-R28-v1 865 (+3.5%) 12 asked, inside a budget. 12 of 52 sites carried, 40 declined in writing
C T-genji-yomogiu-R27-v1 1,012 (+21.1%) 78 asked, fluency subordinate. 51 of 52 sites carried
D materials/decoy.txt 1,013 (+177) 0 A + ornament, matched to C's volume, carrying no deference

How D was built, and the three things that make it a control rather than a second opinion. D is A with 73 short insertions of intensifying, emphatic and doubling material — steadily, exceedingly, quite, whatever, indeed, entirely, decidedly, none whatever, by no means, a great deal more — at the same places C carries its devices. build.py proves three things mechanically and refuses to write the cells if any fails: (i) D is a pure insertion on A (every one of A's 836 tokens appears in D, in order — D never replaces or cuts); (ii) D contains zero tokens from the deference device lists; (iii) D's added volume matches C's on both axes — +177 words against +176, 73 insertion runs against 77 deference tokens, and segment by segment to within two words in five of six segments.

Twelve of the 73 insertions were re-registered after the pre-run critic (§10 F3), which showed the original ornament was colloquial — sheer pique, utterly beaten, what in the world — where C's devices are courtly, so D and C differed in register as well as in deference. The replacements are length-matched and register-neutral (and pique alone, and wholly so, what, then,).

build.py also re-runs R28-v1/counts.py, so B's skeleton constraint is a precondition of this run existing: outside its twelve sites, B is A.

3.1 The segments

Six, defined by paragraph role and verified by anchor in all four arms, so that the same stretch of the story is compared across arms and nothing is cut mid-sentence:

seg paragraphs A B C D C's devices
G1 the Princess dresses 124 129 159 157 18
G2 he enters; his first speech 77 79 92 91 7
G3 the hanging; his second speech 142 153 181 181 14
G4 the narrator's aside; leaving 107 114 140 141 14
G5 poem, third speech, poem, her stir 170 170 204 205 14
G6 the moon; the Falling Flowers 216 220 236 236 10

G5 carries no R28 site, so G5's A and B are the same text under two different labels. This is not waste and it is registered as a measurement: see §6.4.

4. Procedure

  1. python3 build.py — checks the manipulation, writes cells.json (24 cells), assigns each cell an opaque pid from sha256("E-20260814h|<seg>|<arm>") and sorts by it, so arms are interleaved in the dispatch order and no seat meets them grouped.
  2. python3 run_critic.py — one non-Anthropic adversarial pre-run critic (P3 x-ai/grok-4.5), given the design and every one of the 24 passages, before any grading call. Note (bnr): the first call is sized at what is affordable, truncation is expected, and ~20% of the critic budget is held for a continuation call. The run does not proceed on a critic that returned no verdict line. Findings are written into §10 with what was done with each, and amendments are made before dispatch.
  3. python3 run.py — 24 cells × 3 seats = 72 bodies, one cell per request, temperature: 0, dispatched one at a time in a fixed recorded order. Judgment is not parallelised.
  4. python3 analyse.py — computes every number in the result page from run.json.
  5. python3 verify.py — recomputes every reported number by an independent path, imports nothing from analyse.py, and runs mutation tests. Note (bio): the mutations target the per-cell, per-seat scores, not the aggregates.

4.1 Seats

P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5, resolved from config/models.md at run time and logged as provenance. P5 deepseek/deepseek-v4-pro is excluded on any task shape — note (bne). P4 is excluded here because it is the run's most expensive seat and note (bnk) records it consuming a whole cap on hidden reasoning.

P3 writes the critic and grades. This is declared rather than hidden: the critic call sees the design and the passages; the grading calls are stateless, single-cell, arm-blind, and carry none of the design's vocabulary. A seat cannot recognise which arm a passage is if it is shown one passage. If P3's scores turn out to drive any primary on their own, the result page says so and the primary is reported both with and without P3.

4.2 The prompt, byte-identical across every cell but for the passage and its label

You are reading a passage from an English translation of a work of fiction. You are not
told what the original language is, who translated it, or where the passage comes from.

PASSAGE {pid}
------
{passage}
------

Answer with exactly four lines and nothing else.

NAT: an integer 1-7. How close is this to fluent, idiomatic, unmarked literary English
prose of the kind a contemporary publisher would print? 7 = entirely natural English;
1 = strongly marked, awkward or strained English. Judge the English on its own, without
speculating about any original.

REL: an integer 1-7. Setting aside what the passage tells you about who these people are
and what happens to them, how strongly does the WORDING ITSELF mark one person in the
passage as socially above another - deference, rank, station, one person raised or lowered
relative to another by the choice of words rather than by the events described?
7 = the wording marks it insistently; 1 = the wording marks it not at all.

WHO: one of M, F, B, N, U. M = the wording marks the man as the higher. F = the wording
marks the woman as the higher. B = the wording marks both of them as above the narrator
and the reader, without separating them. N = the wording marks no such relation.
U = cannot tell.

WHY: one sentence of at most 25 words, quoting the words that decided REL, or saying that
there were none.

NAT is naturalness (wiki/goodness-senses.md), scored on the target text alone under the post-D-20260802-13 wording, against the unmarked / literary-contemporary register anchor, which is the anchor the baseline's own log E1 names. REL is not a sense on that list and is not proposed as one. It is a descriptive item about what the wording marks, and it is deliberately not perceived-source-carriage: it asks nothing about a source, precisely because that sense's own entry records a 0.40–0.60 false-positive rate on manufactured oddity and RS-20260809h showed a word-scramble beating real prose on it.

Nothing here is a quality judgment and none is licensed. Tier D is NOT PASSED; every score is internal-judgment-only and provisional.

5. Analysis

Per cell, the seat-mean over the three seats. Comparisons are paired across the six segments and tested by exact permutation over sign flips, enumerated in full and never approximated.

Sidedness, fixed here because the first draft got it wrong (critic F8, BLOCKING). Every prediction below registers a direction before dispatch, so every test is one-sided. On six paired segments the smallest attainable one-sided P is 1/64 = 0.015625; on five, 1/32 = 0.03125. Both clear a 0.05 bar; the two-sided minima are 0.03125 and 0.0625, and the five-segment comparison is therefore unreachable two-sided, which is what the first draft claimed to be doing. Both P values are reported for every comparison, and no comparison is called significant on a two-sided P above 0.05 without saying so.

No multiplicity correction is applied across the four predictions, and with six paired segments a unanimous sign is the only route to any P below 0.05. This is a provisional internal pilot and the result page says so.

Every primary is reported twice: with all three seats, and with P3 excluded (critic F19) — P3 wrote the critique and saw all 24 passages, and although the grading calls are stateless the dual report costs nothing.

WHO is analysed as the rate at which a seat names a specific individual (M or F) rather than B, N or U, over the 18 cells of an arm.

6. Predictions, registered before dispatch

Restructured on critic findings F9 and F13. The first draft made REL(C) − REL(D) the primary and called it "the purchase". The critic's first BLOCKING finding is that this is close to a lexical tautology: REL asks in so many words about deference, rank, station, C is built out of her ladyship and was pleased to, and D is mechanically forbidden them. A seat that can read does not need to perceive footing to score C high. It is retained, and it is retained as a manipulation check.

The manipulation checks — registered, expected, and never reported as findings

M1. REL(C) − REL(D) ≥ 1.00, one-sided exact P ≤ 0.05 over six segments. The devices are audible as devices. Near-guaranteed by lexicon. If it fails, the seats are not reading the item as written and every other number on this page is withheld.

M2. NAT(A) − NAT(C) ≥ 1.00, one-sided exact P ≤ 0.05. The maximal arm is read as worse English. Near-tautological — R27 was written to be worse and says so — and demoted for the same reason S134's critic demoted the same shape. Its interest is entirely in the size.

P1 — PRIMARY. Is the naturalness price a price of deference, or a price of bulk?

NAT(D) − NAT(C) ≥ 0.50, one-sided exact P ≤ 0.05 over six segments.

C and D add 176 and 177 words to the same baseline, in the same places, and are matched segment by segment to within two words in five of six segments. They differ in what kind of material was added and in nothing else a seat can see. Neither the construction nor the lexicon fixes the answer, which is why this is the primary and the old P1 is not.

P2 — the middle rung, and the arm's reason to exist

REL(B) − REL(A) ≥ 0.50, one-sided exact P ≤ 0.05 over the five segments where B ≠ A.

Twelve marked tokens in 865 words — one every 72 — is a sixth of C's density, bought for +3.5% of length. Nothing in this project predicts whether that is audible. A failure here is the practitioner-relevant result, not a disappointment: it would say the choice is between an expensive purchase and none, with nothing usable in between.

P3 — WHO: relative grading, or aristocratic register?

Respecified on critic finding F14, which is correct and which the first draft's own §6 text contradicted. C elevates both principals — his lordship and her ladyship — so a run that predicted "more M or F under C" could fail exactly when the interesting thing happened.

P3a: the rate of any social marking — {M, F, B} against {N, U} — is higher under C than under A. Registered, one-sided. P3b, registered as the interpretive split and not as a bar: among C's socially-marked answers, is the modal answer B (both above the narrator and the reader) or M/F (one above the other)? A B-dominant answer is a positive finding, not a failure: it would say the English devices deliver aristocratic register where the Japanese delivers relative grading — a distinction framework/v0.2 §10 does not currently draw anywhere.

P4 — descriptive, and declared partly guaranteed

NAT(A) − NAT(B) ≤ 0.50. B adds 3.5% of words and changes nothing else. Reported, never counted toward any success narrative.

6.4 The noise floor, and the two conditions that withhold P2

G5's A and B are identical texts under different pid labels. The differences there estimate label sensitivity — same seat, same text, different four-character tag — and nothing wider. It is not a test–retest estimate. RS-20260804i measured a related and larger effect (a false attribution moved accuracy from 18–6 to 23–1), which is why it is measured rather than assumed zero.

Registered decision rules. P2 is WITHHELD, whatever its P, if either fires: 1. the G5 same-text |ΔREL| on the seat-mean is ≥ the P2 effect size; or 2. (critic F20) any single seat's G5 same-text |ΔREL| is ≥ 0.25, half the P2 bar.

6.5 One interpretive rule, registered because the critic showed the prompt invites the error

A WHY line that quotes only a rank-noun or a deferential auxiliary establishes that the device was noticed. It does not establish that a relative grading was perceived — that is what WHO is for, and the two are reported separately. Registered before dispatch (critic F5, F11).

7. Failure criteria, fixed before the run

8. What this design cannot establish, written before it is run

  1. One passage, one language pair, one hand, one work. Nothing here generalises to Russian ты/вы or to Bengali, both of which §10.1 records at total loss, and nothing generalises past this translator.
  2. The decoy was written by the hand that wants C to beat D. Its device list was fixed before it was written and its volume match is mechanical, but the choice of what ornament to insert is the lead's and cannot be blinded. A decoy written to be maximally unlike deference would flatter P1; the mitigation is that D is a pure insertion on A with a declared vocabulary, and the limitation stands on the result page.
  3. B is insertion, not composition (R28 §Known limitations). A translator working freely under B's instruction would have rebalanced the sentences around the marks it kept.
  4. REL has never been through Tier D. Neither has NAT under the revised wording, on this material. No score here is calibrated and no quality claim is licensed by any of them.
  5. Both published hands of this chapter are pre-1930 (§10.6), and they are not in this run at all — this run compares four texts by one 2026 hand and says nothing about what English does.
  6. Three seats are three models, not readers. Charter §4: AI-only convergence is weak evidence. The critic's F18 is accepted in full and governs every sentence this run may write: nothing here licenses the word reader, and the strongest entitled claim is about what three uncalibrated model judges detect on one passage. The title of this page was changed for that reason.
  7. B's device edits reword their clauses; they do not merely insert. counts.py proves B is A with exactly the ten declared spans substituted — the critic's F16 claim that B carries extra edits is factually wrong and is overruled on that evidence — but the underlying observation is right and stands here: a LEX carriage such as had left behind for her → had presented to her and left behind changes the clause, not only a word, so B differs from A by slightly more than twelve tokens' worth of style.
  8. D's ornament is register-matched by revision, not by construction. Twelve of its 73 insertions were re-registered after the critic showed the original set was colloquial where C is courtly (§10 F3). That is a better control than the first draft's and it is still a control the interested hand wrote.

9. Pre-flight cost estimate

Today's ledger (config/budget.md, UTC day 2026-08-14): $3.389393 of $5.00 spent across seven sessions; headroom $1.610607.

Worst case built from max_tokens, per note (abc) — the cap the request permits, not an assumed output length — at the prices config/models.md records:

stage seat calls cap worst case
critic P3 $2.00/$6.00 1 + 1 continuation 6,000 / 4,000 $0.0780
grading P1 $1.00/$6.00 24 800 $0.1310
grading P2 $0.75/$3.75 24 1,400 $0.1368
grading P3 $2.00/$6.00 24 800 $0.1440
74 $0.490

Note (bnk)/(bne): an effort pin is a request, not a bound, and a reasoning seat can consume the whole cap without emitting an answer. Declared ceiling for this run: $0.90, which is 56% of today's headroom and leaves it positive under a 1.8× overrun. If the grading stage passes $0.60 the run stops and is reported on the cells completed.

10. Pre-run critic findings

Seat P3 x-ai/grok-4.5, one call, finish_reason: stop, no continuation needed, $0.068592400. Verdict NEEDS REDESIGN. Raw body in critic.json. Twenty findings; eighteen accepted, two overruled with reasons. Every amendment below was made before a single grading call was dispatched, and cells.json was rebuilt afterwards.

# sev finding disposition
F8 BLOCKING The permutation minima are wrong. Two-sided sign-flip minimum on n=6 is 2/64 = 0.03125, not 1/64; on n=5 it is 0.0625, so the five-segment prediction's 0.05 bar was unreachable as written ACCEPTED, and it is the most useful finding in the pass. §5 now tests one-sided on registered directions, which is licensed because every direction is registered pre-dispatch; minima 0.015625 (n=6) and 0.03125 (n=5). Both P values are reported for every comparison
F9 BLOCKING The old primary was nearly lexical. REL names deference, rank, station; C is built from ladyship; D is mechanically forbidden them. A judge that can read scores C high without perceiving footing at all ACCEPTED IN FULL. REL(C) − REL(D) is demoted to manipulation check M1. The new primary is NAT(D) − NAT(C), which neither lexicon nor construction fixes
F3 MAJOR The ornament is colloquial where C is courtly — sheer pique, utterly beaten, what in the world, clung and clung fast, and waited — so D pulls register down while C pulls it up, flattering the old primary from both sides ACCEPTED. Twelve insertions re-registered, length-matched and register-neutral; build.py re-run and the volume match re-proved (+177 vs +176; 73 runs vs 77). The full swap list is in the commit that carries this amendment
F4 MAJOR Interested hand. The mechanical checks bind volume and vocabulary class, not which non-deference words were chosen — the one degree of freedom that decides whether D is a control or a foil ACCEPTED as a limitation that cannot be removed at this cost (§8.2, §8.8). Partially mitigated by F9's demotion: the new primary is a naturalness contrast, far less sensitive to which ornament was picked than REL is
F2 MAJOR Fairer controls exist: replace C's deference lemmas with length-matched neutral formality; or a second ornate rendering with no rank lexicon ACCEPTED and DEFERRED, named as ARM-footing-price step 2 candidate (a). Rebuilding D from C by closed-class replacement is the better control and it is a different, cheaper run
F5 MAJOR REL's prompt is an answer key, and WHY's "quote the words that decided it" invites title-harvesting; D has nothing quotable, so REL(D) is floored by prompt shape ACCEPTED, and handled by demotion plus a registered interpretive rule (§6.5): a WHY line quoting only a device establishes the device was noticed, never that a relative grading was perceived. The prompt itself is not changed — it is frozen, the critic read it as frozen, and rewriting a specification after a critic fires is the move note (bio) forbids
F6 MAJOR Bulk re-opens a channel only C can use: more words ⇒ more surface to quote under WHY ⇒ higher REL ACCEPTED, and answered by the same demotion; the D arm controls bulk for the primary
F14 MAJOR P5 (WHO) could "fail" exactly when the interesting thing happens — C elevates both principals, so a B-dominant answer is the theoretically important result and the old prediction scored it as a miss ACCEPTED IN FULL and respecified as P3a/P3b (§6.3): the bar is {M,F,B} against {N,U}, and a B-dominant split is registered as a positive finding
F13 MAJOR The interesting contrasts are B vs A and WHO, not C vs D; the design mis-labels which prediction makes the run meaningful ACCEPTED — see F9's disposition
F17 MAJOR REL is unvalidated (never through Tier D) and the primary rested on it — circular ACCEPTED, and materially answered: the primary now rests on naturalness, a sense on wiki/goodness-senses.md with a Tier D history. REL carries P2/P3 only, and §8.4 says neither is calibrated
F19 MAJOR The critic seat also grades. Dual-report should be the default, not a contingency ACCEPTED. Every primary is now reported with and without P3 unconditionally (§5)
F18 MAJOR The entitled sentences are weaker than the stated purpose — nothing here licenses "a reader", or generalisation past one passage, one hand, three models ACCEPTED IN FULL. The page title, §1's question and §8.6 are rewritten; the word reader is struck from every claim this run may make, and the same constraint binds the result page and the framework edit
F16 MAJOR B is not pure A + twelve devices: offer him an answer for answer, presented to her for left behind for her — substitutions leak style OVERRULED on the factual half, ACCEPTED on the underlying point. R28-v1/counts.py proves constructively that B is A with exactly the ten declared spans substituted, and each of those spans is a carried site — there are no extra edits. But the observation that a LEX carriage rewords its clause rather than inserting a word is right, and is now §8.7
F11 MINOR R28's rule does not secretly restate NAT — it is fidelity plus budget ACCEPTED (a negative finding, and the one the brief asked for by name under note (bio)). Recorded because a critic clearing the paraphrase question is evidence, not silence
F12 MINOR P2 (old) restates R27's construction; keep it a manipulation check and never narrate it as "the price" ACCEPTED. It is M2 and the word "price" is removed from it
F7 MINOR Run-matching is not perceptual matching: C repeats one formula, D scatters many hapax intensifiers ACCEPTED as a stated limitation. It is also, note, a finding the run may produce: R28's own log says the cost of selective carriage is repetition
F20 MINOR The G5 withhold rule should also fire on seat-level disagreement, not only the seat-mean ACCEPTED and added as §6.4 rule 2 (any single seat's G5 |ΔREL| ≥ 0.25)
F15 MINOR n=6, three model seats, no multiplicity correction — adequate for a provisional pilot only ACCEPTED and stated in §5
F10 MINOR The prompt tells seats not to speculate about a source while C is obvious court translationese; REL may be inflated by "this is a translation of a ranked society" ACCEPTED as a limitation that cannot be fixed without different devices (§8.5)
F1 MINOR Dead-cell rule, temperature: 0 and single-cell dispatch are sound; the cost ceiling is unrelated to validity NOTED. No change

What was NOT done, and why. The critic's minimum-amendment list asks for the REL prompt to be softened (item 7) and for D to be rebuilt from C by closed-class replacement (item 2). Neither is done. The prompt is frozen and was read as frozen; rewriting the instrument after a critic fires is exactly the move note (bio) forbids — amend the claim, not the freeze — so the claim was amended instead, by demoting the prediction the prompt was pre-loading. The C-derived decoy is a better control than the A-derived one and it is a different run: it is written into ARM-footing-price step 2 as candidate (a) rather than smuggled into this one.