Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260803f-craft-carriers/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260803f-craft-carriers
statusfrozen
created2026-08-03
updated2026-08-03
sensesaccuracy, naturalness, voice, style-correspondence, cultural-mediation, affect
purposeReaders of literary fiction in English who cannot read the source, meeting these texts as reading editions rather than as cribs (D-20260801-10).
internal-judgment-onlytrue
provisionaltrue
linksworkshop/translations/odnazhdy-osenyu/R04-v1/translation.md, workshop/translations/odnazhdy-osenyu/R14-v1/translation.md, workshop/translations/odnazhdy-osenyu/R06-v1/translation.md, workshop/regimes/R14-matched-flattening.md, wiki/arms/ARM-first-judgment.md, wiki/findings/results/RS-20260803-a4-set.md, wiki/findings/results/RS-20260802-tierD-verdict.md, wiki/goodness-senses.md, config/models.md

E-20260803f — what carries literary life, when nothing propositional differs

ARM-first-judgment step 3 of 3 (T3). FROZEN 2026-08-03 before any call was dispatched. Both translator's logs were frozen and committed before this file existed: T-odnazhdy-osenyu-R04-v1 at 827a63c, T-odnazhdy-osenyu-R14-v1 at d790968. The A4 freeze condition therefore holds by construction and is not re-argued (ARM-first-judgment §Constraints).

Standing, and it governs every number this design will produce. Tier D was run at S086 and NOT PASSED. No score here carries evidential weight, none may support a framework recommendation, and everything is provisional and internal-judgment-only.

1. The question, and what it teaches

What carries literary life in a translation, independently of what the translation says — and is any of it visible to an evaluation that scores six senses?

The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5), written before the unit was designed: this unit teaches which non-propositional properties of English prose a reader actually registers, by removing a named list of them from a translation and seeing which removals are noticed and named by readers who were told nothing about the list. That is a claim about translating literature. Its second limb — whether the project's own six-sense rating sees the difference — is ARM-first-judgment's clause 3, every result page states plainly what its scores license and what it does not, and is discharged here by demonstration rather than by assertion.

Why this is not "scoring more translations", which step 3 is explicitly not a licence to do. The A4 set scored five filed translations and kept the A4 promise. Nothing here adds to that count. The pair rated in stage 3 was built for this question and exists only to locate the boundary of what the A4 scores mean: Tier D established that the panel separates damage at ceiling; the A4 set established that it does not separate five competent translations. The interval between those two facts has never been measured, and a sentence about what the A4 scores license cannot be written honestly without it.

2. The wire between the limbs, in one sentence

The translation limb generates the object the study limb investigates: two renderings of the same 494 Russian words built to differ in nothing a paraphrase could report, so that any difference a reader finds is a difference in the writing and nothing else. Whether they succeed in that is what stage 1 decides, and this sentence does not pre-empt it (amendment A3, pre-run critic pass 1 finding 2, BLOCKING: the original wording asserted the conclusion the gate exists to test).

3. Materials, frozen

id what words
SOURCE Максим Горький, «Однажды осенью» (1895), ¶84–99, PD, ../../translations/odnazhdy-osenyu/R04-v1/source-ru.txt 494 (RU)
LIVE T-odnazhdy-osenyu-R04-v1, lead, R04, log frozen at 827a63c 669
FLAT T-odnazhdy-osenyu-R14-v1, lead, R14 v0.1, 37 operator sites, log frozen at d790968, repaired at four sites by amendment A1 688

Both English texts are extracted from the frozen pages by materials/extract.py, which writes LIVE.txt and FLAT.txt and recomputes every length figure. Length ratio 1.0284 after A1 (0.9925 before) — FLAT is 2.8% longer, still well inside the ±15% declared tolerance and smaller than the +4.5% confound declared on this project's damage control (RS-20260802-tierD-verdict). Sentence-length SD falls 24.0 → 7.7 under the strict splitter; both splitters are reported because the mean reverses between them and the SD does not.

3a. Amendment A1 — the pre-run critic's first BLOCKING finding, and the repair it forced

Pass 1 (P4, NEEDS-AMENDMENT, six findings, two BLOCKING, all six accepted) found that the F4 operator as written DELETED content rather than restating it, and that the equivalence prompt's own exclusion list would then have told the gate seats to suppress exactly the items FC1 exists to catch. It named four sites from the frozen text. A mechanical check, materials/coverage.py, was then written and run before any further dispatch — every content word present in LIVE and absent from FLAT, exact and after crude stemming — because the finding was computable from the frozen materials alone (note (bhr)).

site the deletion repaired to
¶84 "like an owl!" — the simile of «как сыч» — gone with no replacement, and not in the site table at all "…say nothing, like an owl."
¶89 "The wind howled and moaned" → "was blowing loudly": the moaning erased "was blowing and making a howling and moaning sound"
¶89 "the rain drummed" → "was falling on": the drumming erased "was falling on the boat and making a drumming sound"
¶89 "the waves splashed" → "were moving": the splashing erased "were moving and splashing"
¶93 "Many kisses, past counting, and hot" → "many times and warmly": uncountability and heat both reduced "many times, more than could be counted, and her kisses were hot"

R14's F4 definition was wrong and is amended, not worked around: a deadening operator that deletes is a damage operator. F4 now replaces the concrete verb with a generic verb plus an explicit statement of the manner it carried — the manner is stated rather than enacted, which is the property the operator was for, and nothing is lost.

This is the finding of the run so far and it cost $0.10 to buy. The lead wrote a 37-site log believing the operator was propositionally conservative, and it was not, at five sites, one of which never reached the log. Note (bic).

The operator, in one line: F1 cadence levelling ×5 · F2 figure de-specification ×6 · F3 repetition flattening ×3 · F4 verb deadening ×3 · F5 register levelling ×11 · F6 connective explicitation ×9. Definitions in workshop/regimes/R14-matched-flattening.md; the 37 sites in FLAT's log; the tally parsed from that log by analysis/checks.py, never counted by hand.

4. Seats

Panel roles per config/models.md. Every call is stateless and no seat carries information between calls; overlaps below are therefore declared, not hidden.

stage seats why these
0 pre-run critic P4 moonshotai/kimi-k3 takes no measured role anywhere in this run. effort: low on the first dispatch, note (b)
1 equivalence gate P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 factual adjudication, where S015 measured the panel strongest (20/20 planted false claims rejected). Neither has seen the design
2a naming (PRIMARY) P1, P3, P5 × 2 orders three labs; free text, no categories supplied
2b coding of 2a's free text P2, one call per response P2 is the only seat that is not a namer, so no seat codes its own words. The lead codes the same 36 cells independently as a declared internal-judgment-only check; P2's coding is the primary and the lead's is never substituted for it
3 six-sense rating P1, P2, P5 × 2 texts × 2 passes the A4 set's own jurors, on the A4 set's own prompt format, so RS-20260803-a4-set's measured floors transfer: retest 0.233, paraphrase 0.139, gross damage 4.333 on accuracy

5. Procedure

Stage 1 — the equivalence gate, run first and before anything else is dispatched. Each seat is given SOURCE, LIVE and FLAT and asked to list every place the two English texts differ in propositional content, tagging each item from a fixed vocabulary: negation | quantity | referent | tense-aspect | added-content | removed-content | lexical-specificity | force | other. The prompt states that differences of rhythm, register, figure, punctuation and sentence division are not propositional differences and must not be listed. Order of presentation is swapped between the two seats.

Stage 2a — the naming task, the primary. Each seat sees the two English texts as A and B, in both orders, with no source, no authorship, and no categories: "These are two English renderings of the same passage of Russian prose. Describe, as specifically as you can, how they differ as pieces of writing. Quote from both." Nothing is asked about quality and nothing is asked about the source.

Stage 2b — coding. P2 receives one stage-2a response at a time, together with the six operator definitions (not the site table, not the texts), and answers yes/no per category with the quoted words that name it. Six calls, one per response. The lead codes the same 36 cells from the stored responses, blind to P2's output, as a declared check.

Stage 3 — the six-sense rating. The A4 protocol, unaltered: the juror sees SOURCE and one English text, rates it 1–7 on the six senses of wiki/goodness-senses.md, alone, with authorship stripped and no comparison available. Two passes per (text, juror).

Control C1 — the naming false-alarm floor. Each of the three naming seats additionally receives LIVE against LIVE under the identical stage-2a prompt. A seat that manufactures differences between identical texts cannot be read as having found them between different ones. Three calls.

Judgment is never parallelized. Stages run in order, and stage 3 is not dispatched until stage 1 has returned AND FC1 has been evaluated — amendment A8: "read" is not "adjudicated", and stage 3 is the $0.16 a fired gate exists to protect.

6. Registered predictions

# prediction threshold
P1 The operator set is nameable. Blind seats, given no categories, name most of it ≥ 4 of the 6 categories named by ≥ 2 of 3 seats (a category counts for a seat if named in either order)
P2 Presence-differences are visible and absence-differences are not. Cadence and register differences are on the page in both texts; a de-specified metaphor leaves nothing marked to see. A9 (pass-2 finding N3): the original rationale said "leaves nothing to see", and after A1 that is false of this stimulus — every F2 site now replaces the figure with a literal statement that is present on the page, so the mechanism under test is salience, not absence. A5, registered asymmetry: with three seats the second limb can only fail on 3-of-3, so it has almost no power to be wrong. It is reported as descriptive and is not counted toward the run's primary F1 and F5 named by 3 of 3 seats; F2 named by ≤ 2 of 3
P3 The six-sense rating separates the pair, but far below damage. A6: all six gaps are reported, and the max-over-six selection is named in the result's licence sentence — selecting the largest of six and testing it against a floor inflates the pass rate under noise pooled |LIVE − FLAT| on the largest-gap sense ≥ 0.233 (the A4 retest floor) and ≤ 2.00 (BAR-D's accuracy signal is 4.333)
P4 accuracy does not separate them — the operator was propositionally conservative, checked by the one sense Tier D showed this panel detects at ceiling |LIVE − FLAT| on accuracy ≤ 0.50
P5 naturalness moves toward FLAT. Under its post-D-20260802-13 wording — distance from unmarked literary-contemporary English, on the target alone — the flattened text is by construction nearer the unmarked point FLAT ≥ LIVE on naturalness

P5 is the prediction that matters to the typology and it is registered before any number exists. If the flattened text is called more natural by the rating while the naming seats describe it as the duller piece of writing, that is a fact about the sense, not about the panel — and it is the first direct evidence on whether the struck escape clause left naturalness able to reward flatness.

7. Failure criteria

8. What this run cannot show, registered before it ran

  1. The operator is the lead's own taste. F1–F6 are what the lead believes carries literary life. A high P1 shows the categories are nameable by others, not that they are the right categories. Nothing here surveys the space of craft properties; six were chosen and six were tested.
  2. The readers are language models. Naming a cadence difference is not evidence that a human reader would feel it. The demand pathway is live: a seat asked how two texts differ will look for differences.
  3. One locus, one language pair, one translator, one direction of flattening. No claim about Russian, about Gorky, or about translation in general follows from 669 words.
  4. F5 changes illocutionary force, which is not propositional content but is not nothing. Eleven of thirty-seven sites are F5. The equivalence gate is instructed on propositional content specifically, and the reading of the whole run has to hold that.
  5. Tier D is NOT PASSED. Stage 3's numbers are a description of what this panel does, not a measurement of how good either text is.
  6. The A4 floors are imported, not re-measured within this run — same jurors, same prompt format, different text. Stage 3's two passes give a within-run retest floor as well, and both are reported; if they disagree the imported one is not used.

9. Cost

Worst case built from max_tokens, not from an assumed output length (note (abc)).

stage calls cap worst case
0 critic 2 6,000 $0.234
1 equivalence 2 3,000 $0.056
2a naming + C1 9 2,000 $0.14
2b coding 6 1,500 $0.084
3 rating 12 1,500 $0.16
retry reserve — — $0.25
total 31 $0.92

Today's UTC ledger before this run: $2.517 of $5.00 spent (S094–S098 plus this session's ratification gate at $0.0556), $2.483 headroom. The reservation fits. P5's list price is not what gets billed — the S022 caution — so P5 is priced at 4× list here.