Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260805c-r1-polish/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260805c-r1-polish
statusfrozen
created2026-08-05
updated2026-08-05
sensesstyle-correspondence, accuracy, voice, cultural-mediation, naturalness
provisionaltrue
linkswiki/arms/ARM-r1-fresh-pair.md, framework/v0.1/README.md, workshop/translations/kamizelka/R04-v1/translation.md, workshop/translations/kamizelka/R06-v1/translation.md, wiki/findings/results/RS-20260802e-displaced-marking.md, wiki/findings/results/RS-20260804-yardstick.md, wiki/findings/results/RS-20260804g-yardstick-holds.md, workshop/experiments/E-20260804g-yardstick-repair/design.md, config/models.md, config/budget.md, wiki/goodness-senses.md

E-20260805c — R1 in a language it has never been tried on, and a gate on what counts as a site

ARM-r1-fresh-pair step 1 (T5). Frozen before any seat is addressed. The translation limb — T-kamizelka-R04-v1, the whole of Prus's «Kamizelka», with its log — was frozen first, at 59a5879, and the FORCED and DECOY spans below were written and committed before the yardstick seat was dispatched. Commit order is the guarantee and the verifier checks it.

1. What is being asked

framework/v0.1 §3 registers, as the first of four predictions and the only one that tests the release's only recommendation on new material:

On fresh Class A sites in a new pair, R1 recovers a marking at more than half. Score by the E-20260802e procedure. The failure it should not survive: recovery at or below a third.

R1 itself:

Where the source marks a relation or attitude by a grammatical form the target lacks, the absence of a same-category counterpart is not the absence of the marking. Render the site again under a brief that requires the marking to appear, and let it fall wherever the target does mark such things. Record the loss only if that second attempt fails.

The prediction has been open since the release and two runs have failed to close it. E-20260803c (DE→EN) hid the source and is not the named procedure. E-20260804 (FR→EN) was the named procedure and fired three of its own failure criteria, so it is descriptive only. S106 then removed the explanation S101 offered — RS-20260804g varied the yardstick's authorship alone and R1's separation reproduced at +4 — so the FR→EN null is currently unexplained.

This run's hypothesis about that null is stated now, before any data, because it is what the design is built around. E-20260804 fired F2 (FROZEN recovered at 4 of 7) and F5 (an independent seat's plain translation recovered at 4 of 7). At most of its putative Class A sites the relation was carried by any competent English at all. If there was no loss, there was nothing for R1 to repair, and a run on such sites cannot test R1 whatever it returns.

2. The subject-rule sentence (wiki/tracks.md, continue-prompt.md §4.5)

What does this unit teach about translating literature or evaluating translations? — Whether the one piece of advice this project gives a translator transfers to a language it has never been tried on: when Polish marks a social relation by a grammatical form English does not have — the intimate ty said downward to a tradesman, the diminutive that makes a grown man small, the impersonal past with no one in it — does re-rendering the site so the marking falls on a device English does have put the relation into the English, or does the loss stand?

This is a principal unit and not method work: what is measured is a property of translations. The one apparatus question it settles — what makes a site a site at which the question can be asked — is a gate inside the unit (§4), timeboxed to it, and its result is reported on the result page, not carried into the backlog.

3. Materials — the census, and how it was built

The whole story was translated first, close, under R04, with the log frozen before this design was written. The census was then read off the log: every place where the log records a Polish grammatical form that English has no counterpart for and whose content is social, attitudinal or perspectival. materials/sites.json holds ten candidate sites, S1–S10.

Excluded by design, and named so the exclusion is not invisible:

Contamination. Declared on the translation artifact and not measured: three English renderings exist in print (Jopson 1930–31, Scherer 1955, Sues 1960) and none is freely reachable, so tools/dependence_check.py could not be run. Under CLAUDE.md's standing rule that is said rather than passed over. The design's validity does not turn on it: the claim under test is existential — does an English rendering exist that carries the relation — and a rendering that reproduced a published translator's device would still be an existence proof. Note (bhb) cuts toward the null and is the design's friend: the lead matches itself at up to 37 contiguous tokens across sessions, so a FORCED span that merely reproduces its FROZEN span marks nothing.

4. Stage A — the yardstick, and then the admission gate

A1/A2 — what the source conveys, written by seats shown no English

A1 = P5 deepseek/deepseek-v4-pro, used in no other role. Per site it receives the Polish passage, a staging note of bare narrative fact (who is present, what is happening), and a pointer that quotes the Polish tokens under study and says nothing about them. It receives no English rendering of any kind — not FROZEN, not FORCED, not the gloss, not the log. It returns one sentence stating what the passage conveys over and above what its words denote.

A2 = qwen/qwen3.7-max, the panel's documented first reserve, gets the byte-identical prompt. A2's statements are not graded. They are (i) a stability measurement — is "what this passage conveys" a stable object across two independent readers — coded SAME/OVERLAPPING/DIFFERENT and committed before any grading seat is addressed; and (ii) the declared fallback if A1 fails both attempts.

The leak screen (analysis/screen.py, committed with this design, lexicon inherited verbatim from E-20260804g §4): a statement naming the source's grammatical device tells the graders the answer. A statement that fires is re-requested once with the rule quoted; a statement that fires twice puts its site outside every denominator, named.

The admission gate — new here, and it is the design's whole point

N1 = P4 moonshotai/kimi-k3, used in no other role, is shown the Polish source and the staging note and no English, and writes a plain English rendering of each span: ordinary English, no brief, no knowledge that any study exists. This is the NEUTRAL arm of E-20260802e, produced here first.

Then the three grading seats are shown, per site, A1's relation statement and N1's plain rendering alone, one item per site, and answer YES/NO: would a reader with no access to the source come away with the relation? A site at which ≥ 2 of 3 seats say YES is EXCLUDED before the main run.

Why this is a gate and not a result. R1 is advice about sites where a marking is actually lost. E-20260804 discovered, after paying for a full run, that most of its sites were not such sites. The gate moves that discovery in front of the spend. The count it removes is registered as a prediction (§7 G1) and is reported whatever it is — if it removes most of the census, that is the run's finding and the FR→EN null gets its explanation.

What the gate costs, stated rather than hidden. (i) The NEUTRAL arm is the selector, so it is not carried into the main run and no FORCED-versus-NEUTRAL comparison is available in this design. E-20260804's F5 question is therefore not re-asked here; the gate is its measurement. (ii) DECOY is a lead-written unmarked paraphrase and is not the selector, but on admitted sites it is expected low partly by construction, so F1's DECOY leg is weakened and WRONG carries the discrimination control. (iii) FROZEN is not the selector — it is the lead's close rendering, written months of sessions of habit away from N1's — so F2 remains a genuine unselected control, and it is the one that decides whether the admitted sites were really sites.

5. The arms, and who wrote each one knowing what

arm who wrote it knowing committed
FROZEN lead nothing; written as translation, to no test 59a5879, before this design
FORCED lead, under R1's brief the source and its own log — not A1's statement with this design, before stage A
DECOY lead a length-matched paraphrase required to mark nothing with this design, before stage A
REAIM lead, under R1's brief the source and A1's statement after stage A, declared
POSITIVE lead an explicit gloss authored from A1's statement after stage A
WRONG lead an explicit gloss of a different, plausible relation after stage A

FORCED is written before A1 and that is deliberate. R1 is advice to a translator, and a translator has no oracle telling them what an independent reader will say the passage conveys. A FORCED arm aimed at A1's statement would test whether the lead can hit a target it has been shown. REAIM is that upper bound, kept separate and never reported as R1's score: it answers only is the relation A1 names reachable in English at all by ordinary narrative means. A REAIM span that states the relation instead of carrying it is a failed REAIM; that is what POSITIVE is for.

6. Stage B — grading

Form inherited from E-20260802e §4 and E-20260804g §6.

  1. Three non-Anthropic seats — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5.
  2. Each receives, per admitted site: the source, the gloss, A1's relation statement, and the five renderings unlabelled (FROZEN, FORCED, DECOY, REAIM, POSITIVE, WRONG — six), and answers for each: would an English-speaking reader with no access to the source come away with the RELATION — YES or NO, plus one line of reason.
  3. Presentation order per site is fixed by sha256(site id | seat id | ordering), never chosen by the lead. Two orderings per seat, the second the reverse of the first — six bodies.
  4. A seat counts as recovering the relation only if it answers YES in both of its orderings.
  5. A site counts as recovered under a condition when ≥ 2 of 3 seats recover it.

7. Predictions — registered, before any call

# prediction comparator
G1 (the gate) the admission gate excludes at least a third of the ten candidate sites E-20260804: 4 of 7 by F5
P1p (primary; this is framework/v0.1 prediction 1) FORCED recovered at more than half of admitted sites E-20260802e 5 of 6
P2p FROZEN recovered at ≤ 1 admitted site 0 of 6
P3p (FORCED sites) − (FROZEN sites) ≥ half of n +5
P4p DECOY recovered at ≤ 1 site 0 of 6
P5p POSITIVE recovered at n or n−1 sites 6 of 6
P6p WRONG recovered at 0 sites 0 of 7 (FR→EN)
P7p REAIM recovered at ≥ FORCED new arm

The lead's stated expectation, recorded so it can be wrong. I expect G1 to hold and to be the session's finding: I expect the gate to remove four or five of the ten, and the survivors to be concentrated in S1, S2, S3, S6 and S9 — the intimate ty downward, the two attitudinal diminutives, the agentless past and urzędniczek. I expect S4 and S8 to be excluded, because the log already notes that a little rouble and the lady's husband carry their constructions into English without help. On the survivors I expect P1p to hold at 3 or 4 of 5. If P1p fails on admitted sites, that is the first evidence against R1 from a site population where R1's precondition demonstrably holds, and it is the expensive result.

8. Failure criteria — pre-committed

A null is a result, and here is what each null means. If P1p fails on a properly gated population, framework/v0.1 §3 prediction 1 is falsified in a new pair and §2's traceability paragraph is qualified in place: R1's move did not recover the marking where the marking was demonstrably lost. R1's text would not thereby be refuted — it would be left standing with a declared pair where it does not work, which is the honest reading and is what §7's pair table is for.

9. What this design cannot show

10. Verification

analysis/verify.py, importing nothing from tools/: asserts every FROZEN span byte-identical to workshop/translations/kamizelka/R04-v1/translation.md and every source span byte-identical to workshop/translations/kamizelka/source-pl.txt (whitespace-normalised); asserts by git commit order that materials/sites.json carrying FORCED and DECOY was committed before the first stage-A body; recomputes the hash-derived rotations; recomputes the gate, all eight predictions and all six failure criteria from the raw stored bodies; recomputes the leak screen; re-sums billed cost from the stored bodies against the key-usage delta. At least three mutation tests, each asserting the bytes on disk actually changed (note (bgu)) and restoring them after (note (bhd)).

11. Budget

Declared worst case $1.35, built from max_tokens and never from expected output (note (abc)), with the S022 routing margin.

stage calls max_tokens worst case
stage 0 — seat probe at the run ceiling (note (bhf)/(bit)) 5 12,000 $0.01
pre-run adversarial critic (z-ai/glm-5.2) 1 (+1 reserve) 24,000 $0.30
stage A — A1, A2, N1, 1 retry reserve each 3 (+3) 12,000 $0.30
the admission gate — 3 seats × 1 call 3 12,000 $0.25
stage B — 3 seats × 2 orderings 6 12,000 $0.49

Today's UTC headroom before this run: $3.926537656 of the $5.00 cap, two sessions already run. Lead translation is free and is never ledgered (charter §3, A4): the whole of «Kamizelka» twice, FORCED, DECOY, REAIM, POSITIVE, WRONG, the census, the coding and all verification are lead work at no API cost.


12. Amendments A1–A7, from the pre-run critic (critic.md)

Seat: z-ai/glm-5.2, max_tokens 24,000, used in no other role, one body, 43.8 s, $0.02569214. Nine findings, two BLOCKING, VERDICT: NEEDS-REDESIGN. All nine are accepted, six in full and one in part. Applied before any stage-A body was dispatched; the stage-0 seat probe (5 of 5 alive, $0.002) is the only API call that preceded this section.

A1 — the gate selects on FROZEN, not on N1 (findings 1 and 3, BLOCKING, accepted)

The design gated on the wrong loss and the critic is right. R1's own text says "Record the loss only if that second attempt fails" — the first attempt is the translator's close rendering. So the loss that makes a site a Class A site is FROZEN's loss, not an independent seat's. Gating on N1 also excluded exactly the sites where R1 is most testable (finding 3): a site where FROZEN fails and N1 succeeds is a real loss for the translator and was going to be thrown away.

Amended: a site is admitted iff FROZEN is NOT recovered (fewer than 2 of 3 seats). The gate is applied as pre-registered stratification computed from the main run's own bodies, fixed here before any body exists — which is statistically identical to a separate gate stage, costs one fewer stage, and stops the grading seats seeing the FROZEN span twice.

What this costs, stated: FROZEN is now the selector, so F2 is withdrawn as a failure criterion and P2p and P3p are withdrawn as predictions — they are true by construction. The FROZEN exclusion count becomes the gate result and is reported whatever it is. N1's plain rendering (NEUTRAL) is retained as a secondary measurement and is not a selector.

What this does not repair, and the design now says so rather than implying otherwise. Finding 1's deeper claim is that the run tests whether a briefed bilingual can write English carrying a relation it already understands. That is what R1 claims is possible, so it is the right object; but a positive result is a claim about availability, never about whether a translator would find the device unprompted. §9 already forbids the second reading and A5 below adds the arm that separates R1 does not work from the lead wrote a poor span.

A2 — DECOY withdrawn, PARA substituted (finding 2, BLOCKING, accepted)

A control constructed to mark nothing cannot fail, and P4p was therefore not a prediction. DECOY is withdrawn. In its place: PARA — a free paraphrase of the FROZEN span, written by a seat (N1), with no brief about any relation and no instruction to avoid marking. Its recovery rate is genuinely unknown, so it is a falsifiable control on "any rewrite whatever helps" — which is the comparison R1's advantage has to survive.

The lead-written DECOY spans stay in materials/sites.json, unused, and are shown to no seat. The pre-stage-A commitment of FORCED is untouched and the verifier still checks it.

A3 — the arm count (finding 5, SERIOUS, accepted)

§6 said "five renderings" and listed six. Seven arms, and every seat sees all seven per site: FROZEN, FORCED, FORCEDM, PARA, REAIM, POSITIVE, WRONG. The ordering hash, the grading prompt and the verifier all operate on seven.

A4 — the thresholds stop moving with n (finding 6, SERIOUS, accepted)

"More than half of admitted sites" passes at 3 of 4 (75%) while the comparator is 5 of 6 (83%). Amended, in fixed units:

A5 — a second, independent implementation of R1 (finding 7, SERIOUS, accepted)

One lead-authored FORCED confounds R1 does not work with the lead chose badly at this site, and the critic named two spans it doubts (S2's my Jew, S6's went under the hammer). Added arm FORCEDM: a seat is given R1's text verbatim, the source, the gloss and the staging note — and not A1's statement — and renders the span so the marking appears. Where FORCED fails and FORCEDM succeeds, the site is reported as a lead-quality failure and not as evidence against R1; where both fail, it is evidence against R1 at that site.

A6 — WRONG's plausibility is rated by somebody else (finding 8, MINOR, accepted)

WRONG is the sole discrimination control and its plausibility was set by the lead. N1 — which has never seen A1's statement — rates each WRONG span, before grading, for how plausibly it states something the source passage conveys, 1–5. The ratings are committed before dispatch and reported beside the result. A WRONG arm that rates ≤ 2 across the board makes F1's WRONG leg a trivial pass and the result page says so.

A7 — the staging-note leak, screened rather than re-authored (findings 4 and 9, accepted in part)

Accepted and mostly dissolved by A1: with the gate on FROZEN, the staging notes no longer feed the gate at all, which is the substance of both findings. The residual — A1 reads lead-written staging notes — is answered by running the leak screen over the staging notes and the pointers themselves and committing the result before dispatch. Run at this commit: 0 of 10 staging notes and 0 of 10 pointers fire the banned lexicon. That is a mechanical check on grammatical terminology, not proof the notes are neutral, and §9's declared leak stands.

Not accepted: having a seat author the staging notes. A seat that has not read the story cannot write who is present and what is happening, and a seat that has read it introduces a fresh and larger confound. The limitation is declared instead.

Seat allocation after the amendments, with the no-self-judgment rule checked

role seat grades?
A1 — the yardstick, sees no English P5 deepseek/deepseek-v4-pro no
A2 — independent yardstick, stability + declared fallback qwen/qwen3.7-max no
N1 — NEUTRAL plain rendering, PARA paraphrase, WRONG plausibility rating P4 moonshotai/kimi-k3 no
M1 — FORCEDM, R1-briefed, never shown A1's statement qwen/qwen3.7-max, independent call no
graders P1 gpt-5.6-terra, P2 gemini-3.6-flash, P3 grok-4.5 yes
critic z-ai/glm-5.2 no

Registered contingency: A2 and M1 are the same slug in independent calls, which is not a leak because neither prompt contains the other's output. But if A1 fails both attempts and A2 becomes the yardstick, FORCEDM is authored by the same model as the yardstick and FORCEDM is dropped from the run. Fixed here, before dispatch.

Revised budget

Worst case rises from $1.35 to $1.55: the separate gate stage is removed (−$0.25), N1 gains two calls and M1 one (+$0.25), and the grading bodies carry seven arms instead of six (+$0.20). Today's headroom is $3.926537656.