Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260805h-content-or-marking/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260805h-content-or-marking
statusfrozen
created2026-08-05
updated2026-08-05
sensesstyle-correspondence, voice, cultural-mediation, accuracy
provisionaltrue
internal-judgment-onlyfalse
linkswiki/arms/ARM-r1-fresh-pair.md, wiki/findings/results/RS-20260805c-no-loss-to-repair.md, workshop/experiments/E-20260805c-r1-polish/design.md, framework/v0.1/README.md, wiki/findings/results/RS-20260804-yardstick.md, wiki/findings/results/RS-20260804g-yardstick-holds.md, wiki/findings/results/RS-20260802e-displaced-marking.md, wiki/tracks.md, config/models.md, config/budget.md

E-20260805h — content or marking: what the E-20260802e procedure actually responds to

ARM-r1-fresh-pair step 2 (T5), the arm's closing step. Session S116, 2026-08-05 UTC.

Frozen before any dispatch. Nothing below is edited after the first API call except by a numbered amendment, dated, with its reason and the finding that forced it.

1. The question, and why it is this one

framework/v0.1 §3 prediction 1 says: on fresh Class A sites in a new pair, R1 recovers a marking at more than half, scored by the E-20260802e procedure. Three runs have now failed to discharge it. The third (E-20260805c, S111, PL→EN) failed in the most informative way: its pre-registered admission gate admitted 0 of 8 sites, because the filed close translation already conveyed, to independent readers, the relation an independent reader stated from the Polish alone.

RS-20260805c §4(b) put a second reading on the record and ARM-r1-fresh-pair constraint 6 made it binding on this step:

The procedure scores content, not marking. Six of seven arms preserve the passage's content and only the seventh (WRONG) changes it; WRONG is the only arm that fails. An unbriefed free paraphrase recovered the relation at 8 of 8. Step 2 may not treat a high recovery rate as evidence that a marking is carried until an arm exists that varies the marking with the content held fixed.

This run builds that arm. It is the one measurement that decides whether prediction 1 is testable by its own named procedure, and it is the arm's declared closing condition.

The subject-rule sentence, written before the design

What does this unit teach about translating literature or evaluating translations? — It measures whether a reader's recovery of a social relation from a translation comes from the relation being marked in the English, or from the passage's facts alone: whether the diminutive, the T-form and the honorific that a translator records as losses are doing any work a reader can be shown to detect.

The unit is a principal unit under the subject rule's named exception (wiki/tracks.md): a named deliverable is blocked — framework/v0.1's only recommendation has an undischarged first prediction, and this measurement decides whether it can ever be discharged as worded. The arm says so in its first line.

2. The design in one paragraph

Take the six sites from E-20260805c at which the relation is separable from the propositional content. For each, put to independent readers a minimal pair: the filed close translation (FROZEN, frozen at S111), and a rendering that states every proposition of the source and carries none of the marking (STRIP). Grade both by the E-20260802e procedure, against the same frozen relation statements written at S111 by a seat shown no English at all. If the graders cannot tell the pair apart, the procedure responds to content and not to marking, and prediction 1 is untestable by it.

The wire between the limbs, in one sentence. The translation limb produces the unmarked member of each minimal pair and the proposition inventory that licenses calling it content-identical; the study limb asks whether the project's own grading procedure — and therefore its released framework's first prediction — can tell that member from the marked one.

3. Materials, and what is inherited unchanged

Source. Bolesław Prus, «Kamizelka» (1882), public domain; the same text as E-20260805c.

Inherited without modification, and this is load-bearing:

Reusing the statements and the two controls verbatim makes this run a direct replication of the instrument at S111: POSITIVE scored 48 of 48 judgements and WRONG 4 of 48. Those two numbers are predicted again below, and if they do not come back the primary is withheld.

S5 and S8 remain excluded. They failed the frozen leak screen on both passes at S111 and the frozen F3 rule removed them from every denominator. Nothing here reopens that.

4. Two of the eight sites are excluded, and the demonstration is part of the result

A minimal pair exists only where the relation can be removed without changing what the sentence says. At two sites it cannot, for two different reasons, and both were established by attempting the strip, not by assertion. Both exclusions are declared here, before dispatch.

S1 — the relation is entailed by the propositions

Polish: — No, jeżeli nie możesz oddać za pół rubla, to już idź. Ja więcej nie dam. Propositions: (i) condition — you cannot let it go for half a rouble; (ii) directive — go; (iii) assertion — I shall not give more. The frozen relation statement is "a final ultimatum … the transaction will end the moment the dealer refuses." That is entailed by (ii) and (iii) together. The attempted strip, written at S111 before the statement existed (sites.json DECOY), reads "then that is the end of the matter. I shall not be giving any more" — it drops the directive and strengthens the finality. No rendering was found that keeps (i)–(iii) and loses the ultimatum. Excluded: not a minimal pair.

S6 — English has the device, so the site is not Class A

Polish: sprzedano — the impersonal past with no grammatical agent. R1 is addressed to sites where the target lacks a same-category counterpart. English has one: the agentless passive, which is exactly what the filed translation uses ("the unneeded things were sold off at auction"). Any strip must supply an agent the source withholds, which is new content; the S111 DECOY ("were disposed of at a public auction") is equally agentless and is therefore not a strip at all. Excluded: the site was never Class A, and that is a finding about the S111 census, entered in the result.

5. The six minimal pairs, with the proposition inventory that licenses each

For every site: the propositions of the Polish, enumerated from the source; the marking device the target lacks a counterpart for; and the two spans. STRIP must state every proposition and carry no device from the removed list. Nothing else may differ.

Five of the six STRIP spans are the lead's S111 DECOY spans, committed to E-20260805c/materials/sites.json before stage A ran and therefore before the relation statements existed. Three are used verbatim; three carry a declared amendment. Every amendment is listed with its reason, and no amendment removes a marking beyond the declared device — two of the three make the pair more minimal and one restores a dropped proposition.

site device the target lacks STRIP provenance
S2 diminutive -ek on the ethnic noun Żyd S111 DECOY, amended: it dropped a proposition
S3 pejorative-pitying -czyna on kamizelka, with chora S111 DECOY, amended for minimality
S4 diminutive -ka on rubel S111 DECOY, amended for minimality
S7 second-person singular domyślasz się to the reader S111 DECOY, verbatim
S9 diminutive -ek on urzędnik S111 DECOY, verbatim
S10 honorific plural państwo with genericising takich S111 DECOY, verbatim

S2 — — Pewnie wielmożny pan chce parasol?... — odparł Żydek.

Propositions: (1) the speaker asks whether the gentleman wants an umbrella; (2) he addresses him as wielmożny pan; (3) the question is offered as a surmise (pewnie); (4) the speaker is the Jewish old-clothes dealer; (5) this is a reply to what came before. Removed in STRIP: the diminutive periphrasis little.

Amendment, declared: the S111 DECOY read "replied the dealer", which drops proposition (4). The relation statement is about the label, so a span that removes the noun tests nothing. Restored to the Jew. This is the only span authored after the relation statements were written, and it is the only one to which that objection applies.


S3 — Szaf na zbiory jeszcze nie mam, a nie chciałbym znowu trzymać chorej kamizelczyny między własnemi rzeczami.

Propositions: (1) he does not yet have cabinets for his collection; (2) he would not want, again, to keep the waistcoat among his own things; (3) the waistcoat is in bad condition (chora). Removed in STRIP: the diminutive periphrasis scrap of a and the personifying adjective ailing. At this site the marking is carried jointly by the suffix and the adjective, so the strip removes both; this is the site where the content/marking boundary is thinnest and it is named in the limits.

Amendment, declared: the S111 DECOY ended "among my own possessions"; restored to things so that the pair differs by the marking alone.


S4 — — Da wielmożny pan... rubelka! — odparł, roztaczając mi przed oczyma towar w taki sposób, ażeby okazać wszystkie jego zalety.

Propositions: (1) he names one rouble as the price; (2) he addresses the buyer as wielmożny pan; (3) he spread the goods out before the narrator's eyes; (4) so as to show all their merits. Removed in STRIP: the endearing diminutive on the price.

Amendment, declared: the S111 DECOY read "will give one rouble for it!" without the ellipsis; the ellipsis and for it are restored/removed so the pair differs by a little → one alone.


S7 — Patrząc na to, odrazu domyślasz się, że właściciel odzienia zapewne codzień chudnął i wreszcie dosięgnął tego stopnia, na którym kamizelka przestaje być niezbędną

Propositions: (1) looking at it, the inference comes at once; (2) the owner must have grown thinner every day; (3) and reached at last the point at which a waistcoat stops being necessary. Removed in STRIP: the second-person address (you guess → one guesses). English has no T/V distinction; the marked/unmarked contrast available to it is you against one.


S9 — Był to drobny urzędniczek, który na naczelników wydziałowych patrzył z takim podziwem, jak podróżnik na Tatry.

Propositions: (1) he was a clerk; (2) of low rank (drobny); (3) he looked at his department heads with wonder; (4) as a traveller looks at the Tatras. Removed in STRIP: the diminutive periphrasis minor little. The simile in (4) is source content and stands in both spans, so STRIP removes only the component the target lacks a counterpart for — see §10, limit 3.


S10 — bo służąca przeniosła się do takich państwa, którzy płacili jej trzy ruble na rok i codzień gotowali obiady.

Propositions: (1) the servant moved away to other employers; (2) who paid her three roubles a year; (3) and cooked dinners every day. Removed in STRIP: the genericising phrase the sort of people.

6. Arms, five, graded in one dispatch

arm what it is authored
FROZEN the filed T-kamizelka-R04-v1 rendering S111, before this design existed
STRIP the unmarked member of the pair five spans S111 (pre-statement), three amended above
REPEAT byte-identical to FROZEN, under a different letter —
POSITIVE an explicit statement of the relation S111, verbatim
WRONG an explicit statement of a different relation S111, verbatim

REPEAT is the control this instrument has never had. Two byte-identical spans appear under two letters in every body. A seat that grades them differently is measuring nothing; the disagreement rate is the procedure's own precision, and any FROZEN − STRIP difference has to be read against it. It is declared here so that the duplication cannot later be mistaken for a defect.

Grading procedure, unchanged from E-20260802e as repaired at E-20260804g. Three seats — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 — the same three as S111. Each sees the five spans per site unlabelled as A–E, in an order fixed by sha256(site | seat | ordering) and never chosen by the lead; ordering 2 is the reverse of ordering 1. Six bodies, 30 items each, 180 judgements. A seat counts as recovering a site only if it says YES in both its orderings; a site counts as recovered when ≥ 2 of 3 seats do. max_tokens 20,000 on every grading seat from the start — S111 lost $0.174 to two P2 bodies at a 12,000 ceiling and the A9 fix was a bigger ceiling for that seat; note (bhf).

7. The screens, run before grading, both gates

G-content. Two seats that are not graders — mistralai/mistral-medium-3-5 and qwen/qwen3.7-max, reserve z-ai/glm-5.2 — see the six pairs unlabelled and are asked one question and one only: do these two passages state the same facts — the same people, actions, objects, quantities, and the same words spoken? Ignore every difference of tone, attitude, warmth, politeness, sympathy, emphasis or style. They are shown no relation statement, no source, and no description of the study. A site is admitted only if both seats say the facts are the same.

G-prose. The same seats, in the same call, answer for each of the two passages: is this competent English prose, yes or no? A STRIP span called incompetent by either seat is removed — a rendering that fails as English would fail the grading for a reason that is not the marking.

Why these screens and not a marking screen. No seat can be asked which of these marks the relation without being told the relation, which is the answer. The marking side is therefore established by construction and declared per site in §5, mechanically checkable against the two spans, and the pre-run critic is asked to attack it.

The metalinguistic screen, mechanical. No STRIP or FROZEN span may state the relation rather than enact it — that would make it a POSITIVE. Checked by analysis/screen.py: the content words a span shares with its own relation statement, minus those already present in the source's denotation as glossed in §5, must be empty. Any hit removes the site.

8. Predictions, registered

Let r(arm) be the per-judgement YES rate over the admitted sites (6 judgements per site per arm at full return), and n the number of admitted sites.

# prediction scored how
PM (primary) the signature of a procedure that responds to marking: FROZEN − STRIP ≥ 0.25 with exact paired-permutation P ≤ 0.05 site-level means, labels permuted within site, all 2ⁿ arrangements enumerated
PM′ (primary, the other side) the signature of a procedure that responds to content only: FROZEN − STRIP ≤ 0.10 and STRIP ≥ 0.70 as above
PS STRIP recovers at > half the admitted sites site counts, ≥ 2 of 3 seats
PC1 POSITIVE ≥ n−1 sites and r ≥ 0.90 — replicating S111's 1.000
PC2 WRONG ≤ 1 site and r ≤ 0.20 — replicating S111's 0.083
PN REPEAT and FROZEN disagree in ≤ 2 of 36 paired judgements byte-identical spans
PF FROZEN recovers at all six sites, as at S111

Neither PM nor PM′ can hold with the other. Between them lies a declared indeterminate band, and if the run lands there the page says so and claims nothing.

The lead's registered expectation, written before dispatch and scored afterwards. STRIP ≥ 0.70, FROZEN − STRIP ≤ 0.10 — i.e. PM′, the procedure is content-only. S111's lead expectation was wrong in a specific way and saying so was worth more than the prediction; the same discipline here.

9. Failure criteria

10. Limits, written before the numbers exist

  1. Nothing about quality. Tier D is NOT PASSED; evidence class X3 is inadmissible. This run measures availability — whether a relation is recoverable — and judges no translation. Every sentence of the result is provisional.
  2. Nothing about human readers. Three language models are not a survey (charter §4).
  3. Six sites, one story, one translator, one pair. A census, not a sample. No sampling inference is available and none will be drawn.
  4. The relation statements are composite. At S9 and S10 part of the stated relation is carried by source content that stands in both spans (the Tatras simile; the employers' three roubles). STRIP removes only the component the target lacks a counterpart for, which is what R1 is about — but it means a high STRIP rate at those sites is partly explained by content the design did not remove.
  5. S3 is the thin case. There the marking is carried jointly by a suffix and an adjective present in the source, so the strip is closer to a content change than at the other five sites.
  6. Contamination is declared and NOT measured, exactly as at S111: three English «Kamizelka» renderings exist in print (Jopson 1930–31, Scherer 1955, Sues 1960) and none is freely reachable, so tools/dependence_check.py has no comparator. The design's validity does not turn on independence: both members of every pair are lead-written or lead-frozen and the question is what the graders respond to, not whether the English is original. The declaration is a declaration without a number and is repeated here rather than quietly dropped.
  7. The STRIP spans are the lead's own. Five were committed before the relation statements existed, which removes the worst version of the objection but not all of it: the lead chose which devices count as marking, and §5 states those choices per site so they can be disputed.

11. Pre-flight cost estimate — worst case from max_tokens, per note (abc)

stage seats max_tokens worst case
pre-run critic, up to 2 passes z-ai/glm-5.2, reserve moonshotai/kimi-k3 24,000 $0.13
G-content + G-prose mistral-medium-3-5, qwen3.7-max 8,000 $0.06
grading, 3 seats × 2 orderings P1, P2, P3 20,000 $0.80
total $0.99

Declared worst case: $1.15, allowing for the routing caution in config/models.md (a slug may be served by a provider charging several times list). UTC day 2026-08-05 stands at $3.166196086 of $5.00 with $1.833804 of headroom, so the worst case fits with margin. Lead translation and all analysis are free and are never ledgered (charter §3, A4).

12. Verification

analysis/verify.py imports nothing from analysis/score.py and recomputes every reported number from the stored raw bodies: the parse, the per-seat both-orderings rule, the site counts, the per-judgement rates, the permutation P by exhaustive enumeration, the REPEAT/FROZEN disagreement count, and the billed cost re-summed per request and cross-checked against the key-usage delta. Mutation tests: at least three deliberate corruptions of the stored data, each of which must be caught.

Amendments

(numbered, dated, with the finding that forced each; appended only, never rewritten)

Pre-run critic, 2026-08-05. deepseek/deepseek-v4-pro (P5 — not a grading seat, not a screen seat, and it produces no data in this run), max_tokens 24,000, effort low, 636 s, $0.02631576. NEEDS-AMENDMENT, three findings, one BLOCKING, all three accepted.

The first seat produced nothing and was not billed. z-ai/glm-5.2 at effort high returned no body in 600 seconds and was killed by the wrapper; the key-usage delta across it is exactly 0.000000000. Note (bhf) fires for the twelfth recorded time and in a new mode — not a length body with hidden reasoning, but no response at all — and rule (iii) applied as written: the seat was changed, not the ceiling.

A1 — the content screen's wording (finding 1, BLOCKING, accepted in full)

The G-content question in §7 asked whether the two passages state "the same words spoken". At S4 and S2 the marking sits inside quoted dialogue (a little rouble / one rouble; the little Jew / the Jew), so a seat reading the question literally would call the pair fact-different for a difference of wording that is exactly the variable under test. That could have removed the dialogue sites, dropped the census below F3's floor of five, and withheld the primary for a reason belonging to the screen and not to the materials.

The question is replaced, before any screen dispatch, by the critic's own wording:

Do these two passages state the same facts — the same people, actions, objects, and quantities? For any speech, judge only whether the same propositional content is conveyed, ignoring exact wording.

The instruction to ignore tone, attitude, warmth, politeness, sympathy, emphasis and style stands unchanged. run_screens.py carries the new text and no other change.

A2 — S3 gets a pre-registered sensitivity analysis, not a repaired span (finding 2, SERIOUS, accepted; the prescribed alternative remedy declined in writing)

The critic is right that at S3 the source's chorej kamizelczyny carries a physical component — smallness, a scrap — that STRIP's old and damaged waistcoat does not state, so a difference measured there could be content rather than marking. Of the two remedies offered, the first is declined: adding small to STRIP would put back the very diminutive force the strip exists to remove, which is the defect the repair is meant to avoid.

The second is adopted and registered here, before any datum exists. analysis/score.py and analysis/verify.py report the primary twice — over all admitted sites, and over all admitted sites except S3 — each with its own exact permutation P. If the two disagree about which signature holds (PM, PM′ or the indeterminate band), the result page reports the S3-excluded figure as the primary and says so, because that is the conservative reading. §10 limit 5 stands and is now backed by a number rather than by a caution.

A3 — the two aggregation rules named apart (finding 3, MINOR, accepted)

PS uses the strict rule — a seat recovers a site only when it says YES in both its orderings, and a site is recovered when ≥ 2 of 3 seats do. PM, PM′ and the permutation test use raw per-ordering judgements, six per site per arm. The two answer different questions and neither is derived from the other. analysis/verify.py recomputes each separately and asserts both against scores.json.

A4 — the gate fired at 1 of 6, and what the grading stage may and may not conclude (written BEFORE the grading dispatch)

F3 FIRES. One site of six survives G-content: S7. Both seats admitted S7; both called S2, S4, S9 and S10 fact-DIFFERENT; they split at S3 (G1 DIFFERENT, G2 SAME). G-prose passed every span, 12 of 12. The primary — PM, PM′ and the permutation test — is WITHHELD, and nothing the grading stage returns can restore it. That sentence is written here, before the grading dispatch, so that no number arriving afterwards can be read as reviving a prediction its own gate has already killed.

The grading stage is dispatched anyway, and the reason is stated in advance. It carries two controls that do not depend on the gate — REPEAT, the first measurement this procedure has ever had of its own within-body precision, and the POSITIVE/WRONG replication against S111's 1.000 and 0.083 — and it answers one question the screens raise but cannot settle: whether the content difference the screens name is a difference that changes what a reader takes from the passage. Everything it returns about FROZEN against STRIP is descriptive only, on all six sites, and is reported as such.

Three descriptive predictions, registered now:

# prediction
PD1 STRIP recovers at a per-judgement rate ≥ 0.70
PD2 FROZEN − STRIP ≤ 0.10 over the six sites
PD3 at S7, the one gate-admitted site, FROZEN and STRIP recover at the same seat count

A note on the two screen seats. mistralai/mistral-medium-3-5 returned HTTP 400 twice — "top_p must be 1 when using greedy sampling", the provider-side defect S115 met and named as note (bjh); the runner's S115 repair stored the error body, so it was diagnosed from the first failure instead of by hand. The declared reserve z-ai/glm-5.2 took the seat under note (bfc). It is the same slug that produced no critic body, and it saw nothing of this design: no body ever returned from that dispatch, so it comes to the screen with no more knowledge than any other seat.