Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260801c-anchor-instance/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260801c-anchor-instance
statusfrozen
created2026-08-01
updated2026-08-01
sensesnaturalness, voice, style-correspondence
provisionaltrue
internal-judgment-onlytrue
linkswiki/arms/ARM-anchor-instance.md, wiki/base/anchors/A-mchugh-presence/A-mchugh-presence.md, workshop/experiments/E-20260731f-catalogue-reach/design.md, workshop/experiments/E-20260731f-catalogue-reach/claims.json, workshop/translations/pelsen/R10c-v1/translation.md, workshop/translations/pelsen/opportunity.md, config/models.md, config/budget.md

E-20260801c — does the anchor's second instance attest the same catalogue?

Frozen before any stage-A or stage-F call. The gate (gate/) is a separate registration and ran first; it is not part of this design.

1. The question

A-mchugh-presence is the project's only Tier 1 anchor built on two texts, and it states its own reason: "one text cannot distinguish a register with no markers from an author with no mannerisms." The catalogue was then written as if from one text — E-20260731f/claims.json labels every one of the nineteen propositions C01–C19 with passage: presence-excerpt.txt — and snake-girl-excerpt.txt has never been read by any instrument, deterministic or panel.

So the anchor's stated design has never been executed. Two questions:

  1. Which of C01–C19 do both stored instances attest, and which rest on one?
  2. Whether the shared runs between T-pelsen-R10c-v1 and its published comparator are forced by the Swedish — the mandatory contamination gate on this session's translation limb returned the largest overlap this project has recorded for a lead translation, and a run length alone does not say what produced it.

The wire between the limbs, in one sentence. The translation was written to this anchor's catalogue as a target register, so the propositions that actually decide its wordings are on record; the study limb then measures which of those propositions the anchor's two instances both attest, and therefore whether the translation was aimed at a register or at one author's mannerisms.

2. Materials

Claim set — C01–C19, reused verbatim from E-20260731f/claims.json (sha256 1e9430e4fbfab8d1129e286226a6321341d6c621288c2ecdcf7384a6af692534), read programmatically, never retyped. Not re-decomposed: re-writing the propositions would make the two runs incommensurable, and the point is to run the same instrument on the text it was never run on.

Passages.

tag text words role
P McHugh, "Presence" (2005), opening — presence-excerpt.txt 1,511 the anchor's primary instance
S Kessel, "The Snake Girl" (2008), opening — snake-girl-excerpt.txt 1,505 the anchor's second instance, never read by any instrument
Y Doctorow, Little Brother (2008), opening 1,760 the fixed contrast, identical in both arms

Y is held fixed across the two arms deliberately. E-20260731f assigned contrasts A↔C, B↔A, C↔B; its arm for P used Y. Using Y for S as well keeps the only difference between the arms the target passage, and keeps arm P comparable to the published run.

Translation limb. T-pelsen-R10c-v1 — Söderberg «Pelsen», 1,273 Swedish words → 1,298 English, R10 against this anchor, frozen at 217ce02 with its opportunity list frozen before it at 9fb7ba2. Contamination high, measured after the freeze: longest run 24 tokens, 46 shared 12-grams, 404 of 1,299 tokens (31.1%) inside a shared run ≥ 7 with Stork 1923.

3. Seats

Per config/models.md. All non-Anthropic. Used only on the factual side of the S015 line — is this stated property of English present in this passage — never is this prose good. Tier D has not passed; every verdict here is internal-judgment-only and provisional.

4. Procedure

Stage 0 — deterministic, free, and RUN BEFORE THIS DESIGN WAS FROZEN

Two things, both by importing E-20260731f/analysis/checks.py unmodified and running it against snake-girl-excerpt.txt, which that module already knows how to load and was never asked about:

Stage 0's results were obtained before this design was written and are declared here rather than predicted, because a design cannot register a prediction about a computation it has already run. They are used in two ways only: to build the control block, and as the thing the panel arms are compared against.

Stage A — the per-text attestation (6 calls)

One call per (passage ∈ {P, S}, rater ∈ {P1, P2, P3}). Each call carries the target passage X, the fixed contrast Y, and 22 items: C01–C19 in their frozen wording plus three planted claims, interleaved and unlabelled.

The control block — six cells per arm, every ground truth from stage 0's DATA, per note (bfy):

id claim expected on P expected on S mechanical basis
K1 C14 — forty-eight paragraphs across 1,511 words PRESENT ABSENT 48 / 1,511 against 36 / 1,505
K2 C13 — very short paragraphs, one-sentence paragraphs structural PRESENT PRESENT mean 31.4 / 41.6 words (≤ 45); one-sentence 8 / 5 (≥ 5)
K3 C04 — no dialect spelling PRESENT PRESENT declared dialect patterns: none in either
K4 planted: the narration is in the first person ABSENT ABSENT 0 first-person tokens in narration in either
K5 planted: at least one word is set in italic or bold type ABSENT ABSENT 0 emphasis spans in either stored file
K6 planted: at least one paragraph runs to more than three hundred words ABSENT ABSENT longest paragraph 122 / 169 words

K1 is the only cell whose correct answer DIFFERS between the arms, and it is therefore the only one that can catch a rater answering from a general impression of unmarked contemporary prose rather than from the passage in front of it. K1 is declared to be doing double duty — it is both a control and one of the nineteen findings — and that is stated here rather than discovered later.

Hairline cells excluded from the control block, and why. Stage 0 puts C08 on S at say-rate 0.812 against a PASS threshold of 0.80, and C17 on S at lag-1 autocorrelation 0.303 against a threshold of 0.30. Both are within 0.015 of their thresholds. They are the two most interesting cells in the run and they are not used as ground truth for anything.

Stage F — is the overlap forced? (3 calls)

The instrument is RS-20260728b-forced-run-ru's, unchanged. Three Swedish paragraphs — the ones carrying the four longest shared runs (source ¶4, ¶16, ¶34) — are given alone, with no English, no comparator, and no mention that any published translation exists, and each seat is asked for a contemporary English rendering.

Frozen verdict rule, inherited: a shared run is FORCED iff ≥ 2 of 3 independent seats reproduce ≥ 7 contiguous tokens of it; ELECTIVE otherwise.

Reported alongside, and it is the more informative number: each seat's own longest common run against Stork 1923 and against the lead, over the same three paragraphs.

5. Registered predictions

# prediction what falsifies it
P1 (primary) ≥ 4 of 19 propositions get different majority verdicts on P and S ≤ 3 differ — the pair attests one catalogue and the anchor's two-instance design is doing its job
P2 C08 (say-attribution, no ornamental verbs) is PRESENT on P and ABSENT on S any other pattern. Registered because stage 0 finds gasped and whispered six times in S and zero ornamental verbs in P, and because the hand read finds one of that item's four quotations is S's
P3 C14 differs (K1) it does not — in which case the raters are not reading the passages and F1 has probably already fired
P4 the unattested rate — PRESENT whose quotation is not verbatim in X — is ≤ 0.15 on both arms above 0.15 on either
P5 (closure-defeating, registered as such) — if the arms differ at ≥ 8 of 19, the anchor's sentence "one text cannot distinguish a register with no markers from an author with no mannerisms" must be restated, not annotated, and this arm may not close without doing so
P6 the 24-token run is FORCED fewer than 2 of 3 seats reproduce ≥ 7 tokens of it
P7 the 15-token final-line run ("it has given me the last seconds of happiness i have known in my life") is NOT FORCED 2 or 3 seats reproduce ≥ 7 tokens of it

P1 is the primary. P3 is a control restated as a prediction and buys nothing on its own.

6. Failure criteria — pre-committed

7. Order of operations

  1. Target propositions frozen — S074, sha256 1e9430e4….
  2. Source ingested; opportunity list frozen — 9fb7ba2.
  3. Translation and translator's log frozen — 217ce02.
  4. Comparator fetched and contamination measured — 7a3b39a. (3 before 4 is R10 §1.)
  5. Stage 0 run.
  6. This design frozen.
  7. Pre-run critic; amendments written and committed before any stage-A or stage-F call.
  8. Stage A, then stage F. Raw bodies preserved.
  9. analysis/verify.py, which imports nothing from analyse.py and re-parses every answer from the stored .raw bytes, with mutation tests.

8. What this design cannot do

9. Cost — pre-flight, note (abc), and built from the RETRY STRUCTURE as well as from max_tokens

The gate this session ran overran its own declared worst case by 3%, and the reason was not the token cap: it was the runner's fall-through, which can dispatch a seat up to 3 attempts × 2 slugs. An estimate built from one dispatch per seat is not a worst case. Attempts are therefore capped at 2 per slug for this design and the estimate is built accordingly.

stage shape max_tokens worst case
pre-run critic 1 call, qwen/qwen3.7-max 12,000 $0.16
stage A 2 arms × 3 seats, in ≈ 4,600 8,000 $0.38, ×2 attempts = $0.76
stage F 3 seats, in ≈ 900 4,000 $0.088, ×2 attempts = $0.18
one reserve dispatch — 8,000 $0.06

Declared worst case $1.16. Spent so far today across all sessions $2.008982423; headroom $2.991017577. The gate's $0.105042406 is already inside that figure.