Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260801b-census-author/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260801b-census-author
statusfrozen
created2026-08-01
updated2026-08-01
sensesaccuracy, naturalness, style-correspondence, voice, cultural-mediation, consistency, purpose-fit
provisionaltrue
linkswiki/arms/ARM-option-census.md, workshop/experiments/E-20260801b-census-author/mode-rule.md, workshop/experiments/E-20260801b-census-author/register.md, workshop/experiments/E-20260731e-option-census/design.md, wiki/findings/results/RS-20260731e-option-census.md, wiki/findings/results/RS-20260730i-candidate-reach.md, framework/closure.md, framework/traceability-inventory.md, workshop/translations/kowalski/R04-v1/passageA.md, workshop/translations/kowalski/R04-v1/passageB.md, config/models.md, config/budget.md

E-20260801b — does k survive a change of census author, and is a change of author the same thing as a change of elicitation mode?

ARM-option-census step 2. Frozen before any rater was dispatched. The arm page's step-2 text is the design's starting point and is quoted rather than paraphrased:

"RS-20260730i established k (renderings ruled out) as the only option-count-invariant statistic and this project now recommends quoting it. Step 1 showed the option contents do not reproduce between subjects (three-seat Jaccard 0.170) and never coded a single candidate against a non-lead census. So k has been shown invariant to how long the list is and never to whose list it is."

1. The question, and the one the arm page did not ask

Primary. framework/closure.md §1.4 tells the project to report k — the absolute number of live renderings a bearing candidate rules out — at 0.14, on the ground that it did not move when the option list was halved (RS-20260730i: 0.1395 at four options, 0.1538 at two). Every census that figure was ever computed over was written by the lead.

Does k move when the census author changes?

Secondary, and it is why this session has a translation limb. The four author arms are not author-pure. The lead's census is a moment-of-decision log; the three seats' censuses are on-demand reconstructions by subjects who did not translate. RS-20260731e §7 named that difference and did not measure it. So a LEAD-vs-seats gap in k is author confounded with elicitation mode, and the size of the mode effect is unknown.

Two things follow, and the design uses both:

2. Materials

The author arms. The same 24 Reymont «Śmierć» loci (S01–S24) frozen at S073, under four censuses:

arm census mode provenance
LEAD the lead's own MOMENT E-20260731e/census.json, frozen S073
P1 openai/gpt-5.6-terra ONDEMAND E-20260731e/runs/r1_NONE_P1.txt
P2 google/gemini-3.6-flash ONDEMAND E-20260731e/runs/r1_NONE_P2.txt
P3 x-ai/grok-4.5 ONDEMAND E-20260731e/runs/r1_NONE_P3.txt

The seats' UNPRIMED (NONE) round is used and the FRAMEWORK round is not. A census written after reading the fourteen candidates cannot then be coded against them without circularity. Stated here so the choice is not read as a selection among available arms.

The mode arms. Adam Szymański, «Stolarz Kowalski» (1911), Polish Wikisource, opening paragraph, translated this session under R04 and split at a sentence boundary:

arm passage Polish → English census sites
ONDEMAND A, opening → "…grają na słońcu cudownymi blaski." 393 → 520 reconstructed after the translation froze O01–O20
MOMENT B, "Ale dłużej trochę pobyć tu potrzeba…" → end of ¶1 366 → 525 written at the moment of decision M01–M20

Contamination, measured before any locus was selected and before B existed (tools/dependence_check.py, unmodified, against Benecke & Busch 1921, Project Gutenberg 8378, a comparator located and written to file programmatically and never read):

unit longest shared run 7-grams 12-grams 15-grams verdict
passage A 6 0 0 0 clean
passage B 9 6 0 0 clean
whole span 9 6 0 0 clean

The A/B pair is itself a free datum for a question NEXT.md has been carrying: two adjacent ~400-word passages of one work, one translator, one comparator, one session — 6 tokens and 9. That is a within-condition spread of the longest-run statistic measured for the first time.

The register declaration is register.md, frozen in commit 1bc4b41 before the first English word of either passage. It is a genuine advance declaration for the Szymański passages and a declaration supplied by this design for the Reymont sites, which had none. §7 carries the cost.

3. The coding instrument, and it is not new

The rater prompt is E-20260730i's condition-C prompt, reused with three edits, all recorded here:

  1. the standing register fact reads British rather than American (register.md);
  2. each item tags its source language, including the six controls, so that "no language tag" is not a marker identifying the control block;
  3. each site prints its source sentence as CONTEXT, truncated at 240 characters.

Nothing else changes: the fourteen candidate rows are read verbatim out of the frozen S068 prompt file by build_items.py, the task wording is unchanged, and the output format is unchanged.

Derived quantities, unchanged from E-20260730i §5, with n = options and k = options excluded, over bearing (site, candidate) pairs by three-rater majority:

code condition
INVOKED 0 < k < n
DECIDES k = n − 1
VACUOUS k = n
k the primary statistic: mean options excluded per bearing (site, candidate) pair

E = k/n is not computed. RS-20260730i §1.1 measured it to be option-count-dependent and framework/closure.md forbids it.

3.1 The control block — data-qualified, not endorsed

Note (bfy): "a control item is qualified by DATA, not by a second opinion — either by a prior run in which it behaved as declared, or by building more control items than the criterion needs and reporting which ones behaved." Two consecutive sessions had a critic's blanket sight-unseen endorsement as their weak point.

So this design builds no new control items. All six are reproduced verbatim from E-20260730i's condition-C payload, where P1 passed 6 of 6, P3 passed 6 of 6 and P5 passed 3 of 6 — recorded in that run's analysis/results.json and re-read by build_items.py from the frozen prompt file rather than retyped.

id 730i item kind declared expectation prior behaviour
X1 I06 CTRL-POS C1 bears, excludes exactly 3 of 4 P1 ✓ P3 ✓ P5 ✗
X2 I19 CTRL-POS C1 bears, excludes exactly 3 of 4 P1 ✓ P3 ✓ P5 ✗
X3 I30 CTRL-VAC C1 bears, excludes all 4 P1 ✓ P3 ✓ P5 ✓
X4 I62 CTRL-NEG no candidate bears P1 ✓ P3 ✓ P5 ✗
X5 I43 CTRL-BEAR C3 bears P1 ✓ P3 ✓ P5 ✓
X6 I54 CTRL-BEAR C1 bears P1 ✓ P3 ✓ P5 ✓

Only X6 reads the standing register fact; the other five carry their own note or are register-independent. Under British dialect X6's four options still span four registers, so C1 bears on it exactly as before.

Every control appears in all three payloads, which buys a within-rater consistency check on the control block itself at no extra dispatch.

Note (bgo) is applied literally: the pre-run critic is pointed at the control block first, in its prompt, because "quoting a defect back at yourself does not prevent it — pointing the critic at the control block first does."

3.2 Payloads — a rotation, one arm per site per payload

⚠ AMENDED before any rater call, on the pre-run critic's BLOCKING finding 2 (critic.md, amendment A2). The structure below replaces a three-payload version in which every Reymont site appeared four times inside one payload, once per arm. The critic's objection is accepted in full: ratings 2–4 would have anchored on rating 1, and a rater could have identified the LEAD arm from list length alone. The superseded version is in commit d57ca16.

payload contents items
R1 all 24 Reymont sites, arm = AUTHOR[(i + 0) mod 4] + 6 controls 30
R2 all 24 Reymont sites, arm = AUTHOR[(i + 1) mod 4] + 6 controls 30
R3 all 24 Reymont sites, arm = AUTHOR[(i + 2) mod 4] + 6 controls 30
R4 all 24 Reymont sites, arm = AUTHOR[(i + 3) mod 4] + 6 controls 30
B Szymański O01–O20 and M01–M20 + 6 controls 46

AUTHOR = (LEAD, P1, P2, P3) and i is the site's index in source order. The rotation gives:

All four properties are asserted by build_items.py at build time and re-asserted independently by analysis/verify.py. Item order inside each payload is shuffled with the fixed seed 20260801 declared here and nowhere else; the keymap is written to disk before dispatch.

What the rotation does not fix. In payload B the twenty ONDEMAND items all come from passage A and the twenty MOMENT items all from passage B, so a rater attending to subject matter could separate them. Declared, not repaired.

3.3 Raters and the repeat

Three seats, P1 / P2 / P3 of config/models.md, one call per payload. Reserve declared before dispatch, per note (bfc): deepseek/deepseek-v4-pro.

The drift gate is taken by choice (workshop/experiments/README.md, option 1, and note (bfz)): payload A1 is re-issued byte-identically to one seat in the same session, and the δ it measures is this run's noise floor. Completion length per seat is reported beside it because it is free and (bfz) asks for it.

Any arm difference smaller than the measured floor is reported as inside the floor and not as an estimate. S073 published a null instead of a two-of-three-seat effect on exactly this ground.

4. Registered predictions

Registered before any rater call. The primary threshold is RS-20260730i's own: that run scored its length manipulation on Δ mean k ≤ 0.25 and reported −0.0605. Scoring the author manipulation on the same bar is the only way the two are comparable.

⚠ AMENDED before any rater call, on the pre-run critic's BLOCKING finding 1 (amendment A1). The four-arm range confounds author with elicitation mode, because LEAD is a moment-of-decision census and the three seats are on-demand reconstructions. It is renamed author-or-mode and may never be reported as "the author effect". The clean author contrast — three different authors, elicitation mode held constant — is the three-seat range, and it is promoted to its own prediction P1b.

# prediction what falsifies it
P1 author-or-mode. Range of k across the four author arms at the same 24 loci is ≤ 0.25 a range above 0.25 — k is not invariant to who wrote the census and how, and framework/closure.md's recommendation to quote it must be qualified. A large range is a FALSIFICATION under every reading; there is no wording of P1 on which it is a pass
P1b the clean author contrast. The three-seat range of k — three different authors, all ONDEMAND — is ≤ 0.25 a range above 0.25 — author alone moves k, with mode held constant, and this is the finding that needs no confound argument at all
P2 the decomposition. The three-seat range is smaller than the four-arm range equal or larger. UNRESOLVED, not falsified, if the four-arm range is at or inside the repeat's measured δ (amendment A5): comparing two quantities that are both inside the noise floor is not a result in either direction
P3 mode. |k(MOMENT) − k(ONDEMAND)| on the Szymański loci is ≤ 0.25, the same bar. It may not be subtracted from P1 (amendment A3): the mode arms are on different text, are confounded with passage, and bound an order of magnitude, nothing more above 0.25 — elicitation mode moves k by more than option-list length did
P4 (descriptive) mean n is highest in LEAD any arm above it
P5 frac2 (sites with ≤ 2 live renderings) is 0.000 in every arm, as it was in nine of nine cells at S073 and in the lead's own census any arm above 0 — and DECIDES becomes arithmetically reachable there
P6 closure-defeating, registered as such. If a candidate outside {C1, C2, C3, C5} is INVOKED at ≥ 3 sites by majority in any arm, framework/closure.md §1.4's reach sentence must be restated, not annotated, and this arm may not close without doing so — (a trigger, not a scored prediction)

5. Failure criteria

# fires when consequence
F1 fewer than 2 of 3 raters pass all six controls the run is descriptive only; no k comparison is reported as an estimate
F2 the byte-identical repeat's δ on k exceeds the four-arm range of P1 P1 is unresolved; the instrument cannot see an effect of the size present
F3 any arm returns fewer than 8 bearing (site, candidate) pairs by majority k for that arm is not reported as an estimate and the arm is described only
F4 a rater returns fewer lines than its payload has items, after one retry and one fall-through to the declared reserve that seat's payload is excluded whole and every figure is recomputed on the remaining seats, declared post hoc
F5 the control block's within-rater consistency across the three payloads is below 5 of 6 for a rater that F1 passed the control pass is reported as unstable and F1's verdict is repeated with the instability named

F1 is written to cover the case the null wins. NEXT.md records that S077's registered F4 "as registered does not cover the null winning" and that the stronger reading had to be applied after the fact (note (bgn)). Here: if the range in P1 exceeds 0.25 in the direction that qualifies framework/closure.md, that is a FALSIFICATION and is reported as one — there is no reading of P1 under which a large range is a pass.

6. Procedure

  1. Instrument mode-rule.md frozen (a448ed1) before the span was chosen. ✔
  2. Register declaration frozen (1bc4b41) before the first English word. ✔
  3. Passage A R06 draft (af9c8dd) → R04 (b29062e) → contamination gate → ONDEMAND census (41a01ef) → passage B R06 (81b1af0) → R04 + MOMENT census (ef64a59). ✔
  4. Items built by build_items.py from the frozen censuses and the frozen S068 prompt file; keymap written to disk. ✔
  5. This design frozen.
  6. Independent pre-run critic, pointed at the control block first (note (bgo)); findings recorded in critic.md with accept/decline for each; amendments committed before any rater call.
  7. Three raters × three payloads, raw bodies preserved before parsing (note (bdt)); one byte-identical repeat.
  8. analyse.py, then analysis/verify.py, which imports nothing from analyse.py, re-parses every answer from the stored .raw bytes, and carries mutation tests.

7. What this cannot establish, written before the run

8. Pre-flight cost estimate

Worst case is built from the max_tokens cap the request permits, not from an expected length — note (abc).

stage calls max_tokens worst case
pre-run critic (qwen/qwen3.7-max) — SPENT, $0.071769075 1 12,000 $0.20
payloads R1–R4 × 3 seats 12 6,000 $1.10
payload B × 3 seats 3 6,000 $0.28
byte-identical repeat of R1, one seat 1 6,000 $0.10
declared worst case 17 $1.68

P3 is priced at the worst plausible provider per the S022 routing caution. Today's ledger (UTC 2026-08-01) opened at $0.580589442 of $5.00 with $4.419410558 headroom, so the worst case is 38% of headroom. The worst case was raised from $1.14 to $1.68 by amendment A2, which doubled the author-arm dispatches from six to twelve — the cost of the critic's BLOCKING finding, recorded as such rather than absorbed silently. A stage that will not fit is dropped whole, in the order B → repeat → R4, and the deferral is written into NEXT.md.