Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260731e-option-census/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260731e-option-census
statusfrozen
created2026-07-31
updated2026-07-31
sensescultural-mediation, style-correspondence, voice, naturalness, consistency
linkswiki/arms/ARM-option-census.md, framework/closure.md, framework/traceability-inventory.md, wiki/findings/results/RS-20260730i-candidate-reach.md, wiki/findings/results/RS-20260730h-strict-coverage.md, workshop/experiments/E-20260731e-option-census/census-rule.md, workshop/experiments/E-20260731e-option-census/census.json, workshop/translations/smierc/R04-v1/translation.md, workshop/experiments/README.md, config/models.md

E-20260731e — whose option list is it?

Re-frozen 2026-07-31 after the independent pre-run critic pass (critic.md, qwen/qwen3.7-max, NEEDS-AMENDMENT, three findings — one BLOCKING, one MANDATORY, one ADVISORY, all three accepted; all eight control items ENDORSED sight-unseen). Nothing was dispatched to a rater before this amendment.

Frozen 2026-07-31 (S073) before any call was dispatched. The census instrument was frozen earlier still, at 8c12508, before the source span existed (census-rule.md).

1. Question

Every framework-reach figure this project publishes is computed over an option census written by the lead. framework/closure.md §1.4 reports DECIDES 0 of 126 and k = 0.14; §1.4's boxed correction (S068) reports 0.000 at four options against 0.154 at two; framework/traceability-inventory.md §3 reports which of the fourteen candidates a decision reaches and how the invocations concentrate. All of it is a function of a list of live renderings, and every such list in the repository was written by one translator who has read the fourteen candidates. ARM-candidate-reach said so in its own closure: "It never coded a non-lead option census, so §2's reach figures are still one translator."

RS-20260730h (S067) then showed that the project's other coverage statistic was rank-identical to the mean number of live options the translator wrote down, and RS-20260730i (S068) showed that DECIDES doubles-to-appearing when the option list is halved while k does not move.

So the denominator is doing work, and nobody has asked where the denominator comes from.

Does exposure to the framework change the option list?

If it does, the object every reach statistic is computed over is endogenous to the thing the statistic is measuring.

2. Why this is T5

The object under audit is the synthesis object's own denominator — not a jury, not a calibration, not a typology. It is the S068 shape: auditing what the synthesis object already claims, which wiki/tracks.md records as the one kind of T5 work that is neither downstream of Tier D nor a fourth container. It is not a release push and ARM-option-census's completion criterion forbids it becoming one. No candidate is added, removed or licensed by this design.

3. Materials

4. Arms

Three preambles. The preamble is the only thing that differs between arms; task text, items and order are byte-identical.

arm preamble
NONE none
SHAM 14 lines of generic translation advice with no warrant of any kind, matched to FRAMEWORK on line count and word count (prompts.py, and the realised counts are reported)
FRAMEWORK the fourteen candidate recommendations of framework/traceability-inventory.md §2, verbatim, row text only, as RS-20260728i and RS-20260730i supplied them

The sham is built and measured, not asserted (standing disposition 6). It is matched on shape (14 numbered imperative-or-declarative lines) and on length; its realised word count against FRAMEWORK's is reported in the analysis, and a mismatch above 15% is declared as a confound.

Nothing in either preamble mentions counting, enumerating, listing more, or option lists. The draft sham line "the first rendering that comes to mind is rarely the last one to consider" was removed before freezing because it encourages enumeration directly and would have confounded the sham with the hypothesis; it is replaced by "where the source is ambiguous, the rendering may be ambiguous too." Recorded because a removed item is part of the design.

5. Task, identical in all arms

For each item below you are shown a paragraph of Polish with one locus marked «…». List every distinct English rendering of the marked locus that you would consider live — one you would actually weigh before choosing. Do not choose one. Do not rank them. Give each as a short phrase.

Answer with one JSON object per line: {"id": "...", "options": ["...", "..."]} and nothing else.

6. Seats

Raters P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 (config/models.md). Reserve declared before dispatch (note (bfc)): deepseek/deepseek-v4-pro. Pre-run critic: qwen/qwen3.7-max, probed-but-not-selected, not a rater and not the raters' reserve — the S053 role-collision fix, seventeenth consecutive session. max_tokens 6,000 on raters (note (b), which has fired on empty length returns eighteen times), 8,000 on the critic.

Judgment is not parallelised and none is asked for. Enumerating renderings one would weigh is not a quality judgment about any of them; no seat is asked which is better, and no seat sees the lead's translation or its census.

7. Dispatch plan

round arm seats purpose
1 NONE P1, P2, P3 baseline
2 SHAM P1, P2, P3 preamble-presence control
3 FRAMEWORK P1, P2, P3 the manipulation
4 NONE, byte-identical repeat P1, P2, P3 the same-day noise floor

12 rater calls + 1 critic call. Round 4 is workshop/experiments/README.md's cross-day drift gate, option 1, taken by choice — the first design to take it rather than the exemption. Completion length per seat is reported for every call, as that gate also requires.

8. Statistics

9. The matcher, frozen verbatim (standing disposition 9)

import re
STOP = r"\b(the|a|an|to|of|my|his|her|your|it|is|are|i|me|we|they|he|she|was|were|be|been|that|and)\b"
def norm(s):
    s = s.lower().strip()
    s = s.replace("’", "'")
    s = re.sub(r"[^a-z0-9' ]+", " ", s)
    s = re.sub(STOP, " ", s)
    s = re.sub(r"\s+", " ", s).strip()
    return s

Two renderings match iff norm() is equal. Nothing fuzzier; no stemming, no synonymy.

Amended after the critic pass (ADVISORY finding 3): the stop list gained the personal pronouns and the copula/auxiliary forms, which it had omitted. The change loosens matching in both arms equally, so it cannot move the Pred5 contrast; it raises the absolute containment figures, which were always declared a lower bound.

10. Failure criteria — registered, and two of them can void the run

F1 — the positive control, VOIDING. Pooled over the three arms, for each seat, N_open − N_con ≥ 1.00. Required in at least 2 of 3 seats. If fewer than 2 seats separate the control classes, the seats are not tracking option availability at all and every primary number in §8 is VOID, reported as such and not reinterpreted.

F2 — the noise floor. δ(seat) = |N(seat, NONE, round 1) − N(seat, NONE, round 4)|. Any arm difference whose magnitude is below max_seat δ is reported as inside the noise floor and is not an estimate. This is RS-20260730-grain-clause's F1 imported as a standing condition rather than rediscovered.

F3 — coverage. A seat returning fewer than all 32 item ids in an arm is excluded from that arm, and the exclusion is reported. If fewer than 2 seats survive in any arm, the primary is VOID.

F4 — sham length. If the realised word counts of the SHAM and FRAMEWORK preambles differ by more than 15%, the sham is declared a confounded control and P2's result is reported with that caveat attached. It does not void anything.

11. Predictions, each with its failure condition

# prediction fails if
Pred1 N(FRAMEWORK) > N(NONE) in ≥2 of 3 seats, by more than the F2 floor the direction reverses in ≥2 seats, or the difference lies inside the floor
Pred2 \|N(SHAM) − N(NONE)\| < \|N(FRAMEWORK) − N(NONE)\| in ≥2 of 3 seats otherwise — which would say the effect is the presence of a preamble, not its content
Pred3 the lead's own mean n over the 24 primary items, 4.750, frozen at 6b9b447, exceeds N(NONE) averaged over seats it is lower. Descriptive: one subject, a different elicitation mode
Pred4 frac2 is 0.000 for the lead and > 0 for at least one seat in the NONE arm every arm also returns 0.000
Pred5 containment(FRAMEWORK) > containment(NONE) in ≥2 of 3 seats otherwise. Descriptive, and a lower bound by construction

Pred4 is the prediction that touches a published figure. RS-20260730i measured DECIDES at 0.000 with four live options and 0.154 with two. frac2 is therefore the lever on framework/closure.md §1.4's headline, and the lead's census supplies 0 of 24 sites at n ≤ 2 — which is the third independent time a lead census has produced no n = 2 sites (S068's pre-run critic found the second and it changed that design).

12. What this design does not claim

It does not claim any rendering is better than any other, and no seat is asked. It does not claim the lead's census is wrong. It does not test whether shared runs are forced. And it cannot establish that a non-lead census is the right denominator — three language models enumerating renderings on demand are not a translator writing a log at the moment of choosing, and the two elicitation modes differ in more ways than the manipulation. What it can establish is whether the denominator moves under a manipulation of exactly one variable, which is the only question §1 asks.

13. Pre-flight cost

Worst case $1.20, built from max_tokens (note (abc)) at list out-price × 1.5 provider premium (S022 routing caution), summed over 12 rater calls at 6,000 and 1 critic call at 8,000, with input priced at the realised payload size. Today's headroom before this run: $4.293746048.