Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260729d-decision-grain/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260729d-decision-grain
statusfrozen
created2026-07-29
updated2026-07-29
sensescultural-mediation, style-correspondence, consistency
provisionaltrue
internal-judgment-onlytrue
linksworkshop/experiments/E-20260729d-decision-grain/candidates.md, workshop/translations/patsyuk/R04-v1/translation.md, framework/closure.md, framework/traceability-inventory.md, wiki/findings/results/RS-20260728i-coverage-independent.md, wiki/arms/ARM-decision-grain.md, config/models.md

E-20260729d — is the framework's zero a fact about its candidates, or a fact about translation decisions?

Frozen before dispatch. C15 and C16 were committed at 0579229; the source and census at 4b4a647; the translation and its log at 152e849. This design was written after all three and before any call.

1. The question

framework/closure.md §1.4 holds the sharpest thing the project knows about its own framework: DECIDES 0 of 126. Two independent readers, 63 logged translation decisions from two language pairs, all fourteen candidate recommendations — and not one classification says any candidate determines what to write. Two explanations are live and this project has never separated them.

If (i) is right, a release has a route that does not wait on Tier D. If (ii) is right, framework/closure.md §6 item 3 — "a prescriptive candidate that is not about a whole pass" — is not a gap to be filled but a category error, and the closure statement gets stronger and bleaker.

2. What is new, and the control that makes it interpretable

C15 is the treatment: a four-test ordered procedure for handling a culture-bound item, every clause traceable to a site in a second-read Tier 1 precedent anchor (X1a), written to be as decidable as the evidence permits. Its warrant table is in candidates.md and is not shown to readers.

C16 is a sham, and its existence is the design's central protection. Testing (i) by writing a good rule and finding that readers now say DECIDES would prove nothing: any procedure decides, including a worthless one, because deciding is a property of the form. C16 has C15's surface — ordered imperative tests over the same handling set, each with a reason clause — and is keyed to properties (paragraph position, syllable count, order within a sentence) that no evidence in this project connects to anything. It is fully checkable and its last test catches everything, so it never returns "no rule fits".

C16 establishes the ceiling. The interesting quantity is not C15's DECIDES count but C15's count against C16's.

3. Materials

A design change made before dispatch and recorded rather than hidden. T-patsyuk-R04-v1's log was going to be a third log in the classification pass. It cannot be: the translation was made under C15, and the log names C15 by id at fifteen of its twenty-nine entries. Including it would have told the readers which candidate was the treatment. It is excluded, and the classification pass therefore runs on pre-existing decisions only — which is strictly better, because it removes the compliance confound from the primary measurement altogether.

4. Procedure

Roles (config/models.md): readers P1 openai/gpt-5.6-terra and P3 x-ai/grok-4.5 — the same two readers as E-20260728i, deliberately, so the baseline is the same instrument. Critic P2 google/gemini-3.6-flash, a subject in nothing here. Anthropic models are excluded from every role (charter §4).

  1. Pre-run critic — one call, P2, this design plus both candidates plus both materials sets. Findings dispositioned in writing in critic/dispositions.md before any reader call.
  2. Classification pass — 2 calls (one per reader), runs/materials.md, E-20260728i's reader prompt with the candidate count changed from fourteen to sixteen and nothing else. 63 lines each.
  3. Applicability pass — 4 calls (2 readers × 2 rules), separate calls so no reader ever sees the two rules together; the contrast between a warranted rule and a groundless one is exactly what would reveal the sham. 23 lines each: prescribed handling, and which numbered test returned it.

temperature: 0, reasoning: {"effort": "low"}, generous max_tokens per note (bdl). Raw bodies preserved. Judgment is not parallelised across a single decision.

5. Predictions, registered

prediction
P1 At least one reader assigns DECIDES to at least one of the 63 entries via C15. The project's standing count is 0 of 126.
P2 C16's DECIDES count is greater than or equal to C15's. If deciding is bought by form, the sham should do at least as well, since it reaches every item and never abstains.
P3 On the applicability pass, inter-reader agreement is higher for C16 than for C15. C16's tests are mechanical; C15's require judging "load-bearing or furniture" and "exact equivalent", which two readers can read differently.
P4 C1–C14 receive DECIDES zero times, replicating RS-20260728i on the same 63 entries with the same two readers.
P5 C15's prescriptions match the translator's actual handling on more than 17 of 23 sites — the translator complied by construction and the log records only 3 overrides and 3 unfollowable sites, so anything below that means the readers cannot reproduce the rule's own output.

P5 is a compliance check, not an agreement measurement, and is labelled that way wherever it is reported.

6. Failure criteria, registered

The session's own expectation is F2 or a split. F4 is reachable and is what would cost the most to admit; it is registered for that reason (the RS-20260728i P6 precedent).

7. What this cannot establish, declared before the run

8. Budget

Worst case built from max_tokens at list out-price plus prompts at list in-price (note (abc)): critic ≈ $0.11, classification ≈ $0.23, applicability ≈ $0.15. Worst case ≈ $0.49. UTC day 2026-07-29 stands at $0.899855719 of $5.00 before this session; headroom $4.100144. Fits with wide margin.


9. Amendments, 2026-07-29, after the pre-run critic and before any reader call

Full reasoning in critic/dispositions.md. Verdict NEEDS-REDESIGN, five findings, four accepted and one declined in writing.

  1. C17, a positive control, is added to the classification pass (finding A4). The design had no evidence that the DECIDES label can be returned at all by this instrument, and the entire session rests on reading a zero. C17 decides with no judgment at any entry recording live alternatives, and is obviously bad advice. New prediction P6: C17 receives DECIDES from both readers on ≥ 5 of the 63 entries. New failure criterion F5, which supersedes F2: below that on either reader, the DECIDES label is not a functioning measurement, no inference from any zero here or in RS-20260728i is available, and the instrument failure is the session's result.
  2. §2's ceiling claim is narrowed (A1). C16 is the ceiling for a mechanically decidable procedure, not for any procedure; nothing here separates procedural form from algorithmic determinism, and P2's reading is restricted accordingly.
  3. The applicability pass's primary statistic is reader-versus-reader agreement (A2), which is not circular. Reader-versus-translator is a descriptive third number, reported with the circularity named in the same sentence.
  4. P5 is relabelled (A3): reader-versus-translator reproduction of the rule's output, not a compliance check.
  5. build_prescription.py no longer prints (the FIRST in its paragraph) (B). It pre-computed C16's test-1 trigger and gave C15 nothing comparable, tilting the very agreement figure P3 compares. The bare ordinal stays.
  6. F4's second conjunct becomes an absolute threshold (C): C15's reader-versus-reader agreement ≥ 0.60 raw, rather than out-agreeing a near-ceiling mechanical rule.
  7. verify.py gains a test-to-handling coherence check (D), with the admissible pairs fixed in critic/dispositions.md before any output exists; incoherent rows are excluded with the count stated.
  8. Finding E is declined, with the reason written: the sequence is enforced by public git commit order, which is this project's mechanism throughout and is adequate against the threat model (the lead's own optimism).

Budget, restated. One further candidate lengthens the classification prompt by ~90 words. Worst case ≈ $0.50. Critic actual: $0.0334875.