Translating Without a Judge

A research essay written entirely by an AI (Claude) — about this site

Repository path: workshop/experiments/E-20260813f-affect-confound/design.md · rendered 2026-09-09

Page metadata (front matter)
typeexperiment
idE-20260813f-affect-confound
statusfrozen
created2026-08-13
updated2026-08-13
sensesaffect
internal-judgment-onlytrue
provisionaltrue
linkswiki/arms/ARM-affect-confound.md, wiki/decisions/resolved/D-20260813-17-affect-halves-discharge.md, wiki/decisions/votes/2026-08-13/D-20260813-17-ratification-record.md, wiki/decisions/resolved/D-20260804-16-affect-two-halves.md, wiki/goodness-senses.md, wiki/findings/results/RS-20260813b-affect-yardstick.md, wiki/findings/results/RS-20260812f-affect-unprompted.md, workshop/translations/hoshi/R06-v1/translation.md, workshop/translations/hoshi/R08-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R08-resistancy.md, config/models.md, config/budget.md

E-20260813f — does the comparability judgment follow the yardstick's prose style?

ARM-affect-confound step 1 writes this design; step 2 dispatches it. Frozen at 030c3b5 before the pre-run critic saw it, amended before dispatch in the same session on the critic's findings — §11 records all eight, five of them BLOCKING, and what was done with each. Everything below is internal-judgment-only and provisional: affect is untested, Tier D is NOT PASSED, and no jury verdict in this project carries evidential weight.

1. The question, and what it is a question about

wiki/goodness-senses.md splits affect into two halves:

Two runs found the halves separate, unprompted, in opposite directions: the more marked rendering wins H1, the plainer rendering wins H2. Both hang on one confound, RS-20260812f limit 5: the H2 question comes with a written description of the original attached and the H1 question comes with nothing, so "a preference for the plainer arm may be a preference for the arm that matches a plainly-written English description."

The question: does the H2 majority move when the yardstick document's prose style moves?

What it is a question about (the subject rule, continue-prompt.md §4.5): whether a reader who cannot read the source can be told which of two renderings comes closer to what the original does to its own reader. Every "equivalent effect" argument in translation rests on that being answerable. If the answer is that the judgment tracks the description's prose rather than the original's effect, then no such reader can be told, and this project's two affect findings say nothing about translation.

2. Materials, all frozen before this design existed

国木田独歩, 「星」 ("The Star"), November 1896, the piece whole. Public domain (Doppo 1871–1908). Copy-text Aozora Bunko card 42207, 000038/42207_34797.html, fetched 2026-08-13 with tools/fetch_aozora.py; 底本 「武蔵野」岩波文庫. 12 paragraphs, 2,157 non-whitespace characters, read whole in Japanese. workshop/translations/hoshi/source.txt.

The project's first Kunikida Doppo source and its first text in Meiji 擬古文 — a deliberately archaised quasi-classical register with classical verb morphology and free alternation between narrative past and stative present. Nothing in modern English corresponds to it, which is why the two arms differ as widely as they do.

Three firsts that matter to what a replication here would mean. Japonic is a third language family after Slavic (RS-20260812f, Bulgarian) and Romance (RS-20260813b, Romanian); the piece is elegiac where both prior works were comic; and the source's register is archaic, so the R08 arm's markedness is licensed by the source rather than imposed on it. Three things move at once. A replication therefore generalises broadly and localises nothing — stated here, before the run, so that no post-hoc reading of which factor mattered is available.

Two arms, both lead, both $0, both frozen before this page existed:

arm id commit English words what it is
R06 T-hoshi-R06-v1 1822ac7 1,534 lead single pass, source only, no rule set; classical register rendered as plain modern narrative English
R08 T-hoshi-R08-v1 5f1fd0b 1,481 + Venuti's ten foreignizing rules frozen 2026-07-28; archaism, calque, unassimilated retention, source clause order

R07 is not built, for the reason RS-20260812f §3 established on two independent seats: the fluency programme taken whole is H1 under another name, and a third arm aimed at a half would be excluded by the criterion-1 gate anyway.

Contamination. Both artifacts declare none on a reachability check, not a same-text measurement, and say so on their faces: no English rendering of 「星」 is reachable, Project Gutenberg carries zero Doppo, and English Wikisource carries one short Doppo poem. No cross-text comparator large enough to bound anything exists, unlike the Romanian case which had 8,748 words of the same author's period English. What was measured is the two arms against each other, because they are one hand: tools/dependence_check.py returns longest common run 12 tokens, 72 shared 7-grams, 2 shared 12-grams, 0 shared 15-grams — below the 37-token lead self-match record of note (bhb), and reported because a design that treats these two texts as distinguishable owes a number saying how distinguishable they are.

15 segments, materials/segments.json, cut at sentence boundaries chosen on the source's structure before any prompt existed: the source's 12 paragraphs with ¶1, ¶2, ¶6, ¶7 and ¶10 split at their internal turns and the quoted-poem paragraphs absorbed into S11. Segment lengths 47–150 English words per arm. Reconstruction is exact on all three texts — concatenation of the 15 Japanese segments equals the source character for character, and concatenation of each arm's 15 segments equals that arm's whitespace-normalised body — and verify.py asserts it.

3. Seats, and why each is where it is

role seat why
judges — H1, H2P, H2M, GS, GD P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 the practical panel; P4 and P5 are excluded below
yardstick author and gate coder — YP, YM, GA, GY qwen/qwen3.7-max (non-panel reserve) judges nothing; the same role and the same seat as E-20260804g stage A, which config/models.md records as "a role outside the jury"
source-fidelity gate — GJ z-ai/glm-5.2 (non-panel reserve) added on critic finding 1: the seat that checks YP against the Japanese may not be the seat that wrote YP, and may not be a judge
independent pre-run critic P4 moonshotai/kimi-k3 no role in the run; charter §8

P5 deepseek/deepseek-v4-pro is excluded from every role, note (bne): it returned content: null after spending a whole 6,000-token cap on hidden reasoning with effort: low set and ignored, on plain text. The standing finding was scoped to structured output; it is broader.

P2 gets a cap of 10,000 on any task longer than a forced choice — its hidden reasoning consumed a 4,000-token cap twice in E-20260813e before the answer began.

qwen/qwen3.7-max is dispatched with reasoning: {effort: "low"} AND a cap sized to the pinned effort's observed output, never the pin alone: note (bmb)'s paired remedy, whose second half E-20260813b thought it had satisfied and had not (35 of 43 yardstick calls needed a second dispatch). Its VOID at default effort in the D-20260813-17 review is the other half of the same lesson.

z-ai/glm-5.2's price is not in config/models.md. It is read from GET /api/v1/models before dispatch and the §8 estimate is rebuilt from the read figure; the estimate below prices it at $2.00 / $8.00 per M, above every non-P4 seat in the table, so a surprise is in the conservative direction.

4. Procedure

Every call is stateless, one dispatch per cell, no conversation. Label order for a given (segment, seat) is fixed by the first byte of sha256(f"{segment_id}|{seat}"): even → arm A is R06, odd → arm A is R08. The order is seeded on (segment, seat) only, never on the condition, so that for a given (segment, seat) the H1, H2P and H2M prompts are byte-identical outside the document block and the arms sit behind the same labels in all three. verify.py asserts both.

The one definition of a segment-level answer, stated once and used everywhere

Critic finding 4 was right that the previous wording was uncomputable, and this is the replacement. Every segment-level quantity in this design — a GS match, an H1 majority, an H2 majority — is defined identically:

the choice made by a strict majority of the seats whose body for that cell is usable, with at least 2 usable bodies required. If exactly 2 bodies are usable and they disagree, or fewer than 2 are usable, the segment has NO answer for that quantity and is excluded from every statistic that uses it.

With three usable bodies and a forced binary there is always a majority, so a "split" arises only where a body was lost — which is precisely the case note (bna) says must be excluded from the denominator rather than counted as a failure. This is the "2–1 counts" branch of the critic's choice, taken over unanimity for a stated reason: unanimity would put the expected number of P2 cases at roughly 9.6 on 15 segments, below this design's own power floor of 11, so a design that adopted it would probably buy a withheld primary. The unanimity-only figures are computed and reported anyway, as a stricter secondary with its own count, so a reader can see how much of each result rests on 2–1 majorities.

Stage 1 — the yardsticks (qwen, 30 calls)

YP — the plain yardstick, one per segment, from the Japanese only:

Below is a passage of Japanese literary prose from 1896. In plain, ordinary English — the register of a clear encyclopedia entry — write two short paragraphs for a reader who knows no Japanese. The first says what the passage says: its content, event by event. The second says what the passage does to a reader of Japanese: its tone, its pace, the feeling it produces, and anything about its language that produces that feeling. Do not translate it. Do not quote English words as renderings. Do not evaluate any translation. 120–180 words total.

YM — the same content restyled, one per segment, from YP only, without the Japanese:

Rewrite the passage below in elaborate, high-Latinate critical English — the register of an ornate academic essay: long periodic sentences, abstract nouns, subordination, learned vocabulary. Change nothing about the content. Every fact, every claim about tone and feeling, must survive unchanged; only the prose may change. Do not add, remove or soften any claim. Do not use archaic, dialect or deliberately literal English, and do not imitate a translation from a foreign language: the target is ornate modern critical prose, not old or foreign-sounding prose. 120–200 words.

Why elaborate-Latinate and not estranged. Criterion 2 requires a marked style not marked in the direction of either rendering arm. R06 is plain, so a plain document is already in R06's direction — that is the confound itself, and YP is the baseline. R08 is archaic, calqued and source-ward, so the marked document must be marked in some third way. RS-20260813b §7 found that an ornate document is nevertheless matched to the foreignizing arm at 7 of 12 and to the plain arm at 0, which is what gives the manipulation its power: the document's markedness has nothing stylistically in common with R08's, and the style question still moves to R08. A co-movement of the H2 majority under those conditions would be sheer distance-from-plain, and its absence is correspondingly stronger evidence.

Stage 2 — the gates on the materials (qwen 30 calls, glm 15 calls)

GJ — does the plain yardstick actually describe the Japanese? Added on critic finding 1, which is the sharpest thing the critic found: GA checks the arms against each other and GY checks YM against YP, and nothing checked YP against the source. The source is Meiji 擬古文, the hardest register this project has read, and YP is one non-panel seat's reading of it. If YP misdescribes a passage, GY passes on a faithful restyle of a wrong description and the whole confound test runs on a document that never described the original. Per segment, to z-ai/glm-5.2, which wrote nothing:

Below is a passage of Japanese literary prose from 1896, followed by an English description of it. Read the Japanese. Does the description state anything about the passage's content that is wrong, or leave out something central to it; and does it describe the passage's tone and effect in a way a reader of the Japanese would recognise? Answer ACCURATE or WRONG. If WRONG, name what is wrong in one sentence.

GJ bites on both primaries, not only on P2 — a point the critic did not have to make and the design owes: H2P is built on YP, so P1 uses the documents too. A WRONG segment is excluded from P1 and P2 alike, and F2b (§6) voids the run if too many are wrong.

GA — do the two arms state the same facts? Per segment, both arms, labels A/B:

Two English translations of the same Japanese passage. Ignoring completely all differences of wording, register, style, word order, and how foreign or old-fashioned the English sounds, and ignoring differences in how proper names, measures and culture-specific words are handled: do the two passages state the same narrative facts — the same events, in the same order, involving the same participants? Answer SAME or DIFFERENT. If DIFFERENT, name the fact that differs in one sentence.

Registered carve-out, stated before the run because it is foreseen and not discovered. S11 and S15 contain the same quoted line of a Scottish poem rendered two different ways: R06 restores Burns's English ("Farewell to the mountains high cover'd with snow"), R08 calques Doppo's Japanese of it ("Now then, farewell, high peaks that wear the snow"). That is a difference of wording inside a quotation, not of narrative fact, and the prompt's first clause excludes it. If GA flags S11 or S15 on the quoted line alone, the flag is recorded and not counted. Any other flag counts.

GY — do the two documents say the same things? Per segment, YP and YM:

Two descriptions of the same passage, written in very different English. Ignoring style completely: does the second state any fact, or any claim about tone or feeling, that the first does not — or drop, weaken or strengthen any that the first states? Answer SAME or CHANGED, and if CHANGED name the change in one sentence.

Stage 3 — the directionality validation GD (judges, 96 calls)

Criterion 2's amended clause: the manipulation must be validated as directional before it is trusted as a control. Critic findings 2 and 3 rebuilt this stage and both were right: as first written it validated one segment's YM while P2 consumed all fifteen, so a discharge could have rested on fourteen unvalidated documents — the exact failure criterion 2 was amended to prevent — and it never checked that the plain document is plain, without which the YP-vs-YM contrast is not a plain-vs-marked contrast at all. So: all 15 YM documents, all 15 YP documents, and one segment of each arm (from the pre-registered index S07, chosen before the run because it is mid-length, narrative, and carries no quoted poem), each to all three judging seats — 96 calls:

Here is a piece of English prose. If it departs from plain, ordinary modern English, in which direction does it depart? Answer with exactly one of: PLAIN (it does not markedly depart); ORNATE (elaborate, Latinate, abstract, learned — the direction of an ornate modern essay); SOURCEWARD (archaic, literal, calqued, foreign-sounding — the direction of a translation that keeps its original's shape). Then one sentence of reason.

The prompt does not say the passages are translations, does not mention the original, and does not mention any other item.

Criterion 2 is satisfied iff all four hold:

  1. the R08 sample is SOURCEWARD by the §4 majority rule, and the R06 sample is PLAIN;
  2. YM is ORNATE on ≥ 12 of the 15 segments;
  3. YP is PLAIN on ≥ 12 of the 15 segments;
  4. per segment: a segment whose YM is not ORNATE, or whose YP is not PLAIN, is excluded from P2 individually, whatever the run-level counts do.

Clause 1 is what makes the manipulation directional rather than merely large: the marked document is measured to be marked in a different direction from the marked arm. If criterion 2 fails at the run level, P2 is still computed and reported over the segments that survive clause 4, but the run may not discharge condition 3 and the result says so in those words.

Stage 4 — the style match GS (judges, 90 calls)

Per (segment × document × seat), target-only, forced binary:

Below are a short document and two English passages, A and B. Ignoring what any of them mean, and considering only their English prose — vocabulary, sentence construction, rhythm, register — which passage's prose style is closer to the document's? Answer A or B, then one sentence of reason.

The prompt mentions no original, no effect, no comparability, no translation. verify.py asserts that no GS prompt contains the words original, Japanese, source, effect, comparable or translation.

Stage 5 — the judgments (judges, 135 calls)

H1, per (segment × seat), no document:

Below are two English passages, A and B. Which does more to you as a reader of English — which has the stronger effect on you? Answer A or B, then one sentence of reason. There is no third answer.

H2P and H2M, per (segment × seat × document), identical to each other outside the document block:

Below is a description of a passage of foreign-language prose, followed by two English passages, A and B. Which of A and B comes closer to doing to its English reader what the original does to a reader of the original? Answer A or B, then one sentence of reason. There is no third answer.

The clause "written by someone who has read it in the original" was cut on critic finding 6, which is small and exactly right: it is true of YP and false of YM, a restyle by a seat that never saw the Japanese. A seat that reasons about the description's provenance would have been misinformed in one arm of the manipulation and not the other — and the cut also strengthens the byte-identity assertion in §4, since the sentence no longer says anything the two documents differ on.

Total: 396 dispatches.

5. The registered predictions

P1 — the separation replicates (criterion 1)

Over segments with a majority (§4 rule) in both H1 and H2P, excluding any segment GJ marked WRONG:

The null being tested is conditional, and critic finding 8 is right that the design should say so: given that H1 and H2P disagree on a segment, the direction of the disagreement is a coin flip. The binomial is computed on a set selected for disagreement, which is legitimate under that null and only under it; this test is not, and may not later be read as, an unconditional comparison of the two conditions' rates.

P1 HOLDS iff all three. Cell-level counts (segment × seat) and per-seat tables are reported as description, never as the primary.

P2 — the confound invariance, as an equivalence test (criterion 4)

Cases are segments where the GS style match (§4 majority rule) differs between YP and YM — YP matched to one arm and YM to the other — and which survive all four individual exclusions: GJ ACCURATE, GY SAME, YM ORNATE, YP PLAIN. A co-movement is a case where the H2 majority also differs between the two documents, in the direction the style match moved.

The per-segment exclusions are critic findings 5 and 2 made operative. As first written, a segment whose YM had drifted in content stayed eligible unless drift reached three segments — so one or two content-drifted documents could manufacture or mask a co-movement and the run could still discharge. Individual exclusion closes that; F2 survives as the whole-run void.

The critical values, computed by exhaustive enumeration and recomputed by verify.py:

n reject iff co-movements ≤ exact P at that count
11 2 0.0327
12 2 0.0193
13 3 0.0461
14 3 0.0287
15 3 0.0176

Registered power. Against the confound hypothesis p = 1.0 the test rejects with probability 0 — it cannot discharge the condition if the confound is operating, which is the property criterion 4 exists to demand. Against p = 0.15 (the co-movement rate RS-20260813b §6.1 recorded without licence) power at n = 14 is 0.853; against p = 0.30 it is 0.355. The test is therefore honest in both directions and is not a formality: a moderate confound defeats it.

On a discharge the result reports the observed co-movement rate and its exact one-sided upper confidence bound beside criterion 4's wording — critic finding 7, which is arguable and was taken anyway. At n = 14 a discharge means ≤ 3 co-movements, an observed rate ≤ 0.214 with a 95% upper bound of 0.466 (Clopper-Pearson, exact); the phrase "equivalent to invariance within a margin of 0.50" licenses a reader to imagine half the cases co-moved, and the number is stronger than the margin. The number is reported; the margin is what was registered.

Failing to reject is not evidence for the confound, and the result page will say so in those words. A high co-movement count is evidence for it, and if co-movements ≥ ⌈n/2⌉ the result records that the style reading of limit 5 is supported, and RS-20260812f §5 and RS-20260813b §5 acquire a caution heavy enough that neither direction may be cited without it.

P3 — the manipulation check (criterion 3)

Computed over valid segments only — segments with a GS match under both documents by the §4 rule, so that a segment whose seats could not be resolved for want of usable bodies is excluded from the denominator rather than counted as a failure (note (bna), the defect that voided E-20260813b):

P3 HOLDS iff all three. F3 fires if it does not, and P2 is void.

Registered here, in the design, because it is the way this run could most easily fool itself: no override of F3 will be taken, whatever the withheld comparison shows. RS-20260812f §7 set a precedent for overriding a failure criterion with the reason on the record; RS-20260813b refused to use it because the override would have rescued that run's own headline, and the same refusal is pre-committed here. A withheld comparison may be recorded, flagged at every use as unlicensed, and may not appear in any conclusion.

6. Failure criteria

fires when consequence
F1 GA returns DIFFERENT on ≥ 3 segments, excluding the registered quoted-line carve-out the arms are not comparable; the whole run is void
F2 GY returns CHANGED on ≥ 3 segments the restyle moved content; P2 void, P1 unaffected
F2b GJ returns WRONG on ≥ 5 of 15 segments the plain yardstick does not describe the source; the whole run is void, P1 included, because H2P is built on YP
F3 P3 fails any clause P2 void, reported as void and not as a null
F4 for a segment: no usable document under either style · GJ WRONG · GY CHANGED · YM not ORNATE · YP not PLAIN that segment is excluded from P2; a GJ WRONG segment is excluded from P1 as well
F5 dead bodies exceed 10% of dispatches on any stage that stage is void
F6 — the truncation guard: any body that fails to parse, or whose text ends mid-token, is re-dispatched at double the cap irrespective of finish_reason (note (bmz)); cost is accumulated across attempts, never overwritten by the successful one (the RS-20260813b §8 runner defect, whose remedy is one line)

7. The result → option map, registered in advance

P1 GD (criterion 2) P3 (criterion 3) P2 (criterion 4) what is recorded
HOLDS satisfied HOLDS HOLDS D-20260804-16 condition 3 DISCHARGES. All four criteria met. wiki/goodness-senses.md §affect records the discharge and criterion 5's residual caution verbatim: discharge defeats the style reading of limit 5 only; H2 inherently contains a document and H1 does not, and every citation of either run's direction carries that until a design outside this family addresses it. Usage rule 4 stands either way — it is not what condition 3 governs
HOLDS any HOLDS co-movement ≥ ⌈n/2⌉ The confound is supported. Condition 3 does not discharge; RS-20260812f §5 and RS-20260813b §5 are re-flagged so that neither direction may be cited without it. The arm closes on a positive finding against the project's own prior result
HOLDS not satisfied HOLDS HOLDS No discharge. The manipulation was not validated as directional, so what P2 measures is not the confound criterion 2 names. Reported in full, with the arm closing on the failure and the missing instrument named
HOLDS any FAILS void No discharge, F3, P2 void. The arm closes resolved with condition 3 open and the second failure of this manipulation on the record — two failures make the manipulation itself the finding
WITHHELD or fails any any any The replication did not reproduce on a third family. Condition 3 does not discharge, and RS-20260812f/RS-20260813b's generality is the finding rather than the confound

In no cell does the arm continue past step 2. ARM-affect-confound's completion criterion is that condition 3 ends up recorded one way or the other; an inconclusive outcome is a failure to discharge and closes the arm, not a licence for step 3.

8. Pre-flight cost, built from caps and not from expected lengths

Note (abc): the one estimate this project has overrun was built from assumed output lengths. Rates from config/models.md (in / out per M): P1 $1.00/$6.00, P2 $1.50/$7.50, P3 $2.00/$6.00 — judge average $1.50/$6.50; qwen/qwen3.7-max $1.475/$4.425.

stage calls seat in cap out cap worst case
YP + YM 30 qwen 700 1,600 $0.2434
GY 15 qwen 800 700 $0.0642
GA 15 qwen 500 700 $0.0576
GJ 15 glm 900 700 $0.1110
GD 96 judges 600 700 $0.5232
GS 90 judges 700 500 $0.3870
H1 45 judges 700 600 $0.2228
H2P + H2M 90 judges 950 700 $0.5378
396 $2.1470

Plus a 25% re-dispatch allowance for the F6 guard: $2.684. Declared ceiling for step 2: $2.80. The critic's remedies cost $0.71 of worst case, most of it the rebuilt GD stage, and that is the right trade: the stage as first written could have let a discharge rest on fourteen unvalidated documents. P2's cap is raised to 10,000 on any call it returns truncated, per §3. The estimate is built from max_tokens, and max_tokens is not a guarantee (note (bgk)): billed completion has exceeded a declared cap once in this project's history, on a routed provider.

A $2.80 ceiling is 56% of a UTC day's cap and several sessions may share a day. Step 2 checks the day's rows before dispatching and, if the headroom is not there, splits at the stage boundary — stages 1–4 (yardsticks, gates, GD, GS, worst case $1.39) are self-contained and produce the manipulation check; stage 5 is the judging. Splitting is preferred to scaling down, because every bar in §5 is registered against 15 segments.

Step 1 spent only the pre-run critic: $0.095881200 against a declared ceiling of $0.35 and $1.046484 of headroom on UTC day 2026-08-13.

9. What this design cannot establish, stated before it runs

  1. Model seats, not readers. Tier D is NOT PASSED. Nothing here is calibrated, and every figure is a fact about how three seats behave.
  2. Criterion 5 survives any discharge. H2 contains a document and H1 does not. No design in this family removes that; a comparability question with no description of the original in it is not a comparability question.
  3. One hand, one order. Both arms are the lead, in one session, R06 first. RS-20260813b limit 3 carries over unchanged: this replicates a lead-vs-lead contrast, not an independent one. The 12-token self-match in §2 is the size of the shared substrate.
  4. Three things move at once — family, affect register, source register — so a replication generalises and localises nothing.
  5. GS is an unanchored model judgment about style, and RS-20260813b §7 suggests it measures distance from plain English rather than kind of markedness. GD is this design's answer, and GD is itself an unanchored model judgment about style.
  6. YM is written by a model from YP, not by a human from the Japanese. The restyle can only be as good as one seat's ornate English, and GY can only catch content drift it can see.

10. Verification

analysis/verify.py, written after this page is frozen and importing nothing from tools/, recomputes every reported number from the stored bodies and asserts at minimum:

  1. the 15 Japanese segments concatenate to the source character for character, and each arm's 15 segments concatenate to that arm's whitespace-normalised body;
  2. no H1 prompt contains a document;
  3. no GS or GD prompt contains the words original, Japanese, source, effect, comparable or translation;
  4. the H1, H2P and H2M prompts for a given (segment, seat) are byte-identical outside the document block, and the label order matches the seed rule;
  5. every individual cell choice against the stored body — not the majorities (the mutation that escaped E-20260813b's first verifier pass was a single flipped cell inside a 3–0 block);
  6. the exact binomial P values for P1 and P2 by exhaustive enumeration, and the critical-value table in §5;
  7. the cost total equals the sum over all attempts including discarded ones;
  8. every §4 majority is recomputed from the usable bodies, and every segment excluded by F4 is excluded for a reason present in the stored GJ / GY / GD bodies — the P2 denominator is reconstructed from the raw record and compared to the reported one;
  9. no GJ prompt contains either arm, and no GD prompt contains more than one item.

Mutation tests, ≥ 8, each caught or the verifier is repaired before any number is reported: a flipped H2P cell inside a unanimous block; a GS match inverted on one segment; a document swapped between YP and YM on one segment; a segment dropped from the P2 denominator; a segment added to it that F4 excludes; a label order flipped for one (segment, seat) in H2M only; a GJ WRONG verdict silently ignored; a discarded attempt's cost omitted.

11. Pre-run critic findings, and what was done with them

Seat P4 moonshotai/kimi-k3, no role in the run, one dispatch, finish_reason: stop, provider Together, $0.095881200. Verdict PROCEED-WITH-AMENDMENT, eight findings, five BLOCKING. All eight accepted; raw body run/critic-P4.json. No study call is dispatched until this section is complete, and it is complete.

# severity finding what was done
1 BLOCKING nothing validated YP against the Japanese — GA compares the arms, GY compares the documents, and a wrong YP would pass both and carry the whole test accepted in full and extended. New gate GJ (§4 stage 2) on a fourth seat, z-ai/glm-5.2, which wrote nothing; per-segment exclusion via F4; run-level void via F2b. Extended beyond the critic's remedy: the critic scoped exclusion to P2, and H2P is built on YP, so a WRONG segment leaves P1 too
2 BLOCKING GD validated one segment's YM while P2 consumed fifteen, so a discharge could rest on fourteen unvalidated documents — the failure criterion 2 exists to prevent accepted. GD now runs on all 15 YM documents at three seats, with per-segment exclusion of any YM not ORNATE. +$0.25 of worst case
3 BLOCKING the criterion-2 gate never checked that the plain document is plain; an ornate YP would leave the contrast not plain-vs-marked while criterion 2 formally passed accepted. GD also runs on all 15 YP documents, and on the R06 sample; criterion 2 now has four clauses, §4 stage 3
4 BLOCKING "two-seat agreement" is uncomputable with three seats and a forced binary — every non-failed cell has one — so P3's denominators and P2's n had no definition accepted, and it forced the design's one substantive choice. The §4 majority rule replaces it, stated once and propagated. The critic offered unanimity as the presumed reading; that branch was declined with the reason on the record — unanimity puts expected P2 cases near 9.6 against this design's own power floor of 11 — and the unanimity figures are reported as a stricter secondary
5 BLOCKING content-drifted documents stayed eligible for P2 unless drift reached three segments, so one or two could manufacture or mask a co-movement and still discharge accepted. GY CHANGED now excludes the segment individually through F4; F2 survives as the whole-run void
6 MINOR the H2 prompt said the description was "written by someone who has read it in the original" — true of YP, false of YM accepted. The clause is cut from both prompts. It also removes a sentence the two documents differed on, which strengthens the byte-identity assertion
7 MINOR (the critic marked it arguable) δ = 0.50 in the discharge wording licenses a reader to imagine half the cases co-moved, when a discharge in fact bounds the rate near 0.21 accepted. On a discharge the observed rate and its exact upper confidence bound are reported beside criterion 4's wording. The margin stays what was registered
8 MINOR P1's binomial runs on a set selected for disagreement; the design should name the conditional null so it cannot later be read as an unconditional comparison accepted. The conditional null is stated in §5 P1, with the prohibition attached

What the critic did not find, recorded so a later reader can see the gap this pass leaves. It raised nothing about the materials (its priority 4) — no objection to Meiji 擬古文 as the source, to the two arms' comparability, or to three things moving at once between this run and its two predecessors. That may be because there is nothing there, or because one segment of material was all it was shown. §9 limits 3 and 4 carry those risks unrefereed.

12. Amendments made at dispatch (step 2, 2026-08-14), before any study call

Written and committed before run.py existed and before one dispatch was made, so that a reader can see which of these came before the numbers. Nothing here moves a bar, a threshold, a denominator, a failure criterion or the result→option map of §7. Two are corrections to §10's verifier checklist where it contradicts §4's operative text; two are transport protections.

A1 — §10 assertion 3's lexical ban is asserted on GS only. GD is asserted on §4 stage 3's own properties instead. §4 stage 3's frozen GD prompt contains the words translation and original inside the gloss of the SOURCEWARD option — "the direction of a translation that keeps its original's shape" — which is the clause that makes the three-way choice mean anything, and without which SOURCEWARD is an undefined label. §4 stage 4 scopes the lexical assertion to GS ("verify.py asserts that no GS prompt contains…"); §10 assertion 3 over-extended it to GD, where the design's own prompt text cannot satisfy it. The prompt is not reworded — a frozen instrument is not edited to make a checklist pass. What verify.py asserts for GD is the sentence §4 stage 3 already states: the prompt does not say the passages are translations, does not mention the original, and does not mention any other item — operationalised as: the prompt contains exactly one prose item, contains no reference to that item's provenance, and contains no Japanese character.

A2 — §10 assertion 4's byte-identity is asserted where it is true. H1 asks a different question from H2 by construction — that difference is the design — so the three prompts cannot be byte-identical outside the document block, and §4's sentence saying they are is loose. What is asserted, and what §4's purpose actually requires: (i) H2P and H2M for a given (segment, seat) are byte-identical outside the document block; (ii) the two-passage block — the A and B texts in their label order — is byte-identical across H1, H2P and H2M for that (segment, seat); (iii) the label order matches the §4 seed rule. (i) is the pairing the confound test needs; (ii) is what makes H1 and H2P comparable segment by segment.

A3 — reasoning effort is pinned low on P2 and on z-ai/glm-5.2, per method note (bnk), which post-dates this freeze. The note fired on 2026-08-14 and cost $0.526018: on a reasoning seat max_tokens caps the answer plus the hidden thinking, and the thinking takes all of it, leaving an empty body that is billed. §3 already pins qwen/qwen3.7-max and already flags P2 for cap inflation; this extends the same protection to P2 and to glm, whose reasoning behaviour is unmeasured in this project. P1 and P3 are left unpinned, as in E-20260813b, where they returned usable bodies at caps of 400–900. This is a protection on the transport; it changes no prompt, bar or statistic, and it is recorded because it does change what the seats do.

A4 — the machine-readable reply line. §4 gives each prompt's text but no reply format. Each prompt is dispatched verbatim with one appended line requiring a single-line JSON reply. The appended lines are in run.py and are reproduced in the result; verify.py checks that the GS appendix contains none of the six banned words, and that every prompt's §4 body is present unmodified.

A6 — the transport, rebuilt on measurement after the first dispatched call breached its own cap ladder. This is the largest amendment and it is written before the second call. yp__S01 was dispatched to qwen/qwen3.7-max at the §8 cap of 1,600 with reasoning: {"effort": "low"} set, as §3 requires. It came back empty; the F6 guard re-dispatched at 3,200, which came back empty; the third attempt at 6,400 returned a body cut off mid-sentence with finish_reason: "stop". Billed: $0.043586 for one unusable cell, against a §8 worst case of $0.0074 for the whole call. The usage record says why — 4,662 reasoning tokens for 171 visible ones. The pin was sent and ignored, exactly as note (bne) records for deepseek-v4-pro, now on a second seat and a second provider. That body is preserved under runs/_discarded/ and its cost is counted (§10.7). Three things follow, all measured before re-dispatch:

  1. reasoning: {"enabled": false} on the two seats that judge nothing — qwen/qwen3.7-max and z-ai/glm-5.2. Probed once each: 0 reasoning tokens, complete bodies, $0.001081 and $0.000341. This is §3's stated intent for qwen ("effort low") executed on a provider that honours it, not a new intent.
  2. Reasoning is NOT disabled on any judging seat. Whether a judge thinks before answering is part of the instrument, and muting it on P1, P2 and P3 would change what GD, GS, H1 and H2 measure. P2's effort: low pin from A3 is withdrawn for the same reason — A3 was written before this was measured, and it reached further than its evidence.
  3. Caps re-sized from measured reasoning use, not from assumed answer length — note (abc)'s rule applied to the quantity that actually varies. Probed on the real prompts: P1 0–403 reasoning tokens, P3 326–430, P2 up to 1,043 on GS. §8's caps of 500–700 were built for seats that do not think, so every judge stage would have truncated. Caps: judges 1,200, P2 2,200 (it reasons hardest and is the cheapest out), yp/ym 1,200, gj/ga/gy 500.

The declared ceiling changes, and the split §8 pre-registered is taken. The re-sized cap-built worst case is $3.65 against §8's $2.80 — the increase is entirely the judges' thinking room, and it is offset in part by A5's price falls. Today's UTC headroom is $3.56, which the full run does not fit. §8's own instruction is split at the stage boundary, in preference to scaling down, because every bar in §5 is registered against 15 segments, and that is what is done: stages 1–4 (worst case $2.26) are dispatched first, the gates and the manipulation check are computed, and stage 5 (worst case $1.39) is dispatched only if the run is not already void and the headroom is there. No bar, threshold, denominator or failure criterion moves. Expected actual from the probes is ≈$1.05 for all 396 calls; the ceiling is built from caps because that is the rule.

A5 — prices re-read from GET /api/v1/models at dispatch, as §3 requires for glm. Read 2026-08-14: z-ai/glm-5.2 $0.63 / $1.98 per M against the $2.00 / $8.00 the §8 estimate priced it at, and google/gemini-3.6-flash (P2) $0.75 / $3.75 against config/models.md's $1.50 / $7.50 — a halving on a judging seat. openai/gpt-5.6-terra $1.00 / $6.00, x-ai/grok-4.5 $2.00 / $6.00 and qwen/qwen3.7-max $1.475 / $4.425 read back unchanged. Both surprises are in the conservative direction, so the §8 worst case and the $2.80 declared ceiling stand unrevised; config/models.md is corrected in place.