Repository path: workshop/experiments/E-20260813f-affect-confound/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260813f-affect-confound |
| status | frozen |
| created | 2026-08-13 |
| updated | 2026-08-13 |
| senses | affect |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-affect-confound.md, wiki/decisions/resolved/D-20260813-17-affect-halves-discharge.md, wiki/decisions/votes/2026-08-13/D-20260813-17-ratification-record.md, wiki/decisions/resolved/D-20260804-16-affect-two-halves.md, wiki/goodness-senses.md, wiki/findings/results/RS-20260813b-affect-yardstick.md, wiki/findings/results/RS-20260812f-affect-unprompted.md, workshop/translations/hoshi/R06-v1/translation.md, workshop/translations/hoshi/R08-v1/translation.md, workshop/regimes/R06-lead-single-pass.md, workshop/regimes/R08-resistancy.md, config/models.md, config/budget.md |
E-20260813f — does the comparability judgment follow the yardstick's prose style?
ARM-affect-confound step 1 writes this design; step 2 dispatches it. Frozen at 030c3b5
before the pre-run critic saw it, amended before dispatch in the same session on the critic's
findings — §11 records all eight, five of them BLOCKING, and what was done with each. Everything
below is internal-judgment-only and provisional: affect is untested, Tier D is NOT PASSED,
and no jury verdict in this project carries evidential weight.
1. The question, and what it is a question about
wiki/goodness-senses.md splits affect into two halves:
H1— what the English does to a reader of English.H2— how comparable that is to what the source does to a reader of the source.
Two runs found the halves separate, unprompted, in opposite directions: the more marked rendering
wins H1, the plainer rendering wins H2. Both hang on one confound, RS-20260812f limit 5:
the H2 question comes with a written description of the original attached and the H1 question
comes with nothing, so "a preference for the plainer arm may be a preference for the arm that
matches a plainly-written English description."
The question: does the
H2majority move when the yardstick document's prose style moves?
What it is a question about (the subject rule, continue-prompt.md §4.5): whether a reader who
cannot read the source can be told which of two renderings comes closer to what the original does
to its own reader. Every "equivalent effect" argument in translation rests on that being
answerable. If the answer is that the judgment tracks the description's prose rather than the
original's effect, then no such reader can be told, and this project's two affect findings say
nothing about translation.
2. Materials, all frozen before this design existed
国木田独歩, 「星」 ("The Star"), November 1896, the piece whole. Public domain (Doppo
1871–1908). Copy-text Aozora Bunko card 42207, 000038/42207_34797.html, fetched 2026-08-13 with
tools/fetch_aozora.py; 底本 「武蔵野」岩波文庫. 12 paragraphs, 2,157 non-whitespace characters,
read whole in Japanese. workshop/translations/hoshi/source.txt.
The project's first Kunikida Doppo source and its first text in Meiji 擬古文 — a deliberately archaised quasi-classical register with classical verb morphology and free alternation between narrative past and stative present. Nothing in modern English corresponds to it, which is why the two arms differ as widely as they do.
Three firsts that matter to what a replication here would mean. Japonic is a third language
family after Slavic (RS-20260812f, Bulgarian) and Romance (RS-20260813b, Romanian); the piece
is elegiac where both prior works were comic; and the source's register is archaic, so the
R08 arm's markedness is licensed by the source rather than imposed on it. Three things move at
once. A replication therefore generalises broadly and localises nothing — stated here, before
the run, so that no post-hoc reading of which factor mattered is available.
Two arms, both lead, both $0, both frozen before this page existed:
| arm | id | commit | English words | what it is |
|---|---|---|---|---|
R06 |
T-hoshi-R06-v1 |
1822ac7 |
1,534 | lead single pass, source only, no rule set; classical register rendered as plain modern narrative English |
R08 |
T-hoshi-R08-v1 |
5f1fd0b |
1,481 | + Venuti's ten foreignizing rules frozen 2026-07-28; archaism, calque, unassimilated retention, source clause order |
R07 is not built, for the reason RS-20260812f §3 established on two independent seats: the
fluency programme taken whole is H1 under another name, and a third arm aimed at a half would
be excluded by the criterion-1 gate anyway.
Contamination. Both artifacts declare none on a reachability check, not a same-text
measurement, and say so on their faces: no English rendering of 「星」 is reachable, Project
Gutenberg carries zero Doppo, and English Wikisource carries one short Doppo poem. No cross-text
comparator large enough to bound anything exists, unlike the Romanian case which had 8,748 words of
the same author's period English. What was measured is the two arms against each other, because
they are one hand: tools/dependence_check.py returns longest common run 12 tokens, 72 shared
7-grams, 2 shared 12-grams, 0 shared 15-grams — below the 37-token lead self-match record of note
(bhb), and reported because a design that treats these two texts as distinguishable owes a number
saying how distinguishable they are.
15 segments, materials/segments.json, cut at sentence boundaries chosen on the source's
structure before any prompt existed: the source's 12 paragraphs with ¶1, ¶2, ¶6, ¶7 and ¶10 split
at their internal turns and the quoted-poem paragraphs absorbed into S11. Segment lengths 47–150
English words per arm. Reconstruction is exact on all three texts — concatenation of the 15
Japanese segments equals the source character for character, and concatenation of each arm's 15
segments equals that arm's whitespace-normalised body — and verify.py asserts it.
3. Seats, and why each is where it is
| role | seat | why |
|---|---|---|
judges — H1, H2P, H2M, GS, GD |
P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5 |
the practical panel; P4 and P5 are excluded below |
yardstick author and gate coder — YP, YM, GA, GY |
qwen/qwen3.7-max (non-panel reserve) |
judges nothing; the same role and the same seat as E-20260804g stage A, which config/models.md records as "a role outside the jury" |
source-fidelity gate — GJ |
z-ai/glm-5.2 (non-panel reserve) |
added on critic finding 1: the seat that checks YP against the Japanese may not be the seat that wrote YP, and may not be a judge |
| independent pre-run critic | P4 moonshotai/kimi-k3 |
no role in the run; charter §8 |
P5 deepseek/deepseek-v4-pro is excluded from every role, note (bne): it returned
content: null after spending a whole 6,000-token cap on hidden reasoning with effort: low
set and ignored, on plain text. The standing finding was scoped to structured output; it is
broader.
P2 gets a cap of 10,000 on any task longer than a forced choice — its hidden reasoning
consumed a 4,000-token cap twice in E-20260813e before the answer began.
qwen/qwen3.7-max is dispatched with reasoning: {effort: "low"} AND a cap sized to the pinned
effort's observed output, never the pin alone: note (bmb)'s paired remedy, whose second half
E-20260813b thought it had satisfied and had not (35 of 43 yardstick calls needed a second
dispatch). Its VOID at default effort in the D-20260813-17 review is the other half of the
same lesson.
z-ai/glm-5.2's price is not in config/models.md. It is read from GET /api/v1/models
before dispatch and the §8 estimate is rebuilt from the read figure; the estimate below prices it
at $2.00 / $8.00 per M, above every non-P4 seat in the table, so a surprise is in the
conservative direction.
4. Procedure
Every call is stateless, one dispatch per cell, no conversation. Label order for a given
(segment, seat) is fixed by the first byte of sha256(f"{segment_id}|{seat}"): even → arm A is
R06, odd → arm A is R08. The order is seeded on (segment, seat) only, never on the
condition, so that for a given (segment, seat) the H1, H2P and H2M prompts are byte-identical
outside the document block and the arms sit behind the same labels in all three. verify.py
asserts both.
The one definition of a segment-level answer, stated once and used everywhere
Critic finding 4 was right that the previous wording was uncomputable, and this is the
replacement. Every segment-level quantity in this design — a GS match, an H1 majority,
an H2 majority — is defined identically:
the choice made by a strict majority of the seats whose body for that cell is usable, with at least 2 usable bodies required. If exactly 2 bodies are usable and they disagree, or fewer than 2 are usable, the segment has NO answer for that quantity and is excluded from every statistic that uses it.
With three usable bodies and a forced binary there is always a majority, so a "split" arises only
where a body was lost — which is precisely the case note (bna) says must be excluded from the
denominator rather than counted as a failure. This is the "2–1 counts" branch of the critic's
choice, taken over unanimity for a stated reason: unanimity would put the expected number of
P2 cases at roughly 9.6 on 15 segments, below this design's own power floor of 11, so a design
that adopted it would probably buy a withheld primary. The unanimity-only figures are computed
and reported anyway, as a stricter secondary with its own count, so a reader can see how much of
each result rests on 2–1 majorities.
Stage 1 — the yardsticks (qwen, 30 calls)
YP — the plain yardstick, one per segment, from the Japanese only:
Below is a passage of Japanese literary prose from 1896. In plain, ordinary English — the register of a clear encyclopedia entry — write two short paragraphs for a reader who knows no Japanese. The first says what the passage says: its content, event by event. The second says what the passage does to a reader of Japanese: its tone, its pace, the feeling it produces, and anything about its language that produces that feeling. Do not translate it. Do not quote English words as renderings. Do not evaluate any translation. 120–180 words total.
YM — the same content restyled, one per segment, from YP only, without the Japanese:
Rewrite the passage below in elaborate, high-Latinate critical English — the register of an ornate academic essay: long periodic sentences, abstract nouns, subordination, learned vocabulary. Change nothing about the content. Every fact, every claim about tone and feeling, must survive unchanged; only the prose may change. Do not add, remove or soften any claim. Do not use archaic, dialect or deliberately literal English, and do not imitate a translation from a foreign language: the target is ornate modern critical prose, not old or foreign-sounding prose. 120–200 words.
Why elaborate-Latinate and not estranged. Criterion 2 requires a marked style not marked in
the direction of either rendering arm. R06 is plain, so a plain document is already in R06's
direction — that is the confound itself, and YP is the baseline. R08 is archaic, calqued and
source-ward, so the marked document must be marked in some third way. RS-20260813b §7 found
that an ornate document is nevertheless matched to the foreignizing arm at 7 of 12 and to the plain
arm at 0, which is what gives the manipulation its power: the document's markedness has nothing
stylistically in common with R08's, and the style question still moves to R08. A co-movement
of the H2 majority under those conditions would be sheer distance-from-plain, and its absence is
correspondingly stronger evidence.
Stage 2 — the gates on the materials (qwen 30 calls, glm 15 calls)
GJ — does the plain yardstick actually describe the Japanese? Added on critic finding 1,
which is the sharpest thing the critic found: GA checks the arms against each other and GY
checks YM against YP, and nothing checked YP against the source. The source is Meiji
擬古文, the hardest register this project has read, and YP is one non-panel seat's reading of it.
If YP misdescribes a passage, GY passes on a faithful restyle of a wrong description and the
whole confound test runs on a document that never described the original. Per segment, to
z-ai/glm-5.2, which wrote nothing:
Below is a passage of Japanese literary prose from 1896, followed by an English description of it. Read the Japanese. Does the description state anything about the passage's content that is wrong, or leave out something central to it; and does it describe the passage's tone and effect in a way a reader of the Japanese would recognise? Answer
ACCURATEorWRONG. IfWRONG, name what is wrong in one sentence.
GJ bites on both primaries, not only on P2 — a point the critic did not have to make and
the design owes: H2P is built on YP, so P1 uses the documents too. A WRONG segment is
excluded from P1 and P2 alike, and F2b (§6) voids the run if too many are wrong.
GA — do the two arms state the same facts? Per segment, both arms, labels A/B:
Two English translations of the same Japanese passage. Ignoring completely all differences of wording, register, style, word order, and how foreign or old-fashioned the English sounds, and ignoring differences in how proper names, measures and culture-specific words are handled: do the two passages state the same narrative facts — the same events, in the same order, involving the same participants? Answer
SAMEorDIFFERENT. IfDIFFERENT, name the fact that differs in one sentence.
Registered carve-out, stated before the run because it is foreseen and not discovered. S11 and
S15 contain the same quoted line of a Scottish poem rendered two different ways: R06 restores
Burns's English ("Farewell to the mountains high cover'd with snow"), R08 calques Doppo's
Japanese of it ("Now then, farewell, high peaks that wear the snow"). That is a difference of
wording inside a quotation, not of narrative fact, and the prompt's first clause excludes it. If
GA flags S11 or S15 on the quoted line alone, the flag is recorded and not counted. Any other
flag counts.
GY — do the two documents say the same things? Per segment, YP and YM:
Two descriptions of the same passage, written in very different English. Ignoring style completely: does the second state any fact, or any claim about tone or feeling, that the first does not — or drop, weaken or strengthen any that the first states? Answer
SAMEorCHANGED, and ifCHANGEDname the change in one sentence.
Stage 3 — the directionality validation GD (judges, 96 calls)
Criterion 2's amended clause: the manipulation must be validated as directional before it is
trusted as a control. Critic findings 2 and 3 rebuilt this stage and both were right: as
first written it validated one segment's YM while P2 consumed all fifteen, so a discharge could
have rested on fourteen unvalidated documents — the exact failure criterion 2 was amended to
prevent — and it never checked that the plain document is plain, without which the YP-vs-YM
contrast is not a plain-vs-marked contrast at all. So: all 15 YM documents, all 15 YP
documents, and one segment of each arm (from the pre-registered index S07, chosen before the
run because it is mid-length, narrative, and carries no quoted poem), each to all three judging
seats — 96 calls:
Here is a piece of English prose. If it departs from plain, ordinary modern English, in which direction does it depart? Answer with exactly one of:
PLAIN(it does not markedly depart);ORNATE(elaborate, Latinate, abstract, learned — the direction of an ornate modern essay);SOURCEWARD(archaic, literal, calqued, foreign-sounding — the direction of a translation that keeps its original's shape). Then one sentence of reason.
The prompt does not say the passages are translations, does not mention the original, and does not mention any other item.
Criterion 2 is satisfied iff all four hold:
- the
R08sample isSOURCEWARDby the §4 majority rule, and theR06sample isPLAIN; YMisORNATEon ≥ 12 of the 15 segments;YPisPLAINon ≥ 12 of the 15 segments;- per segment: a segment whose
YMis notORNATE, or whoseYPis notPLAIN, is excluded fromP2individually, whatever the run-level counts do.
Clause 1 is what makes the manipulation directional rather than merely large: the marked document
is measured to be marked in a different direction from the marked arm. If criterion 2 fails at the
run level, P2 is still computed and reported over the segments that survive clause 4, but the
run may not discharge condition 3 and the result says so in those words.
Stage 4 — the style match GS (judges, 90 calls)
Per (segment × document × seat), target-only, forced binary:
Below are a short document and two English passages,
AandB. Ignoring what any of them mean, and considering only their English prose — vocabulary, sentence construction, rhythm, register — which passage's prose style is closer to the document's? AnswerAorB, then one sentence of reason.
The prompt mentions no original, no effect, no comparability, no translation. verify.py
asserts that no GS prompt contains the words original, Japanese, source, effect,
comparable or translation.
Stage 5 — the judgments (judges, 135 calls)
H1, per (segment × seat), no document:
Below are two English passages,
AandB. Which does more to you as a reader of English — which has the stronger effect on you? AnswerAorB, then one sentence of reason. There is no third answer.
H2P and H2M, per (segment × seat × document), identical to each other outside the
document block:
Below is a description of a passage of foreign-language prose, followed by two English passages,
AandB. Which ofAandBcomes closer to doing to its English reader what the original does to a reader of the original? AnswerAorB, then one sentence of reason. There is no third answer.
The clause "written by someone who has read it in the original" was cut on critic finding 6,
which is small and exactly right: it is true of YP and false of YM, a restyle by a seat
that never saw the Japanese. A seat that reasons about the description's provenance would have
been misinformed in one arm of the manipulation and not the other — and the cut also strengthens
the byte-identity assertion in §4, since the sentence no longer says anything the two documents
differ on.
Total: 396 dispatches.
5. The registered predictions
P1 — the separation replicates (criterion 1)
Over segments with a majority (§4 rule) in both H1 and H2P, excluding any segment GJ marked
WRONG:
H1majority →R08andH2Pmajority →R06at the segment level;- discordant segments ≥ 6 (the power floor — below it the primary is withheld);
- exact two-sided binomial on the direction split of the discordant segments,
P< 0.05, with the excess toR06underH2P.
The null being tested is conditional, and critic finding 8 is right that the design should say
so: given that H1 and H2P disagree on a segment, the direction of the disagreement is a
coin flip. The binomial is computed on a set selected for disagreement, which is legitimate under
that null and only under it; this test is not, and may not later be read as, an unconditional
comparison of the two conditions' rates.
P1 HOLDS iff all three. Cell-level counts (segment × seat) and per-seat tables are reported
as description, never as the primary.
P2 — the confound invariance, as an equivalence test (criterion 4)
Cases are segments where the GS style match (§4 majority rule) differs between YP and YM
— YP matched to one arm and YM to the other — and which survive all four individual
exclusions: GJ ACCURATE, GY SAME, YM ORNATE, YP PLAIN. A co-movement is a case
where the H2 majority also differs between the two documents, in the direction the style match
moved.
The per-segment exclusions are critic findings 5 and 2 made operative. As first written, a
segment whose YM had drifted in content stayed eligible unless drift reached three segments —
so one or two content-drifted documents could manufacture or mask a co-movement and the run could
still discharge. Individual exclusion closes that; F2 survives as the whole-run void.
- Invariance margin δ = 0.50. Null hypothesis
H0: p ≥ 0.50, one-sided exact binomial; reject atP< 0.05. P2HOLDS (co-movement is equivalent to invariance within the margin) iffH0is rejected.- Power floor: if the number of cases
n< 11,P2is WITHHELD for lack of power and is not reported as either outcome. This is criterion 4's "must have registered power" made a rule rather than a hope.
The critical values, computed by exhaustive enumeration and recomputed by verify.py:
n |
reject iff co-movements ≤ | exact P at that count |
|---|---|---|
| 11 | 2 | 0.0327 |
| 12 | 2 | 0.0193 |
| 13 | 3 | 0.0461 |
| 14 | 3 | 0.0287 |
| 15 | 3 | 0.0176 |
Registered power. Against the confound hypothesis p = 1.0 the test rejects with probability
0 — it cannot discharge the condition if the confound is operating, which is the property
criterion 4 exists to demand. Against p = 0.15 (the co-movement rate RS-20260813b §6.1 recorded
without licence) power at n = 14 is 0.853; against p = 0.30 it is 0.355. The test is
therefore honest in both directions and is not a formality: a moderate confound defeats it.
On a discharge the result reports the observed co-movement rate and its exact one-sided upper
confidence bound beside criterion 4's wording — critic finding 7, which is arguable and was taken
anyway. At n = 14 a discharge means ≤ 3 co-movements, an observed rate ≤ 0.214 with a 95% upper
bound of 0.466 (Clopper-Pearson, exact); the phrase "equivalent to invariance within a margin of 0.50" licenses a reader
to imagine half the cases co-moved, and the number is stronger than the margin. The number is
reported; the margin is what was registered.
Failing to reject is not evidence for the confound, and the result page will say so in those
words. A high co-movement count is evidence for it, and if co-movements ≥ ⌈n/2⌉ the result
records that the style reading of limit 5 is supported, and RS-20260812f §5 and
RS-20260813b §5 acquire a caution heavy enough that neither direction may be cited without it.
P3 — the manipulation check (criterion 3)
Computed over valid segments only — segments with a GS match under both documents by the §4
rule, so that a segment whose seats could not be resolved for want of usable bodies is excluded
from the denominator rather than counted as a failure (note (bna), the defect that voided
E-20260813b):
- (a) at least 8 segments have a
GSmatch under both documents; - (b) the style match differs between
YPandYMat ≥ 0.60 of those segments; - (c)
YPis matched toR06at ≥ 0.80 of the segments whereYPhas a match.
P3 HOLDS iff all three. F3 fires if it does not, and P2 is void.
Registered here, in the design, because it is the way this run could most easily fool itself: no
override of F3 will be taken, whatever the withheld comparison shows. RS-20260812f §7 set a
precedent for overriding a failure criterion with the reason on the record; RS-20260813b refused
to use it because the override would have rescued that run's own headline, and the same refusal is
pre-committed here. A withheld comparison may be recorded, flagged at every use as unlicensed, and
may not appear in any conclusion.
6. Failure criteria
| fires when | consequence | |
|---|---|---|
F1 |
GA returns DIFFERENT on ≥ 3 segments, excluding the registered quoted-line carve-out |
the arms are not comparable; the whole run is void |
F2 |
GY returns CHANGED on ≥ 3 segments |
the restyle moved content; P2 void, P1 unaffected |
F2b |
GJ returns WRONG on ≥ 5 of 15 segments |
the plain yardstick does not describe the source; the whole run is void, P1 included, because H2P is built on YP |
F3 |
P3 fails any clause |
P2 void, reported as void and not as a null |
F4 |
for a segment: no usable document under either style · GJ WRONG · GY CHANGED · YM not ORNATE · YP not PLAIN |
that segment is excluded from P2; a GJ WRONG segment is excluded from P1 as well |
F5 |
dead bodies exceed 10% of dispatches on any stage | that stage is void |
F6 |
— | the truncation guard: any body that fails to parse, or whose text ends mid-token, is re-dispatched at double the cap irrespective of finish_reason (note (bmz)); cost is accumulated across attempts, never overwritten by the successful one (the RS-20260813b §8 runner defect, whose remedy is one line) |
7. The result → option map, registered in advance
P1 |
GD (criterion 2) |
P3 (criterion 3) |
P2 (criterion 4) |
what is recorded |
|---|---|---|---|---|
| HOLDS | satisfied | HOLDS | HOLDS | D-20260804-16 condition 3 DISCHARGES. All four criteria met. wiki/goodness-senses.md §affect records the discharge and criterion 5's residual caution verbatim: discharge defeats the style reading of limit 5 only; H2 inherently contains a document and H1 does not, and every citation of either run's direction carries that until a design outside this family addresses it. Usage rule 4 stands either way — it is not what condition 3 governs |
| HOLDS | any | HOLDS | co-movement ≥ ⌈n/2⌉ | The confound is supported. Condition 3 does not discharge; RS-20260812f §5 and RS-20260813b §5 are re-flagged so that neither direction may be cited without it. The arm closes on a positive finding against the project's own prior result |
| HOLDS | not satisfied | HOLDS | HOLDS | No discharge. The manipulation was not validated as directional, so what P2 measures is not the confound criterion 2 names. Reported in full, with the arm closing on the failure and the missing instrument named |
| HOLDS | any | FAILS | void | No discharge, F3, P2 void. The arm closes resolved with condition 3 open and the second failure of this manipulation on the record — two failures make the manipulation itself the finding |
| WITHHELD or fails | any | any | any | The replication did not reproduce on a third family. Condition 3 does not discharge, and RS-20260812f/RS-20260813b's generality is the finding rather than the confound |
In no cell does the arm continue past step 2. ARM-affect-confound's completion criterion is
that condition 3 ends up recorded one way or the other; an inconclusive outcome is a failure to
discharge and closes the arm, not a licence for step 3.
8. Pre-flight cost, built from caps and not from expected lengths
Note (abc): the one estimate this project has overrun was built from assumed output lengths.
Rates from config/models.md (in / out per M): P1 $1.00/$6.00, P2 $1.50/$7.50, P3
$2.00/$6.00 — judge average $1.50/$6.50; qwen/qwen3.7-max $1.475/$4.425.
| stage | calls | seat | in cap | out cap | worst case |
|---|---|---|---|---|---|
YP + YM |
30 | qwen | 700 | 1,600 | $0.2434 |
GY |
15 | qwen | 800 | 700 | $0.0642 |
GA |
15 | qwen | 500 | 700 | $0.0576 |
GJ |
15 | glm | 900 | 700 | $0.1110 |
GD |
96 | judges | 600 | 700 | $0.5232 |
GS |
90 | judges | 700 | 500 | $0.3870 |
H1 |
45 | judges | 700 | 600 | $0.2228 |
H2P + H2M |
90 | judges | 950 | 700 | $0.5378 |
| 396 | $2.1470 |
Plus a 25% re-dispatch allowance for the F6 guard: $2.684. Declared ceiling for step 2:
$2.80. The critic's remedies cost $0.71 of worst case, most of it the rebuilt GD stage, and
that is the right trade: the stage as first written could have let a discharge rest on fourteen
unvalidated documents. P2's cap is raised to 10,000 on any call it returns truncated, per §3.
The estimate is built from max_tokens, and max_tokens is not a guarantee (note (bgk)): billed
completion has exceeded a declared cap once in this project's history, on a routed provider.
A $2.80 ceiling is 56% of a UTC day's cap and several sessions may share a day. Step 2 checks
the day's rows before dispatching and, if the headroom is not there, splits at the stage
boundary — stages 1–4 (yardsticks, gates, GD, GS, worst case $1.39) are self-contained and
produce the manipulation check; stage 5 is the judging. Splitting is preferred to scaling down,
because every bar in §5 is registered against 15 segments.
Step 1 spent only the pre-run critic: $0.095881200 against a declared ceiling of $0.35 and $1.046484 of headroom on UTC day 2026-08-13.
9. What this design cannot establish, stated before it runs
- Model seats, not readers. Tier D is NOT PASSED. Nothing here is calibrated, and every figure is a fact about how three seats behave.
- Criterion 5 survives any discharge.
H2contains a document andH1does not. No design in this family removes that; a comparability question with no description of the original in it is not a comparability question. - One hand, one order. Both arms are the lead, in one session,
R06first.RS-20260813blimit 3 carries over unchanged: this replicates a lead-vs-lead contrast, not an independent one. The 12-token self-match in §2 is the size of the shared substrate. - Three things move at once — family, affect register, source register — so a replication generalises and localises nothing.
GSis an unanchored model judgment about style, andRS-20260813b§7 suggests it measures distance from plain English rather than kind of markedness.GDis this design's answer, andGDis itself an unanchored model judgment about style.YMis written by a model fromYP, not by a human from the Japanese. The restyle can only be as good as one seat's ornate English, andGYcan only catch content drift it can see.
10. Verification
analysis/verify.py, written after this page is frozen and importing nothing from tools/,
recomputes every reported number from the stored bodies and asserts at minimum:
- the 15 Japanese segments concatenate to the source character for character, and each arm's 15 segments concatenate to that arm's whitespace-normalised body;
- no
H1prompt contains a document; - no
GSorGDprompt contains the words original, Japanese, source, effect, comparable or translation; - the
H1,H2PandH2Mprompts for a given (segment, seat) are byte-identical outside the document block, and the label order matches the seed rule; - every individual cell choice against the stored body — not the majorities (the mutation that
escaped
E-20260813b's first verifier pass was a single flipped cell inside a 3–0 block); - the exact binomial
Pvalues forP1andP2by exhaustive enumeration, and the critical-value table in §5; - the cost total equals the sum over all attempts including discarded ones;
- every §4 majority is recomputed from the usable bodies, and every segment excluded by
F4is excluded for a reason present in the storedGJ/GY/GDbodies — theP2denominator is reconstructed from the raw record and compared to the reported one; - no
GJprompt contains either arm, and noGDprompt contains more than one item.
Mutation tests, ≥ 8, each caught or the verifier is repaired before any number is reported:
a flipped H2P cell inside a unanimous block; a GS match inverted on one segment; a document
swapped between YP and YM on one segment; a segment dropped from the P2 denominator; a
segment added to it that F4 excludes; a label order flipped for one (segment, seat) in H2M
only; a GJ WRONG verdict silently ignored; a discarded attempt's cost omitted.
11. Pre-run critic findings, and what was done with them
Seat P4 moonshotai/kimi-k3, no role in the run, one dispatch, finish_reason: stop,
provider Together, $0.095881200. Verdict PROCEED-WITH-AMENDMENT, eight findings, five BLOCKING.
All eight accepted; raw body run/critic-P4.json. No study call is dispatched until this
section is complete, and it is complete.
| # | severity | finding | what was done |
|---|---|---|---|
| 1 | BLOCKING | nothing validated YP against the Japanese — GA compares the arms, GY compares the documents, and a wrong YP would pass both and carry the whole test |
accepted in full and extended. New gate GJ (§4 stage 2) on a fourth seat, z-ai/glm-5.2, which wrote nothing; per-segment exclusion via F4; run-level void via F2b. Extended beyond the critic's remedy: the critic scoped exclusion to P2, and H2P is built on YP, so a WRONG segment leaves P1 too |
| 2 | BLOCKING | GD validated one segment's YM while P2 consumed fifteen, so a discharge could rest on fourteen unvalidated documents — the failure criterion 2 exists to prevent |
accepted. GD now runs on all 15 YM documents at three seats, with per-segment exclusion of any YM not ORNATE. +$0.25 of worst case |
| 3 | BLOCKING | the criterion-2 gate never checked that the plain document is plain; an ornate YP would leave the contrast not plain-vs-marked while criterion 2 formally passed |
accepted. GD also runs on all 15 YP documents, and on the R06 sample; criterion 2 now has four clauses, §4 stage 3 |
| 4 | BLOCKING | "two-seat agreement" is uncomputable with three seats and a forced binary — every non-failed cell has one — so P3's denominators and P2's n had no definition |
accepted, and it forced the design's one substantive choice. The §4 majority rule replaces it, stated once and propagated. The critic offered unanimity as the presumed reading; that branch was declined with the reason on the record — unanimity puts expected P2 cases near 9.6 against this design's own power floor of 11 — and the unanimity figures are reported as a stricter secondary |
| 5 | BLOCKING | content-drifted documents stayed eligible for P2 unless drift reached three segments, so one or two could manufacture or mask a co-movement and still discharge |
accepted. GY CHANGED now excludes the segment individually through F4; F2 survives as the whole-run void |
| 6 | MINOR | the H2 prompt said the description was "written by someone who has read it in the original" — true of YP, false of YM |
accepted. The clause is cut from both prompts. It also removes a sentence the two documents differed on, which strengthens the byte-identity assertion |
| 7 | MINOR (the critic marked it arguable) | δ = 0.50 in the discharge wording licenses a reader to imagine half the cases co-moved, when a discharge in fact bounds the rate near 0.21 | accepted. On a discharge the observed rate and its exact upper confidence bound are reported beside criterion 4's wording. The margin stays what was registered |
| 8 | MINOR | P1's binomial runs on a set selected for disagreement; the design should name the conditional null so it cannot later be read as an unconditional comparison |
accepted. The conditional null is stated in §5 P1, with the prohibition attached |
What the critic did not find, recorded so a later reader can see the gap this pass leaves. It raised nothing about the materials (its priority 4) — no objection to Meiji 擬古文 as the source, to the two arms' comparability, or to three things moving at once between this run and its two predecessors. That may be because there is nothing there, or because one segment of material was all it was shown. §9 limits 3 and 4 carry those risks unrefereed.
12. Amendments made at dispatch (step 2, 2026-08-14), before any study call
Written and committed before run.py existed and before one dispatch was made, so that a reader
can see which of these came before the numbers. Nothing here moves a bar, a threshold, a
denominator, a failure criterion or the result→option map of §7. Two are corrections to §10's
verifier checklist where it contradicts §4's operative text; two are transport protections.
A1 — §10 assertion 3's lexical ban is asserted on GS only. GD is asserted on §4 stage 3's own
properties instead. §4 stage 3's frozen GD prompt contains the words translation and
original inside the gloss of the SOURCEWARD option — "the direction of a translation that keeps
its original's shape" — which is the clause that makes the three-way choice mean anything, and
without which SOURCEWARD is an undefined label. §4 stage 4 scopes the lexical assertion to GS
("verify.py asserts that no GS prompt contains…"); §10 assertion 3 over-extended it to GD,
where the design's own prompt text cannot satisfy it. The prompt is not reworded — a frozen
instrument is not edited to make a checklist pass. What verify.py asserts for GD is the sentence
§4 stage 3 already states: the prompt does not say the passages are translations, does not mention
the original, and does not mention any other item — operationalised as: the prompt contains exactly
one prose item, contains no reference to that item's provenance, and contains no Japanese character.
A2 — §10 assertion 4's byte-identity is asserted where it is true. H1 asks a different
question from H2 by construction — that difference is the design — so the three prompts cannot
be byte-identical outside the document block, and §4's sentence saying they are is loose. What is
asserted, and what §4's purpose actually requires: (i) H2P and H2M for a given (segment, seat)
are byte-identical outside the document block; (ii) the two-passage block — the A and B
texts in their label order — is byte-identical across H1, H2P and H2M for that (segment,
seat); (iii) the label order matches the §4 seed rule. (i) is the pairing the confound test needs;
(ii) is what makes H1 and H2P comparable segment by segment.
A3 — reasoning effort is pinned low on P2 and on z-ai/glm-5.2, per method note (bnk), which
post-dates this freeze. The note fired on 2026-08-14 and cost $0.526018: on a reasoning seat
max_tokens caps the answer plus the hidden thinking, and the thinking takes all of it, leaving
an empty body that is billed. §3 already pins qwen/qwen3.7-max and already flags P2 for cap
inflation; this extends the same protection to P2 and to glm, whose reasoning behaviour is
unmeasured in this project. P1 and P3 are left unpinned, as in E-20260813b, where they
returned usable bodies at caps of 400–900. This is a protection on the transport; it changes no
prompt, bar or statistic, and it is recorded because it does change what the seats do.
A4 — the machine-readable reply line. §4 gives each prompt's text but no reply format. Each
prompt is dispatched verbatim with one appended line requiring a single-line JSON reply. The
appended lines are in run.py and are reproduced in the result; verify.py checks that the GS
appendix contains none of the six banned words, and that every prompt's §4 body is present
unmodified.
A6 — the transport, rebuilt on measurement after the first dispatched call breached its own cap
ladder. This is the largest amendment and it is written before the second call. yp__S01 was
dispatched to qwen/qwen3.7-max at the §8 cap of 1,600 with reasoning: {"effort": "low"} set,
as §3 requires. It came back empty; the F6 guard re-dispatched at 3,200, which came back
empty; the third attempt at 6,400 returned a body cut off mid-sentence with
finish_reason: "stop". Billed: $0.043586 for one unusable cell, against a §8 worst case of
$0.0074 for the whole call. The usage record says why — 4,662 reasoning tokens for 171 visible
ones. The pin was sent and ignored, exactly as note (bne) records for deepseek-v4-pro, now
on a second seat and a second provider. That body is preserved under runs/_discarded/ and its
cost is counted (§10.7). Three things follow, all measured before re-dispatch:
reasoning: {"enabled": false}on the two seats that judge nothing —qwen/qwen3.7-maxandz-ai/glm-5.2. Probed once each: 0 reasoning tokens, complete bodies, $0.001081 and $0.000341. This is §3's stated intent forqwen("effort low") executed on a provider that honours it, not a new intent.- Reasoning is NOT disabled on any judging seat. Whether a judge thinks before answering is
part of the instrument, and muting it on
P1,P2andP3would change whatGD,GS,H1andH2measure.P2'seffort: lowpin from A3 is withdrawn for the same reason — A3 was written before this was measured, and it reached further than its evidence. - Caps re-sized from measured reasoning use, not from assumed answer length — note (abc)'s
rule applied to the quantity that actually varies. Probed on the real prompts:
P10–403 reasoning tokens,P3326–430,P2up to 1,043 onGS. §8's caps of 500–700 were built for seats that do not think, so every judge stage would have truncated. Caps: judges 1,200,P22,200 (it reasons hardest and is the cheapest out),yp/ym1,200,gj/ga/gy500.
The declared ceiling changes, and the split §8 pre-registered is taken. The re-sized cap-built worst case is $3.65 against §8's $2.80 — the increase is entirely the judges' thinking room, and it is offset in part by A5's price falls. Today's UTC headroom is $3.56, which the full run does not fit. §8's own instruction is split at the stage boundary, in preference to scaling down, because every bar in §5 is registered against 15 segments, and that is what is done: stages 1–4 (worst case $2.26) are dispatched first, the gates and the manipulation check are computed, and stage 5 (worst case $1.39) is dispatched only if the run is not already void and the headroom is there. No bar, threshold, denominator or failure criterion moves. Expected actual from the probes is ≈$1.05 for all 396 calls; the ceiling is built from caps because that is the rule.
A5 — prices re-read from GET /api/v1/models at dispatch, as §3 requires for glm. Read
2026-08-14: z-ai/glm-5.2 $0.63 / $1.98 per M against the $2.00 / $8.00 the §8 estimate priced
it at, and google/gemini-3.6-flash (P2) $0.75 / $3.75 against config/models.md's
$1.50 / $7.50 — a halving on a judging seat. openai/gpt-5.6-terra $1.00 / $6.00,
x-ai/grok-4.5 $2.00 / $6.00 and qwen/qwen3.7-max $1.475 / $4.425 read back unchanged. Both
surprises are in the conservative direction, so the §8 worst case and the $2.80 declared ceiling
stand unrevised; config/models.md is corrected in place.