Repository path: workshop/experiments/E-20260809a-trajectory-ja/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260809a-trajectory-ja |
| status | frozen |
| created | 2026-08-09 |
| updated | 2026-08-09 |
| senses | style-correspondence, voice, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-trajectory.md, framework/v0.2/README.md, workshop/translations/kokoro/R06-v1/translation.md, workshop/experiments/E-20260808h-trajectory/design.md, wiki/findings/results/RS-20260808h-trajectory.md, config/models.md, config/budget.md |
E-20260809a — the second language pair, and the gate that withheld the first one, repaired
Frozen before dispatch. The translation limb it hangs on (T-kokoro-R06-v1, «こころ» 上・十三)
was written from the Japanese alone and committed at 9d4de2f, before this design existed.
1. The question, and what step 1 left owing
framework/v0.2 §8 Q-d: where a source marks a relation by a change of form across a text — a
switch of pronoun, an onset of honorifics, a drop of a title — does English carry the change, and
which devices carry it?
ARM-trajectory step 1 (E-20260808h, S139) put that question to Russian and the run withheld its
own answer: the source-side gate returned +0.667 against a registered +1.00, F1 fired, P2
was withheld. Note (bkv) names the reason and it is arithmetic, not design — the gate got 2
ratings per site where the primary got 6. Step 2 of the arm is written to fix exactly that and to
add the second pair.
So this design does two things, and the first one is a repair the arm declared in advance:
- Stage R — the Russian gate, re-measured at parity. S139's twelve MAIN sites, unchanged texts, the same question, the same bar of +1.00, now read on 6 ratings per site instead of 2.
- Stage J — the second pair, crossed on device. Japanese, where the relation between two speakers is marked in a place English does not have at all: the politeness inflection of every finite predicate.
2. What the translation limb contributed, and the wire
The lead translated «こころ» 上・十三 whole — the walk at Ueno in which Sensei, who has held the narrator at the courteous distance of あなた + です・ます for the length of their acquaintance, breaks twice: at ¶18 into plain forms and the familiar 君 («――君、黒い長い髪で縛られた時の 心持を知っていますか»), at ¶21 into unbroken plain forms («悪い事をした… あなたを焦慮していたの だ»), and then deliberately rebuilds the polite frame to close the scene («とにかく恋は罪悪ですよ、 よござんすか»).
The translator's log, frozen at 9d4de2f, records what English could and could not do with that.
Its D2 put the whole trajectory on contraction, sentence length and directive force, the only
exponents English has. Its D3 records that ¶18's 君 was rendered by adding two words that are not
in the Japanese ("see here, my boy") — converting a pronoun into a vocative, a change of device.
Its D4 records the other 君, at ¶23, and that nothing in the English marks it at all, because
there it is a subject pronoun and English deletes it.
The wire, one sentence. The translation limb generates the problem — a translator facing a relational change carried by predicate morphology and pronoun choice, neither of which the target language has — and the study limb asks which of the devices a translator might reach for instead is actually recoverable by a reader who is told nothing.
3. Contamination, and one thing this run does not have
T-kokoro-R06-v1 declares suspected and unmeasured, because no published English of «こころ»
is freely reachable (Gutenberg holds only Botchan and Kusamakura for Sōseki; McClellan 1958 and
McKinney 2010 are in copyright). The note (bhb) repository check was run before the section was
selected and returns zero: the project has no prior rendering of this novel.
Two consequences, both stated rather than smoothed. (i) The lead writes no arm of this experiment, as at S139 — the three hands are independent models — and §十三 is excluded from the site pool, so no hand renders the lead's passage. (ii) S139's stage 8, the published-translator census, has no analogue here. That census — Garnett letting Chekhov's mid-sentence correction go — was one of the most informative things in the previous result, and this run cannot do it. It is a real loss and it is the price of choosing the language pair over the comparator.
4. Materials
Source. 夏目漱石『こころ』(1914), Aozora Bunko 000148/773_14560, 底本『こころ』集英社文庫,
PD-old-70 (Sōseki d. 1916). materials/kokoro.txt, 1,331 paragraphs, 160,234 non-whitespace
characters, decoded cp932 and cleaned by tools/fetch_aozora.py.
Candidate pool, generated mechanically by pool.py, committed before any span was read for
content. A candidate is a maximal contiguous paragraph run within one section such that:
- 250–600 non-whitespace characters;
- ≥ 4 quoted utterances, containing ≥ 6 classifiable sentence-final predicates;
- every classifiable predicate is of one register — polite (です/ます family) or plain — so no span carrying an authorial register mix can enter the pool;
- ≥ 2 classifiable predicates in each character-count half;
- outside sections 上十二 / 上十三 / 上十四, which the lead read while choosing the translation locus.
The pool is 20 candidates — 16 POLITE, 4 PLAIN. The imbalance is a property of the novel: part 上 is the narrator with Sensei and his wife, and it is polite on both sides; the plain-register dialogue is in part 中, with the narrator's family. Widening the character window to 200–700 and to 150–900 was tried before any span was read and returned 4 PLAIN candidates both times, so the window stayed where it was written. The design is built around 4, and says below what that costs.
A second mechanical census, also before reading: address terms inside dialogue across the whole novel — あなた 96 · 先生 75 · 君 37 · お前 37 · 奥さん 17, against name-plus-suffix forms in the low single figures. This is why the positive control below is a title/pronoun contrast and not the name-form contrast S139 used: «こころ» does not have one.
5. Procedure
Stages run in the order given. Every dispatch runs in the background — note (bid) has fired three times on a shell timeout killing a call OpenRouter had already billed, and its corollary has been recommended three times and applied once.
snapshot open—GET /api/v1/key.- Stage 0 — two pre-run critics, on different labs (note (bhf): one critic pass is not
adversarial coverage when the seat can return nothing). C1
mistralai/mistral-medium-3-5; C2nvidia/nemotron-3-ultra-550b-a55bdispatched with"reasoning": {"enabled": false}— note (bkw), applied before the failure rather than after it, the first time this project has done so. Neither is a hand, a rater or a gate seat. - Stage R — the Russian re-gate. S139's sixteen Russian spans, byte-identical, read from
../E-20260808h-trajectory/run/sites.json. Three seats × two calls, sixteen items per call, one arm per site per call, alternating by call so that every site receives 3 ratings of arm F and 3 of arm S — 6 per site, the same as S139's English primary. Same prompt as S139 stage 6, sameOK/FLAGgrammaticality field, same bar. - Stage 1 — independent scene screen. All 20 candidates to one non-hand, non-rater seat
(
mistralai/mistral-medium-3-5), neutral prompt naming no device and no hypothesis: for each passage, how many people take part in the conversation (TWO/MORE-THAN-TWO/UNCLEAR), the name or description of the person spoken to, the name or description of the speaker. This removes the lead's discretion over the judgement the manipulation depends on. - Stage 2 — site selection, by a rule fixed here and applied to stage 1's output without further
judgement.
- Every
TWOPLAIN candidate becomes a MAIN site, up to 4. - Walk theTWOPOLITE candidates in pool-index order, assigning alternately to MAIN and to POS, MAIN first, until MAIN holds 4 POLITE sites and POS holds 4. A POLITE candidate is POS-eligible only if its tail carries ≥ 1 address token (先生 / 奥さん / あなた / 君 / お前); one that is not is passed to MAIN and the alternation continues. - Result: 12 sites — MAIN 8 (4 POLITE, 4 PLAIN), POS 4 (POLITE). If fewer than 4 PLAIN survive the screen the MAIN deficit is made up from POLITE and the imbalance is reported. - Stage 3 — the manipulation. For each site the split point is the paragraph boundary
nearest the character-count midpoint. Only text at or after the split is touched. In both blocks
arm F is the printed text and arm S is the manipulated tail — a simplification of S139, where
POS manufactured both arms because its tails often carried no vocative; here every POS tail
already carries an address token, so replacement suffices and the arm-F baseline is the same kind
of thing in both blocks.
- MAIN — the predicate-register trajectory. Arm S: every sentence-final predicate inside a
quoted utterance at or after the split is converted to the other register (です・ます ↔ plain),
with the agreement that conversion requires (copula, negation, tense, request forms, the
explanatory のです/んだ). No word outside a predicate is added, removed or changed, and no
address term is touched. Direction: POLITE sites move toward familiarity, PLAIN sites move
toward formality.
- POS — the address-term trajectory, the positive control. Arm S: every address token at or
after the split is moved between the title class (先生 / 奥さん) and the second-person
pronoun class (あなた / 君 / お前), the variant chosen to fit the dyad the screen named.
Predicates are left alone. Direction is 2 and 2: the first and third POS sites in index order
move toward familiarity (title → pronoun), the second and fourth toward formality (pronoun →
title); a site can serve a direction only if its tail carries a token of the source class, and
the rule falls through in index order until it can.
- Orientation. Every rating is oriented so that positive = in the direction the manipulation
went. This is what lets the two directions be added rather than cancelled, and what stops a
rater's standing tendency to read later text as cooler from counting as an effect. Imported
unchanged from S139 §5.4, whose A3 amendment defended it against the pre-run critic.
- Manipulated texts and a residual checker's output are committed to
run/sites.jsonbefore dispatch. Every hand edit is an ordered phrase rule inoverrides.jsonwhose find-string is a hard error if absent. - Stage 4 — the hands. Three, none the lead, none Anthropic: H1
deepseek/deepseek-v4-pro, H2openai/gpt-5.6-terra, H3x-ai/grok-4.5— the same three as S139, for comparability. Each receives all 12 spans in one call, numbered, in site order, every source paragraph numbered[site.paragraph]with one output line required per input line (S139's A2 repair, imported), and is asked only to translate. No hand sees both arms of any site: with three hands each site is rendered three times, split 2/1 between arms, the singleton arm alternating with the site's position and the singleton hand rotating (assign.py). - Stage 5 — the reading probe, the measure. Three seats, none a hand: R1
google/gemini-3.6-flash, R2qwen/qwen3.7-max, R3z-ai/glm-5.2. Two calls each, 12 items per call, one rendering per site per call, rendering index(r + 2k + i) mod 3at site i for rater r, call k — so no call contains two renderings of one span, every rendering is rated at least once, and arms mix within every call. Calls are stateless. Each item is printed asPART ONE/PART TWOat the true split. The prompt is S139's, verbatim, so that the two pairs are measured on one scale:Below are N short passages of English prose fiction. They are unrelated to one another. For each passage, consider how the speaker treats the person he or she is speaking to. Comparing the later part of the passage with the earlier part, does the way the speaker treats that person become more distant and formal, become warmer and more familiar, or stay about the same? Answer with exactly one of:
MUCH MORE DISTANT,SOMEWHAT MORE DISTANT,ABOUT THE SAME,SOMEWHAT MORE FAMILIAR,MUCH MORE FAMILIAR. Then quote up to twelve words from the passage that most influenced your answer, or writeNONE. Output exactly N lines and nothing else, each<number>|<answer>|<quote>.
Coding: MUCH MORE DISTANT +2 · SOMEWHAT MORE DISTANT +1 · ABOUT THE SAME 0 ·
SOMEWHAT MORE FAMILIAR −1 · MUCH MORE FAMILIAR −2, then oriented per §5.5.
8. Stage 6 — the Japanese source-side gate, at parity. Three seats × two calls, twelve Japanese
items per call, one arm per site per call, alternating by call: 6 ratings per site, 3 per arm —
the same precision as the English primary. This is the whole of note (bkv)'s repair. Same
question, same five-point scale, plus a per-item OK / FLAG:<what> field for anything
ungrammatical or unnatural in the Japanese.
The seats are the three hands (deepseek-v4-pro, gpt-5.6-terra, grok-4.5), as at S139 and
for the reason S139 gave: they read the Japanese source and never any English rendering, so no
seat judges its own output, and calls are stateless so nothing carries between stages. The three
remaining non-hand, non-rater seats are kimi-k3 and nemotron-3-ultra, which have returned
zero-content bodies in four of their last six load-bearing roles, and mistral-medium-3-5, which
config/models.md records as weakest of the panel on Japanese texture. Note (bkv)'s corollary
is that a gate is the wrong place to economise seats, and competence on the source language is the
competence this gate needs.
9. Stage 7 — mechanical census, $0. Over the post-split English of every rendering: contraction
rate, mean sentence length, honorific/address tokens (sir, Sensei, Mr, Mrs, madam,
my dear, my boy, old man), imperative count. Arm S against arm F, oriented per site.
10. snapshot close; reconcile per-response usage.cost against the key delta.
Judgment is not parallelised. No model rates its own output; the hands, the English raters, the screen, and the critics are four disjoint sets of seats, and the source gate overlaps the hands only on the source language.
6. Predictions, registered
P0 — the Japanese source-side gate. On the 8 MAIN sites, the oriented source-side mean of arm S minus arm F is ≥ +1.00, on 6 ratings per site.
P0R — the Russian re-gate, the repair. On S139's 12 MAIN sites, the oriented source-side mean of arm S minus arm F is ≥ +1.00, on 6 ratings per site. The bar is unchanged from S139; only the precision behind it changes. Registered before dispatch, and reported whatever it returns.
P1 — the instrument gate. On the 4 POS sites, the oriented English-side mean of arm S minus arm F is ≥ +0.75.
P2 — the primary, null-favouring. On the 8 MAIN sites, the oriented English-side mean of arm S minus arm F is ≤ +0.40.
P2 holds → three independent hands do not put a source's predicate-register trajectory into their English in any form a reader recovers, while the same readers recover an address-term change from the same hands on the same material. P2 fails → the change is carried.
Why the Japanese MAIN block is a stronger test than the Russian one was. Russian marks the relation in one pronoun paradigm; Japanese marks it on every finite predicate in every sentence, which in these tails is six to fourteen separate morphemes. If breadth of exponence in the source were enough to force a translator's hand, MAIN would carry here and did not have to carry there. A holding P2 is therefore evidence about kind, not merely about quantity.
On the control, and note (bkt). (bkt) requires a control that manipulates the same component of the grammar as the phenomenon. English has no politeness inflection, so a strictly exponent-matched control does not exist — S139's A5 made this argument for the Russian pronoun and it holds a fortiori here. The address-term contrast is the nearest thing inside the same system: the same relation, changed at the same point in the same passages, judged by the same seats on the same scale, realised through a device English possesses. Q-d names it explicitly — "a drop of a title" — so the control is not off to one side of the question, it is one of the question's three named cases.
P3 — descriptive, no prediction. Stage 7's mechanical census.
Registered in advance: what stage R licenses. If P0R ≥ +1.00, S139's withheld P2 (+0.134) is
released with the status released-on-repair — explicitly weaker than a primary released by a
gate registered blind, because the quantity was already visible when this repair was registered, and
the result page must say so wherever the figure appears. If P0R < +1.00, S139's P2 stays
withheld permanently, and the finding is that the Russian pronoun manipulation could not be shown
legible in the source even at three times the precision — which is an answer about the manipulation,
not about English.
7. Failure criteria, frozen
- F1 — P0 < +1.00 → the Japanese manipulation is not legible as a change of manner even to a reader of the Japanese, and P2 is withheld.
- F2 — P1 < +0.75 → the instrument cannot see a carried change of address at all, and P2 is withheld.
- F3 — the arm-F oriented mean on MAIN is itself ≥ +0.75 → the unmanipulated baseline is already being read as a change; P2 is withheld and the baseline is reported.
- F4 — a hand returns other than 12 numbered items, or renumbers, or drops a paragraph line → one re-dispatch; on a zero-content body note (bkw)'s reasoning switch is tried on the same seat before re-seating. A second failure drops the hand and every quantity is recomputed on two, with the loss reported.
- F5 — more than 2 unparseable lines in a rater call → one re-dispatch; a second failure drops that call and the coverage loss is reported.
- F6 — the Japanese gate returns
FLAGon more than 2 arm-S MAIN items → the manipulation is defective; those sites are dropped and P2 is recomputed, with the drop reported. - F7 — the residual checker finds any surviving predicate of the departing register in an arm-S MAIN tail, or any surviving address token of the departing class in an arm-S POS tail → that site is repaired before dispatch or dropped; the count is reported either way.
- No criterion is weakened after it fires. If a gate fires the number is still reported.
8. Pre-flight cost estimate
Built from max_tokens, not from expected length (note (abc)).
| stage | calls | model | cap | worst case |
|---|---|---|---|---|
| 0 critics | 2 | mistral-medium-3-5 · nemotron-3-ultra | 10,000 | $0.12 |
| R Russian re-gate | 6 | deepseek · gpt-5.6-terra · grok-4.5 | 4,000 | $0.10 |
| 1 screen | 1 | mistral-medium-3-5 | 5,000 | $0.04 |
| 4 hands | 3 | deepseek · gpt-5.6-terra · grok-4.5 | 16,000 | $0.30 |
| 5 English raters | 6 | gemini-3.6-flash ×2 · qwen3.7-max ×2 · glm-5.2 ×2 | 8,000 | $0.35 |
| 6 Japanese gate | 6 | deepseek · gpt-5.6-terra · grok-4.5 | 8,000 | $0.30 |
| re-dispatch reserve | 4 | — | — | $0.39 |
| declared worst case | $1.60 |
Today's UTC ledger (2026-08-09) stands at $0.00 of $5.00. The worst case fits with $3.40 to spare. Stages 7, the whole translation limb, the pool, the manipulation and all analysis cost $0.
9. What this cannot show
- Three 2026 models are not three human translators. A holding P2 is these hands did not carry it, not it is uncarriable — and unlike S139, this run has no published human hand at all to set beside them (§3).
- A small study cannot prove a null. P2's meaning comes entirely from P1 standing beside it, which is why F2 exists.
- Arm S of a MAIN site is not Sōseki. A manufactured register break may be less motivated than an authorial one, and an unmotivated one is one a hand might normalise away. The lead's translation of §十三 is the only authorial break in view and no hand renders it.
- The MAIN block's two directions rest on 4 sites each, and the POS block on 4 sites total. The pool gave 4 PLAIN candidates and no parameterisation gave more.
- Block is not fully crossed with register: MAIN is 4 POLITE + 4 PLAIN, POS is 4 POLITE. P2 and P1 are therefore also reported split by register, per S139's A3.
- The gate seats are the hands. They read only Japanese and calls are stateless, but they are not a fourth independent lab.
- One work, one author, one language pair added.
10. Amendments, after the pre-run critics and before dispatch
Both critic seats returned, and note (bkw) was applied before the failure for the first time in
this project. mistralai/mistral-medium-3-5 returned NEEDS-REDESIGN, 10 findings for
$0.033836; nvidia/nemotron-3-ultra-550b-a55b, dispatched with "reasoning": {"enabled": false},
returned NEEDS-REDESIGN, 10 findings for $0.011126 — the same seat that spent $0.046932 on
a zero-content body as the registered critic at S139. Two labs, two verdicts, and the two agree on
three of the four blocking issues. Both bodies are in run/.
A1 — the pool's binding parameter, relaxed before any site was selected
The frozen rule returned 4 plain-register candidates, and the first screen called two of them
MORE-THAN-TWO (correctly: they are three-party family scenes), leaving 2 — too few for a
direction-balanced MAIN block. MIN_PREDS was relaxed 6 → 5; MIN_UTTS was tested and is not
the binding parameter (relaxing it alone changes nothing). One parameter, one step, no span read for
content. Pool 20 → 27, plain-register 4 → 8. The superseded pool and its screen are preserved
at run/pool-v1.json and run/screen-v1.json.
A2 — the screen runs on three seats, majority rule (critic 2 finding 8). Accepted.
mistral-medium-3-5, gpt-5.6-terra, nemotron-3-ultra. A dispatch error is recorded rather than
smoothed: the third seat first dispatched was z-ai/glm-5.2, which is one of the three English
rater seats and should not have been a screen seat under §5's disjointness claim. It returned a
zero-content body ($0.039174, ledgered as waste) and was replaced by nemotron-3-ultra. Calls are
stateless, so no information could have passed from a screen call to a rating call, but the seat was
replaced regardless. nemotron's screen is degenerate — TWO on all 27 — and contributes no
discrimination; the majority is therefore effectively carried by the other two, and where they
disagree (2 candidates) the vote goes to TWO. Four candidates are excluded on a MORE-THAN-TWO
majority.
A3 — the Japanese gate is measured on six seats, and the independence objection is given teeth
Both critics flagged that the gate seats were the hands (critic 1 BLOCKING 1.1, critic 2 finding 3). Calls are stateless, so the "remembers its own translation" mechanism they propose does not exist; the real objection — shared weights, hence correlated insensitivity — stands, and the panel contains no three competent Japanese seats that are neither hands nor raters. So the gate runs on all six available seats, twelve calls, 12 ratings per site, 6 per arm:
- hand block —
deepseek-v4-pro,gpt-5.6-terra,grok-4.5; - independent block —
kimi-k3,nemotron-3-ultra,mistral-medium-3-5, the last two dispatched with reasoning disabled from the start per (bkw).
P0 is the pooled figure, and the independence objection is registered as a veto: if the two
sub-blocks fall on opposite sides of the +1.00 bar, P2 is withheld whatever the pooled number
says. That converts the objection from something the design argues away into something that can stop
the run.
A4 — the residual check becomes a positive check (critic 2 finding 5). Accepted, and it fired.
The check no longer asks "has any predicate of the departing register survived" but "is every
classifiable predicate in the arm-S tail of the arriving register". It failed immediately on site
24, which is A4b: pool.py's polite pattern cannot see a polite marker followed by から/けれど/が,
so 「横着な了簡ですからね」 reads as plain. The pool stays frozen as it ran; the residual check uses
the corrected pattern. The correction also reclassifies one selected site — 9, 上十七 ¶1–7, whose
tail carries a polite 「叱り付けられそうですから」 — so that span is not register-uniform and never
satisfied the pool's own criterion. It is dropped from the selection, not manipulated. Dropping
it costs a site and is the conservative direction: an impure site adds noise, and noise makes a
null-favouring P2 easier to hold.
A5 — お前 is excluded from conversion toward politeness (critic 2 finding 2). Accepted.
The critic is right that Japanese register normally co-varies across predicate and pronoun, and that お前 + です・ます is pragmatically defective. あなた and 君 are attested with both registers in Sōseki's own hand (§十三 ¶18 君 + polite, ¶21 あなた + plain), which is the evidence that those two are compatible; お前 is not. A candidate whose tail carries お前 and is converting toward politeness is skipped mechanically. One site (25) was excluded by this rule.
A6 — a minimum manipulable unit (critic 1 ADVISORY 3.3). Accepted.
A site must carry ≥ 2 quoted utterances in the tail, ≥ 1 in the head, and ≥ 4 paragraphs, so the manipulation is never a single token and neither half is a fragment. Three candidates excluded.
A7 — the positive control is rebuilt, because the material refused the first one
The frozen POS design swapped a title for a second-person pronoun in the tail. Applied to the selected sites it collapsed: most 先生 tokens in these tails are third-person reference, not address (the narrator talking to Sensei's wife about Sensei), and where the tail's address token is a pronoun it is Sensei's あなた/君 to the narrator, for whom no title exists. Only 2 of 6 POS sites could carry the manipulation at all, and only in one direction.
Replaced before any arm was rendered, with S139's own POS structure: a courtesy formula
prepended to the first tail utterance — formal 「失礼ですが、」 against familiar 「なあ、」 — identically
placed, both arms manufactured, differing in exactly one constituent, which flip.py asserts by
reconstructing arm F from arm S. Direction is 3 formal / 3 familiar. Lexical politeness is the
nearest device to predicate politeness that English possesses, which is what the control is for; that
POS arm F is therefore not the printed text is the same concession S139 made, for the same reason,
and F3 reads on MAIN only.
A8 — what P1 licenses is narrowed in writing (critic 2 finding 1, critic 1 ADVISORY 1.3)
The two blocks differ in exponence density: MAIN converts every finite predicate in the tail (2–8 morphemes per site), POS inserts one constituent. P1 passing therefore shows that the instrument detects a relational change carried by a device English has — not that it is sensitive to a change of MAIN's density. The per-site manipulated-constituent counts are reported beside every figure. This is registered as a limit, not answered.
A9 — hand assignment balanced (critic 1 BLOCKING 4.1). Accepted, and it came out exact.
assign.py searches the singleton-hand choice at each site to minimise the spread of arm-S and
plain-register counts across hands. Achieved: 6 F / 6 S for every hand, 1 plain-register arm S
each — spread (0, 0).
A10 — stage R's denominator (critic 2 finding 7). Clarified.
All 16 Russian spans are shown in every call, because the prompt is S139's and is asserted
byte-for-byte identical to the stored G1 request before dispatch (regate_ru.py refuses to
build a spec otherwise). P0R is computed on the 12 MAIN sites only, exactly as S139's P0 was.
A11 — overruled, with reasons
- Critic 1 BLOCKING 2.2 (P1 ≥ +0.75 against P2 ≤ +0.40 is a "logical inconsistency"): the bars are already nested in the direction the critic asks for, 0.75 > 0.40.
- Critic 1 BLOCKING 1.2 (raters not validated for sensitivity): that is what P1 is. A separate
pilot would be the same measurement at extra cost, and
F2already withholds the primary if it fails. - Critic 1 BLOCKING 2.1 (the +1.00 bar is easier for Japanese, so lower it): the bar must be the same in both pairs or the two are not comparable, and lowering a gate before dispatch on the speculation that it will be cleared too easily is the one move a gate may not make — S139's A5, in the same words. That Japanese should clear it more readily is the design's stated expectation and is why a holding P2 would mean something.
- Critic 2 finding 4 (POS needs 8 sites): accepted in part — POS went from 4 to 6. The pool does not support 8 without taking sites from MAIN, which is already at 6.
- Critic 1 BLOCKING 4.2 (cap rater calls at 10 items): 12 items is below S139's 16, where all raters returned complete bodies.
A12 — deviations at dispatch, recorded rather than smoothed
- Stage R fired note (bhf) at the cap that worked at S139 on the identical prompt. Three of six
calls returned dead or truncated at
max_tokens4,000 —deepseektwice with zero content,gpt-5.6-terratruncated at 129 characters — although the byte-identical prompt at the same cap returned clean bodies from both seats at S139. Per (bhf)'s own distinction (raise the cap for the seat that has performed the role cleanly, change the seat that has not) the cap was raised to 12,000 and the seats kept. Two of the three recovered. The third,deepseekcall 1, died again at 12,000 and was recovered by (bkw)'s reasoning switch on the same seat for $0.0010549 — a cap raise that failed and a reasoning switch that worked, on one seat, in one stage. - The utterance counter treats a quoted word (e.g. 「先生」と呼び掛けた) as an utterance. It affects the A6 counts only, never register classification, and is reported.