Repository path: workshop/experiments/E-20260731-anchor-second-read/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260731-anchor-second-read |
| status | frozen |
| created | 2026-07-31 |
| updated | 2026-07-31 |
| senses | style-correspondence, voice, accuracy |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-anchor-second-read.md, wiki/base/anchors/A-sasaki-kuroneko/A-sasaki-kuroneko.md, wiki/base/anchors/A-chekhov-pari/A-chekhov-pari.md, wiki/base/anchors/A-beowulf-ingeld/A-beowulf-ingeld.md, workshop/translations/black-cat/R04-v2/translation.md, workshop/experiments/E-20260731-anchor-second-read/census.md, wiki/findings/theory/TH-20260724-translation-distance-axes.md, config/models.md |
Design — the anchor shelf's second read, and whether its most load-bearing rate is about a language or about a translator
Frozen before any API call. ARM-anchor-second-read step 1, T4. Arm A ran before this document
existed and is reported here as executed; Arm B is the part this freeze binds.
0. What changed the arm's premise before the design was written
ARM-anchor-second-read names A-shaw-spider-thread, A-garnett-vanka and A-chekhov-pari as the
three unaudited anchors to take, and lists eight anchors as "not audited, and not in any closed arm's
criterion." Two of the three named, and three of the eight listed, were audited at S015.
RS-20260725-anchor-verification audited A-shaw-spider-thread (50 quoted correspondences, all
attesting), A-garnett-vanka (85, all attesting) and A-baudelaire-chat-noir, string by string, with
a blind arm and a decoy-controlled adjudication arm; it retracted one claim and corrected three. The
arm's own table says so two lines above the list that contradicts it.
The genuinely unaudited set is therefore: A-chekhov-pari, A-sasaki-kuroneko, A-mchugh-presence,
A-mansfield-garden-party, A-doctorow-little-brother, and A-beowulf-ingeld outside §4.3. The
last three are condition-2 material (feature catalogues with no quantitative claims), which is the
arm's step 2. Step 1's three are A-chekhov-pari, A-sasaki-kuroneko and A-beowulf-ingeld
(non-§4.3), and the arm page is corrected rather than followed.
1. Question
Arm A (executed, free). Do the three anchors' quoted correspondences and count assertions survive a deterministic second read? — the denominator the arm was constituted to produce.
Arm B (this freeze). A-sasaki-kuroneko §2 reports that Japanese repays English's lost masculine
pronoun at 5 of 15 sites, 33%, and reads that as C1 holding with English as the source:
The category was repaid, and the repayment is intermittent: 5 of 15 sites, 33%. … So the shape C1 describes is present, and present in the direction the theory page said it had no evidence for.
Is 33% a property of the language pair, or a property of 佐々木? The count is verified — Arm A recomputes every figure in that section exactly. What is untested is the inference from a rate measured on one translator to a claim about a pair of languages. S067 and S068 found the same inference failing twice in the framework — a coverage rate and a decision rate that turned out to be measuring the translator's own option list rather than the thing they were quoted for. This asks whether the evidence base has the same defect.
2. Materials
- Source. Poe, "The Black Cat", body paragraphs 4–9, 822 words, from the stored file. Contains
eleven of the thirteen Pluto masculine-pronoun sites and two generic-human ones (census, frozen
at
4cbea57). The other two Pluto sites and both second-cat sites are in ¶14, outside this span and already rendered by both 佐々木 andT-black-cat-R04-v1. - 佐々木's Japanese, blocks 3–8, already stored. Overt 彼 at these eleven sites: 4.
T-black-cat-R04-v2, the lead's rendering, frozen at52c5887with its log, contamination measured at4dbe561. Overt 彼 at these eleven sites: 0. Declared primed toward agreement (the lead knew the aggregate 5-of-15 before translating), so its zero is measured against its bias.- Three independent non-lead renderings, produced by this arm.
3. Procedure
Renderers: P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5, P5 deepseek/deepseek-v4-pro.
P2 google/gemini-3.6-flash is the declared reserve for any seat that fails or returns non-Japanese.
P4 moonshotai/kimi-k3 is the pre-run critic and may not be a renderer — the S053 role-collision fix,
twelfth session running.
Each renderer receives the 822 English words and one instruction: render into modern literary Japanese
prose, one Japanese paragraph per English paragraph. The prompt does not mention pronouns, 彼,
佐々木, the anchor, this project, or any comparison. max_tokens 4,000, temperature default, one
call per seat, raw bodies stored under runs/.
Scoring is by script, then hand-adjudicated and printed. For each rendering, at each of the
thirteen sites in the span (eleven cat, two human), the scorer records whether the corresponding
Japanese clause contains an overt 彼 referring to that site's referent. Adjudication rules, frozen:
彼女,彼ら,彼方are not counted. A彼whose referent is the narrator's wife, the pets as a group, or a generic person at sites 13–14 is counted only in the human column.- A rendering that merges or splits paragraphs is aligned by sentence; if a site's clause cannot be
located, it is scored
UNLOCATABLEand named, never scored zero. - A rendering that is not Japanese (< 90% of characters CJK/kana after stripping punctuation) is rejected and the reserve seat is dispatched.
4. Predictions, frozen
- P-1. The three non-lead renderers do not all return the same cat-site 彼 count. Fails if all three counts are identical — which would be evidence that the rate is fixed by the language.
- P-2. At least one non-lead renderer returns 0 overt 彼 at the eleven cat sites.
- P-3. For every renderer, the human-site rate is greater than or equal to the cat-site rate. Fails if any renderer uses 彼 for the cat more readily than for the generic man — which would mean animacy is not what the anchor thinks it is.
- P-4 (memory control). No non-lead rendering's longest common character run against 佐々木's
blocks 3–8 exceeds the lead's 17, measured by
tools/dependence_check_cjk.pywith the same reference cells. If one does, that seat's pronoun choices are confounded with recall and are reported separately rather than pooled.
What the arm can conclude and what it cannot. With four independent renderings plus 佐々木 the arm can say whether 33% is reproducible. It cannot say what a human Japanese translator would do, and it cannot rank any rendering for quality — Tier D has not passed and no quality judgement is elicited anywhere in this design. Charter §4's bar on treating AI convergence as validation applies: the arm's usable direction is falsifying the claim that the rate is fixed, not confirming it.
Failure criterion for the arm as a whole. If two or more seats are rejected as non-Japanese or
UNLOCATABLE at more than three sites, the arm reports that the instrument did not work and returns
no rate.
5. Budget
Pre-flight, built from max_tokens (note (abc)), at list prices in config/models.md, with P5 priced
at the worst plausible provider (the S022 routing caution, ~$3.30/M out):
| call | max_tokens | worst case |
|---|---|---|
| critic P4 (in ≈ 3,500) | 12,000 | $0.191 |
| renderer P1 (in ≈ 1,300) | 4,000 | $0.032 |
| renderer P3 | 4,000 | $0.027 |
| renderer P5 | 4,000 | $0.015 |
| reserve P2 (if it fires) | 4,000 | $0.032 |
| total | $0.297 |
UTC day 2026-07-31 opens with the full $5.00. A run that does not fit is split or deferred.
6. Verification
analysis/verify.py, importing nothing from analysis/, recomputes: every stage-1 verdict from the
anchor pages and stored texts; every count assertion; every site score from the stored raw bodies;
the cost sum from the raw bodies; and at least two mutation tests that must be caught.
Amendment A1 — 2026-07-31, after the pre-run critic. All six findings accepted.
Critic: qwen/qwen3.7-max, one call, in 5,125 / out 7,870, stop, 143.7 s, provider Alibaba,
$0.042384125. Verdict NEEDS-REDESIGN, six findings: two BLOCKING, two MANDATORY, two
ADVISORY. All six accepted; none declined. Raw at runs/critic.raw.
The seat is not P4, and why. moonshotai/kimi-k3 was dispatched first, per the S053 role-collision
fix (the critic may not be a renderer, and P2 is the renderers' reserve). It held the socket open
past twenty minutes without returning and was abandoned — note (b), and note (beh) exactly:
a fall-through chain protects against a call that returns a failure and does nothing against one
that hangs. qwen/qwen3.7-max is probed-but-not-selected (config/models.md), is not a renderer, and
is not the renderers' reserve, so the role-collision fix holds. No billing for the abandoned launch
is visible in the key delta; it is declared here in case it settles later.
What the critic broke, and it is the design's own question
Findings 1 and 2, both BLOCKING, and they are right. Language models rendering English into
Japanese are not a stand-in for an independent translator for this feature, because 彼 for an
animal is precisely the Meiji–Shōwa 翻訳調 calque their pretraining is saturated with: "if the LLMs
return 33%, it does not mean the language pair demands it; it means the historical corpus they
memorized used it 33% of the time. If they return 0%, it may just mean modern RLHF has penalized
translationese." And the design's dichotomy is false in the other direction too: "if the
English–Japanese pair inherently permits multiple valid referent-tracking strategies … high variance
among translators is actually a property of the language pair's solution space, not just individual
whim."
The question "is 33% a property of the pair or of 佐々木" is therefore WITHDRAWN, before the run, and is not answered anywhere below. What replaces it is two things the critic's own finding 3 points at, one of which needs no model at all.
Amended Arm B-0 (free, no API) — the three-way split the anchor asserts and never counted
Finding 3, MANDATORY: "In Japanese, repeating the noun … is a standard, highly frequent strategy for maintaining referent tracking where English uses a pronoun. If a renderer uses lexical repetition, it scores 0 for 彼, falsely categorizing it as a zero-pronoun (null) choice."
A-sasaki-kuroneko §2 asserts, in prose and without a count, that "the other ten sites take a bare
noun (その猫, その動物, かわいそうな動物) or nothing at all." Nobody has counted that split —
not for 佐々木, not for anyone — and the distinction inside it is the anchor's own theory. C1's
clause is "meaning carried by grammar is not translated, it is transcoded into lexis": a lexical
noun at a pronoun site is transcoding into lexis. A zero is not. So a rate that lumps them
together is not a repayment rate.
Arm B-0 classifies every one of the thirteen sites, for 佐々木 and for the lead, into
KARE / NOUN / ZERO, prints the clause it classified, and reports the three-way split. It costs
nothing and it is the arm's primary result.
Amended Arm B-1 (the API arm) — narrowed to the one contrast the confound does not reach
P-1 and P-2 are retained and demoted: they are reported as whether any independent renderer reproduces an overt-pronoun rate near 33%, explicitly labelled as a measurement of model corpus priors, not of the language pair, in the critic's own words, quoted on the result page.
P-3 is promoted to the arm's primary API question, because it is WITHIN-renderer and a shared corpus
prior affects both of its arms equally. Sites 13–14 are a generic human he; sites 2–12 are a cat.
If a renderer writes 彼 for the man and not for the cat, that is about animacy inside one text by one
renderer — not about whether that renderer remembers 佐々木. This is the contrast the anchor's reading
needs and the one the confound does not reach.
P-4 is moved out of the predictions and into the procedure as a quality-control filter (finding 5, ADVISORY, accepted), and its scope is narrowed (finding 4, MANDATORY, accepted): it bounds verbatim character copying and nothing else. It cannot catch structural copying of 佐々木's pronoun-drop pattern with different vocabulary, and it cannot catch translationese priors acquired without copying 佐々木 at all. The dependency-parse check the critic asks for is not reachable here — the environment is Python standard library only, with no Japanese parser — so the limitation is declared rather than repaired, and no independence claim rests on P-4.
Finding 6, ADVISORY, accepted. Every claim about "the language pair" is withdrawn from this arm. Four renderings do not describe English→Japanese. What they can do is show that a number the shelf quotes as a finding is not reproduced by anyone, which is a statement about the number's robustness and not about Japanese.