Repository path: workshop/experiments/E-20260729f-inheritance-census/design/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260729f-inheritance-census |
| status | frozen |
| created | 2026-07-29 |
| updated | 2026-07-29 |
| senses | cultural-mediation, style-correspondence, accuracy |
| provisional | true |
| internal-judgment-only | true |
| links | wiki/arms/ARM-evidence-audit.md, wiki/base/anchors/A-yosano-yomogiu/A-yosano-yomogiu.md, wiki/findings/results/RS-20260729-drift-window-verify.md, workshop/translations/malory-worship/R04-v1/translation.md, workshop/translations/malory-worship/census-frozen.md, config/models.md, config/budget.md |
E-20260729f-inheritance-census — is a culture-bound-item census a set?
ARM-evidence-audit step 2. Frozen before any call is dispatched. Amendments made after the
pre-run critic pass are appended in §10 and are the only changes permitted after this point.
1. The question
A-yosano-yomogiu §2 makes the project's sharpest claim about intralingual translation:
given the identical free option, Yosano substitutes at five sites and Shibuya at one. Two translators of the same language, offered the same costless copy, choose oppositely.
Two things have to hold for that sentence to mean what it says. The sites have to be a set — a list a second reader would also produce — and the calls at each site have to be right. This experiment tests both, and declares in advance which of the two it expects to fail.
2. Why §2 and not §1 — the arm's own decision rule, discharged
ARM-evidence-audit step 2 says the choice between the うるはし thread (§1) and the
cultural-mediation substitution count (§2) is decided by what step 1 learned about rater
reliability. Step 1 (RS-20260729-drift-window-verify) learned three things:
| what was measured | result |
|---|---|
| cell-level factual scoring, with an attesting quotation required | reproduces — 0.934 pooled, 0.978 on a frozen site list |
| classification of items into classes | does not — four-class α 0.654 |
| the census — which items belong in the set at all | does not — Jaccard 0.4375, and step 1's own conclusion was "the anchor, whose rule was never written down at all, has no reason to do better" |
§1's site list is mechanically determinate — the string うるはし occurs three times in 951
characters and a second reader cannot get a different three. It is therefore already on the
reproducible side of that line, and it is audited here without any API call, by exact string
match (Arm A). §2's twelve-row table is a census drawn by one reader after reading all three
renderings, with no inclusion rule written down anywhere. That is the arm's answer: the money goes
where step 1 says the risk is.
3. Materials
| # | material | state |
|---|---|---|
| M1 | yomogiu-murasaki-classical.txt — the classical passage, 951 non-whitespace characters excluding the two section headers |
stored, verified |
| M2 | yomogiu-yosano-1939.txt — Yosano 1938–39, the whole chapter, 12,868 non-whitespace characters |
stored, verified |
| M3 | Shibuya Eiichi's 現代語訳 | copyrighted, not stored, not recoverable — see §7 |
| M4 | workshop/translations/malory-worship/source-me.txt — Malory XVIII.xxiv, 821 ME words |
stored, frozen 6d79d35 |
| M5 | the lead's frozen 20-item census on M4, with its written inclusion rule | frozen 6d79d35, before the translation was drafted |
4. The inclusion rule
Given verbatim to every censuser, and identical for both passages except for the naming of the source culture:
A culture-bound item is a common noun or fixed noun phrase occurring in the passage that denotes a thing, material, practice, office, social role, or named work belonging specifically to the source culture and period, and whose referent a non-specialist reader of the target would not reliably identify. Proper names of persons and places are excluded. Each item is listed once, by its form in the source, however many times it occurs.
This rule is a reconstruction. A-yosano-yomogiu §2 never states one; the rule above is
inferred from the twelve items its table actually contains. That is declared rather than hidden,
and it is conservative in the direction that matters: a written rule should reproduce better
than the unwritten one the anchor used, so a low agreement figure here is an upper bound on the
anchor's.
5. Arms
Arm A — mechanical attestation. No API call.
Every Japanese string quoted on the anchor page, matched against M1 and M2 by exact string match
after whitespace normalisation and removal of markdown emphasis. The 951-character figure. The count
of うるはし in M1. The count of the modern reflex 麗 in M2. Which of the anchor's twelve §2 items
occur inside M1.
DECLARED, because it changes how Arm A's predictions may be read: Arm A was run before this design was frozen. The 34 bracketed quotations, the 麗 count, the 951 figure and the presence/absence of the twelve §2 items were all seen first. Predictions P1 and P2 are therefore post hoc and are labelled so wherever they appear. Nothing about Arms B or C was seen before the freeze, and P3–P8 are prior.
Arm B — census reproducibility. 4 calls.
Two models independently enumerate culture-bound items, given the rule of §4 verbatim and shown no translation of any kind:
- B1 — on M1 (Genji, classical Japanese), P1 and P3.
- B2 — on M4 (Malory, Middle English), P1 and P3.
The positive control, in the same call and after the main list, separately tagged: a narrow census with a near-determinate answer — on M1, every plant name in the passage; on M4, every number expressed in words. Ordering is main-list-first, declared, and the direction of any contamination is conservative: the main list is produced before the narrow one exists.
No cap is placed on the number of items returned. A cap would corrupt the quantity being measured.
Arm C — cell scoring on the Yosano half. 2 calls.
The scoring set is built mechanically as the union of (the anchor's §2 items that occur in M1) and (every item returned by either B1 census). P1 and P3 are each given M1, M2 whole, and the item list, and score each item COPY / SUBSTITUTE / OMIT:
- COPY — Yosano's text contains the item's own written form (a prefix or suffix attached to it still counts as a copy: 紙屋紙 → 古紙屋紙 is a copy).
- SUBSTITUTE — Yosano renders the referent with a different word.
- OMIT — Yosano's text does not render the referent at all.
Every cell requires an attesting quotation from M2, or from M1 for OMIT. This is the factual-adjudication form step 1 measured at 0.934 and is the only part of this design inside the S015 licence (§8).
Arm D — the Shibuya column. No call, and no attempt.
Declared unrecoverable. See §7.
6. Registered predictions
| # | prediction | prior? |
|---|---|---|
| P1 | every Japanese string on the anchor quoted from M1 or M2 attests by exact match | post hoc |
| P2 | the anchor's "42 of 42 attest ... against the stored files" cannot be true as written for its Shibuya quotations, because Shibuya is not a stored file | post hoc |
| P3 | on M1, Jaccard(lead's in-passage census, each independent census) < 0.60 — step 1's own registered threshold | prior |
| P4 | the fork. Jaccard(P1's census, P3's census) > max(Jaccard(lead, P1), Jaccard(lead, P3)) on M1. If it holds, the lead's census is idiosyncratic. If it fails, "culture-bound item" does not define a set for anyone, which is the stronger and worse finding | prior |
| P5 | the positive control. On the narrow census, Jaccard(P1, P3) ≥ 0.75 on both passages | prior |
| P6 | Arm C cell agreement between the two scorers ≥ 0.85 | prior |
| P7 | the Yosano substitution rate recomputed over the union set differs from the anchor's 5/12 = 0.417 by more than 0.10 | prior |
| P8 | on M4, Jaccard(lead's frozen census, each independent census) < 0.60 — i.e. writing the rule down first, and freezing it before translating, does not rescue the census | prior |
7. What is not attempted, and why
The Shibuya column cannot be audited and will not be. His 現代語訳 is copyrighted and this
project does not store copyrighted text whole (charter §7 / copyright hygiene). No rater can be
given it, so no cell in it can be re-scored, and the anchor's "Shibuya at one" is permanently a
single reader's unverifiable observation. ARM-evidence-audit named this obstacle at birth as a
known one rather than a discovery. Two things follow and both are stated on the record rather than
worked around:
- The anchor's headline comparison is between an auditable column and an unauditable one. Half of "five versus one" can be checked and half cannot, ever, by anyone.
- Re-fetching the page to check a cell is not a way round this. It would be a fresh
consultation of a copyrighted text, ledgerable in
wiki/base/consulted.md, and it would leave the finding exactly as unverifiable for the next reader.
8. What licenses the calls
RS-20260725-anchor-verification (S015): on factual adjudication all four panel members used
scored 0.886–0.917 against planted false claims. Arm C is factual — is this string present,
does this word appear in that form — and is inside that licence.
Arm B is not a factual task. It is a classification, and ARM-evidence-audit's own constraint
says "every step of this arm must stay on the factual side of that line." Arm B is admissible only
because it adjudicates nothing: it measures whether two readers produce the same set, and every
number it yields is a statement about the instrument. No output of Arm B is evidence that any item
is or is not culture-bound, and no correction to the anchor may rest on one (F4).
Tier D has not passed. No quality judgment is made, sought, or licensed anywhere in this design.
9. Failure criteria — registered, and they withhold rather than explain
| # | criterion |
|---|---|
| F1 | the positive control gate. If P5 fails on a passage, no census figure from that passage is reported as a measurement of the anchor; it is reported as a measurement of the instrument only, and the withheld figure is printed with the reason |
| F2 | any Arm C cell whose attesting quotation does not match the stored file by exact string match after whitespace normalisation is discarded. If more than 10% of cells are discarded, the whole scoring arm is withheld |
| F3 | a census returning fewer than 5 or more than 60 items on a passage is declared degenerate and excluded, with its number printed |
| F4 | no correction is made to the anchor on the strength of a model output alone. Every correction must be reproducible by exact string match against a stored file, checked by hand, and re-checked by verify.py. This is step 1's rule and it is what made its three corrections stick |
| F5 | the Shibuya column may not be corrected, defended, or scored |
| F6 | the lead does not judge T-malory-worship-R04-v1, and no arm here scores it |
10. Models, roles and cost
| role | model | why |
|---|---|---|
| pre-run critic | P4 moonshotai/kimi-k3, max_tokens 16,000 |
a subject in nothing here. P1 and P3 are the censusers and scorers, and a reader may not critique the instrument it is about to be. This is the declared fix for step 1's role collision, in which the critic's own model produced the alternative census. max_tokens 16,000 per note (bdl) — the cap is the remedy for note (b) |
| censuser ×2, scorer ×2 | P1 openai/gpt-5.6-terra, P3 x-ai/grok-4.5 |
the same two as step 1, deliberately, so the census figures are directly comparable to Jaccard 0.4375 |
| reserve | P2 google/gemini-3.6-flash |
if it fires, the rejected body is written to a separately named file and verify.py reads the accepted one — note (bdt) |
Pre-flight estimate, built from the max_tokens cap the request actually permits (note (abc)):
| call | cap | worst case |
|---|---|---|
| critic (P4) | 16,000 out + ~10k in | $0.27 |
| census ×2 (P1) | 8,000 out each | $0.25 |
| census ×2 (P3) | 8,000 out each | $0.11 |
| scoring (P1) | 12,000 out | $0.22 |
| scoring (P3) | 12,000 out | $0.11 |
| total | ≈ $0.96 |
Against $3.551884 of headroom on UTC day 2026-07-29 after five sessions. Routing caution (S022): billed can exceed list by ~4× on some slugs; P1 and P3 have honoured list for seventeen consecutive runs, P4 routed to Fireworks at 35% of its cap in S057.
11. Amendments after the pre-run critic pass
Critic: P4 moonshotai/kimi-k3, provider Fireworks, stop, in 5,852 / out 2,719, $0.0875115,
max_tokens 16,000 (35% of cap). Verdict NEEDS-AMENDMENT, eleven findings — five BLOCKING, four
MANDATORY, two ADVISORY. All eleven accepted; one sub-clause of finding 8 corrected on a fact.
Nothing above this line is edited; every change is stated here and the dispatched prompts are
rebuilt from it. Note (rr), sixteenth consecutive session.
A1 (finding 1, BLOCKING) — a matching rule, frozen before dispatch
Jaccard is a set operation on strings, and without a stated matching rule the rule chooses the result. Frozen:
- Both lists NFKC-normalised;
「」『』・, whitespace, markdown emphasis and the honorific prefixes御/おstripped. English items lowercased, leadingthe/a/anstripped, trailing's/sstripped from each token. - exact — the normalised strings are identical.
- lenient — one normalised string contains the other, the shorter being ≥ 2 characters (Japanese); or one item's content-token set is a subset of the other's (English).
- Matching is greedy longest-first, so no item is matched twice.
- Every Jaccard in this experiment is reported under both rules, always both. No headline may quote one without the other.
A2 (finding 2, BLOCKING) — the aligned span, established and frozen before dispatch
Scoring against the whole 12,868-character Yosano chapter manufactures COPYs: a common noun will be found somewhere whether or not Yosano rendered the referent at the site. The aligned span was therefore located by hand before any scoring call, by its two boundary sentences, and frozen:
- start —
ただ少しの助力でもしようとする人をも持たない女王であった。(renders the classical's openingはかなきことにても、見訪らひきこゆる人はなき御身なり。) - end —
こんなふうに末摘花は古典的であった。(renders the classical's closingかやうにうるはしくぞものしたまひける。) - offsets 1796–2857, 1,061 non-whitespace characters against the classical's 951.
Arm C scores against this span, not against the chapter. The whole-chapter figure is also computed, and both are reported, because the difference between them is the size of the artefact the critic named.
Recorded because a control that fires and finds nothing is still a result: every one of the
anchor's fifteen in-passage items has its Yosano counterpart inside the span, and 地方官 — the
substitute for 受領 — is outside it, consistent with 受領 being outside the classical
passage. The anchor's §2 calls survive the restriction. They now do so by measurement.
A3 (finding 3, BLOCKING) — P4 replaced by a registered 2×2
The original P4 was a two-way fork whose unnamed third outcome is the one that exonerates the
anchor, and it would have been absorbed into "P4 fails" and reported as the worse finding. That is a
registered asymmetry against the object of study and it is withdrawn. Let J_mm = Jaccard(P1, P3)
and J_l = max(Jaccard(lead, P1), Jaccard(lead, P3)), both on M1, under each matching rule:
J_l ≥ 0.60 |
J_l < 0.60 |
|
|---|---|---|
J_mm ≥ 0.60 |
(iii) the census reproduces and the anchor's site list is vindicated | (i) two independent readers converge and the lead does not: the lead's census is idiosyncratic |
J_mm < 0.60 |
(iv) incoherent — the lead agrees with both while they disagree with each other; report as an artefact of the matching rule and withhold | (ii) nobody has a set: the category does not define one, and the lead is not singled out |
A4 (finding 4, BLOCKING) — the "positive control" is demoted to what it is
The critic is right on both counts: Task 2 tests category enumeration rather than rule application,
and "every plant name" is not determinate in this passage (草, 木草, 薮 are all arguable), so
it could have failed for reasons having nothing to do with instrument reliability and fired F1 on
spurious grounds.
- M1's Task 2 is replaced by a mechanically determinate one: list every occurrence of the exact
string
うるはしin the passage, each with the ten characters that follow it. The answer is three, and it is checkable by string match. - M4's Task 2 is unchanged (every number expressed in words) and the mismatch in control type between the two passages is declared.
- §5's description of Task 2 as a "positive control" is withdrawn. It is an output-integrity and enumeration-compliance check: it shows the model read the passage and can return a list in the required format. It says nothing about rule-fidelity, and passing it is not evidence that Arm B's main lists are sound.
- F1 is rewritten accordingly (see A8).
A5 (finding 5, BLOCKING) — the lead's M1 census, frozen as an explicit list
The anchor's twelve-row table contains 16 distinct lexemes, of which 15 occur in M1; the
absent one is 受領, which the anchor's own table marks as coming from the adjacent §1-2. Frozen:
The lead's M1 census (15 items), the primary denominator for P3, P4 and P7: 浅茅 蓬 葎
寝殿 野分 禅師の君 紙屋紙 数珠 唐守 藐姑射の刀自 かくや姫の物語 陸奥紙 御厨子 総角 下衆.
Row-to-lexeme mapping (the anchor groups three of its rows): row 浅茅 / 蓬 / 葎 = 3 lexemes; row
唐守 / 藐姑射の刀自 / かくや姫の物語 = 3 lexemes; the other ten rows = 1 each. In-passage: 11 rows,
15 lexemes, 4 substitutions under either grouping. Every figure is reported at both granularities,
because the anchor's own headline rate is not invariant under its own grouping — 4/11 = 0.364 at row
level, 4/15 = 0.267 at lexeme level, against the published 5/12 = 0.417.
A6 (finding 6, MANDATORY) — P7 on three sets, with the weight declared
P7 is reported on the anchor's 15, on the union, and on the intersection of the two model censuses. The anchor's own 15 carries the interpretive weight, because it is the only one of the three that answers did the anchor score its own sites wrong; the union answers the different question what does the set look like if you do not get to pick it, and cannot distinguish "the anchor scored wrong" from "the set changed". The write-up says so wherever the union figure appears.
A7 (finding 7, MANDATORY) — the "conservative upper bound" sentence is withdrawn
§4's claim that a written rule should reproduce better than the anchor's unwritten one, so a low figure here is an upper bound on the anchor's, does not follow and is deleted. The conditions differ in more than one way — the anchor's author knew the text and applied his own tacit rule; the models get a reconstructed rule cold — and the direction of the bias is unknown.
A8 (findings 8 and 9, MANDATORY) — disclosure, relabelling, and F1/F3 rewritten
- P2 is relabelled an observation, not a prediction, and is excluded from any count of predictions scored. It is an entailment of M3's non-storage, not something that could have come out otherwise.
- Arm A contamination disclosure. Two design decisions were informed by pre-freeze Arm A
results, and only two: (a) the 15-item lead census of A5, and (b) the distinction between
in-passage and all-table items in §3 and P7, which exists because Arm A found
受領absent. Corrected on a fact: the critic also names the union-set construction and F3's bounds as Arm A–informed, and neither is — the union construction is carried from step 1's design and the 0.60 threshold is step 1's own registered P7, both of which predate Arm A. - F1 rewritten. If the Task 2 enumeration-compliance check fails on a passage, that passage's census lists are treated as possibly corrupt output: they are printed, and every census figure from that passage is reported as a measurement of the instrument only, never of the anchor.
- F3 rewritten. Bounds widened to fewer than 5 or more than 80, justified: the passage is 951 characters, roughly 400 words, so 80 items is one per five words and is degenerate by construction. Procedure, registered: if a census is degenerate, both counts are printed, P3 and P4 are marked not estimable on that passage, and the reserve model's census is not substituted post hoc.
A9 (findings 10 and 11, ADVISORY) — parameters and an assumption relabelled
- All census and scoring calls run at
temperature0,reasoning.effortlow, recorded here and in the run records. Reproducibility is measured at those settings and no other. - §5's claim that the main-list-first ordering makes any Task 2 contamination conservative is relabelled an assumption. A long Task 1 list could equally prime over-inclusion in Task 2. Untested, and it is not offered as a reason to believe anything.