Repository path: workshop/experiments/E-20260730i-candidate-reach/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260730i-candidate-reach |
| status | frozen |
| created | 2026-07-30 |
| updated | 2026-07-30 |
| senses | accuracy, naturalness, style-correspondence, voice, cultural-mediation, consistency, purpose-fit |
| provisional | true |
| links | wiki/arms/ARM-candidate-reach.md, framework/closure.md, framework/traceability-inventory.md, wiki/findings/results/RS-20260728i-coverage-independent.md, wiki/findings/results/RS-20260730h-strict-coverage.md, wiki/findings/results/RS-20260729d-decision-grain.md, workshop/translations/cartomante/source-pt.txt, workshop/regimes/R04-lead-close.md, config/models.md, config/budget.md |
E-20260730i — how far do the fourteen candidates reach, once the option list is counted?
FROZEN before the source text was read for translation and before either of the two archived
decision logs was re-read under the rule below. ARM-candidate-reach step 1.
1. Question
framework/closure.md §1.4 reports DECIDES: 0 of 126 and calls it "the sharpest thing this
page knows" and "a fact about the candidates." RS-20260730h (S067) has just shown that the
project's other coverage statistic — the R07 D-rate — was rank-identical across four runs to the
mean number of live renderings the translator wrote down, i.e. it measured the option list rather
than the rules.
DECIDES and INVOKED are the strict and weak readings of one relation and both are functions of
the option count: with n live renderings, a candidate that rules out one can return INVOKED at
any n and can return DECIDES only at n = 2.
Question. Is §1.4's zero a property of the fourteen candidates, or of how many options were on the table when the 63 decisions were logged?
2. Standing and discipline
provisional: true. The raters are panel models; the panel is NOT CALIBRATED. This is classification against a written list, not a quality judgment — the task shapeRS-20260728iused andconfig/models.md(S015) records the panel as strong on in the failing direction. Rater disagreement is informative; rater agreement is weak evidence.- The lead never judges its own translation (charter §5). No quality claim about
T-cartomante-R04-v1is made, elicited or implied anywhere in this design. - Every evaluative sentence in the translator's log carries
internal-judgment-only; the log is frozen in git before this design's condition C is built. - Charter §5.5: nothing here calibrates anything.
3. Declared confounds and priming, before the run
- The lead read Log A in full at session start, while establishing whether the archived logs record option lists at all, and before this design was written. Log B was not read. Condition A is therefore reported split, with Log B as the blind arm, and no pooled proportion from A is quoted without both halves beside it.
- The lead saw the first ~80 characters of each paragraph of the Portuguese source while establishing the span and checking its orthography. No full paragraph was read before this design was frozen.
- The published English (Isaac Goldberg, "The Fortune-Teller", in Brazilian Tales, Boston:
Four Seas, 1921; Project Gutenberg #21040) is downloaded and unopened. Its table of contents
was displayed. The contamination gate runs on Unit A before the rest is drafted, per the
standing selection-gate rule in
CLAUDE.mdand method note (bcd). - The direction of any correction favours the project, and this is the confound to be most
suspicious of. On
R07the option-count artifact made a published rate look too high; here the same arithmetic would make the zero look too damning. The result page must state this.
4. Materials
4.1 The translation limb — T-cartomante-R04-v1
- Source. Machado de Assis, «A cartomante», Várias Histórias (Rio de Janeiro: Laemmert, 1896), pp. 9–29; first published in Gazeta de Notícias, 1884. Public domain (author d. 1908).
- Span, fixed here before reading: paragraphs 32–56 of the extracted text — from "Imaginariamente,
viu a ponta da orelha de um drama…" through the fortune-teller's descent to the street. 1,221
Portuguese words. Stored at
workshop/translations/cartomante/source-pt.txt(sha256prefix9286a29d9acbd0e8). - Witnesses. W1 = pt.wikisource transcription of the 1896 edition (modern orthography, mixed
proofreading status). W2 = Google/Internet Archive OCR of the 1896 first edition
(
variashistorias00assigoog, original orthography, heavy OCR noise). W2 is collated against W1 for substantive lexical divergence only, not for orthography or diacritics; divergences are listed on the translation artifact. - Why this material, under criteria fixed before the passage was chosen: (a) a language pair the project has never worked in — PT→EN is new — so the candidates are tested away from the pairs they are evidenced on; (b) original and a published English both free, so the contamination gate can actually run; (c) realia-dense and register-mixed (a hired cab, mil-réis, an Italian fortune-teller's Italian, sustained free indirect discourse), so several candidates have a chance to bear; (d) no earlier lead translation has touched it.
- Regime.
R04v1.0 (lead close translation, source only). - Unit A for the contamination gate = paragraphs 32–35 of the span (≈394 words), translated and frozen before paragraphs 36–56 are drafted.
4.2 The census, and the order it is built in
This ordering is the design's load-bearing part and any deviation from it voids condition C.
- Translate, and at every site where more than one rendering was live, write down every rendering
the translator was actually willing to use — the option list — at the moment of decision, before
consulting
framework/traceability-inventory.md. - Freeze the translation and the log in git.
- Only then: the lead records its own candidate assignment per site, in a separate file, committed separately.
- Only then: build condition C's items from the frozen census.
A site enters the census iff n ≥ 2 live renderings were named. Sites with n = 1 are not
decisions between options and are excluded — the class RS-20260730h §3 found in four of the 39
published R07 D codes and which its pre-run critic caught in that session's own control.
4.3 The archived logs for condition A
T-mare-au-diable-R04-v1 (Log A, 21 entries, FR→EN) and T-takasebune-R04-v1 (Log B, 42
entries, JA→EN) — the two logs RS-20260728i classified, verbatim, entry by entry.
5. Conditions
Condition A — what option counts the 126 classifications were computed over (panel, 3 raters)
Each rater receives all 63 entries verbatim, in log order, and for each returns:
n_named— the number of distinct renderings of one site the entry explicitly names as having been live, counting the chosen one. Explicit naming only: a rendering must appear as a quoted or italicised string, or be named unambiguously ("the paraphrase", "the calque"). No reconstruction, no inference about what the translator might also have considered.NOT-A-SITE-CHOICEwhere the entry is not a choice between renderings of one site (a whole-text policy, a cross-text consistency note, a statement of what was not resolved).
Controls in condition A, interleaved and unlabelled: 2 synthetic entries naming exactly 4 renderings, 2 naming exactly 2, and 2 that are whole-text policies with three renderings mentioned as examples but not as options for one site (a rater who counts strings rather than options fails these).
Condition B — the lead's own coding (free, lead)
Per census site, the lead records which candidates it would cite and, per candidate, which options that candidate excludes. Committed after the census freeze and before any panel call. This is the comparator, not the measurement.
Condition C — per-option candidate coding (panel, 3 raters)
Each rater receives, per item: the Portuguese source string with sentence context, a literal gloss,
the full option list, and the texts of all fourteen candidates from
framework/traceability-inventory.md §2 — including C12, with its text as written and no note
that Tier D's failure makes it inadmissible (RS-20260728i §1's rule: withholding it rigs the
count).
For each item the rater returns, for each candidate: BEARS or NOT; and for each BEARS
candidate, for each option in the list: EXCLUDES or PERMITS.
Derived per (site, bearing candidate), with n = options and k = options excluded:
| derived code | definition |
|---|---|
INVOKED |
0 < k < n |
DECIDES |
k = n − 1 |
VACUOUS |
k = n — the candidate excludes every rendering the translator was willing to use |
E |
k / n — the option-count-normalised exclusion fraction |
Controls in condition C, interleaved and unlabelled:
CTRL-POS×2 — a site withn = 4constructed so that a named candidate excludes exactly three.n ≥ 3is required: a control atn = 2would be passed by a rater who excludes one option by reflex, which is the defectRS-20260730h's critic caught in that session's own control.CTRL-VAC×1 — a site where a named candidate excludes all four options.CTRL-NEG×1 — a site withn = 3where no candidate bears (a choice between two synonyms of a common verb with no register, realia, grammar or consistency dimension).
A rater passes the controls iff all four are coded correctly.
Repeat control. One rater's condition-C payload is re-issued byte-identically in a fresh call; cell-level agreement is reported.
6. Registered predictions
| # | prediction | what falsifies it |
|---|---|---|
| P1 | (replication) panel-majority DECIDES rate over real (site, bearing-candidate) pairs is ≤ 0.10 |
a rate above 0.10 — §1.4's zero does not replicate on a third language pair |
| P2 | (replication) site-level weak coverage — ≥1 candidate INVOKED at the site — falls in 0.30–0.60, the band RS-20260728i §3 found on two pairs |
a rate outside the band |
| P3 | the confound. Over bearing pairs, Spearman ρ(k, n) ≤ +0.25 and mean k ∈ [0.8, 1.6] — a bearing candidate excludes about one option regardless of how many there are | ρ > 0.25 with mean k scaling in n — the candidates do more work when there is more to do, and the zero is a fact about them |
| P4 | the consequence. DECIDES rate at n = 2 is strictly greater than at n ≥ 3 |
equal or lower rates |
| P5 | (retrospective) in condition A, ≥ 50% of the 63 entries are n_named ≥ 3, n_named ≤ 1, or NOT-A-SITE-CHOICE — i.e. most of the 126 classifications were made where DECIDES was arithmetically out of reach or where the option set is not recorded at all |
under 50% |
| P6 | closure-defeating, registered as such. If a candidate outside {C1, C2, C3, C5, C6, C7, C9} is INVOKED at ≥ 3 sites by panel majority, then framework/closure.md §1.4's "7 of 14 were never invoked at all" must be restated, not annotated, and this arm may not close on step 2 without doing so |
— (it is the trigger, not a prediction to be scored) |
P3 is the primary. P1 and P2 are replications and are reported as such; a run that only replicated them would have bought nothing.
7. Failure criteria — pre-committed
- F1. Repeat-control cell agreement < 0.85 → no figure from condition C is reportable, including P1–P4.
- F2. Fewer than 2 of 3 raters pass all four condition-C controls → nothing from C is reportable. (The S064 instrument scored 1 of 3 on its controls; the S067 rebuild scored 4 of 4 for all three. This design inherits the S067 question shape — per-option judgments, not a multi-way classification — for that reason.)
- F3. Condition-A rater agreement on
n_named, as Krippendorff α on the ordinal counts, < 0.50 → condition A reports no proportion as an estimate; the distribution is described and P5 is scoredUNRESOLVED. - F4. Fewer than 12 real census sites survive §4.2 → condition C is descriptive only and no rate is reported as an estimate.
- F5. Any deviation from §4.2's ordering → condition C is void, not adjusted.
8. Panel seats and reserves (note (bfc): a reserve is declared for every seat before dispatch)
| role | seat | slug | reserve |
|---|---|---|---|
| pre-run critic | P4 | moonshotai/kimi-k3 |
x-ai/grok-4.5 (P3), which is not a rater in condition A |
| rater 1 | P1 | openai/gpt-5.6-terra |
P2 google/gemini-3.6-flash |
| rater 2 | P3 | x-ai/grok-4.5 |
P2 |
| rater 3 | P5 | deepseek/deepseek-v4-pro |
P2 |
If P3 serves as critic reserve it may not also serve as rater 2; rater 2 then falls to P2 and rater 1's reserve becomes P2 as well, with the collision declared on the result page.
Note (b), seventeen firings: P5 (deepseek-v4-pro) has returned finish_reason: length with an
empty body and unreturned reasoning repeatedly. Its call carries max_tokens at the stage cap and
no reasoning field; on an empty body the reserve fires immediately and the wasted cost is
ledgered.
9. Pre-flight cost estimate — built from max_tokens, not from expected output (note (abc))
| stage | calls (incl. reserves) | max_tokens |
worst case |
|---|---|---|---|
| pre-run critic | 1 + 1 | 8,000 | $0.25 |
| condition A | 3 + 2 | 4,000 | $0.30 |
| condition C | 3 + 2 | 8,000 | $0.40 |
| repeat control | 1 | 8,000 | $0.12 |
| total | $1.07 |
Day headroom at session start: $1.7727 on the per-request sum, $1.7289 on the conservative
key-delta reading (config/budget.md, S067's closing rows). The estimate fits either. If a
stage overruns, condition A defers and is reported as not run.
Everything else in this experiment is free: the translation, the census, the collation, the contamination gate, the item build, the analysis and the independent verifier.
10. Verification
analysis/verify.py imports nothing from analyse.py, recomputes every reported number from the
stored raw bodies, and carries mutation tests — deliberately corrupted inputs that the checks
must catch (RS-20260730h ran four and caught four). Verifier output goes on the result page with
its check count and failure count.