Repository path: workshop/experiments/E-20260814d-elevation-resolution/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260814d-elevation-resolution |
| status | frozen |
| created | 2026-08-14 |
| updated | 2026-08-14 |
| senses | style-correspondence, voice, cultural-mediation |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-elevation-resolution.md, workshop/regimes/R26-light-ennoblement.md, workshop/regimes/R25-ennoblement.md, workshop/translations/flipperne/R26-v1/translation.md, wiki/findings/results/RS-20260813c-ennoblement-direction.md, wiki/findings/results/RS-20260813g-register-quadrants.md, framework/v0.2/README.md, config/models.md, config/budget.md |
E-20260814d — a register axis with three rungs, and a gate that can fail short of ceiling
Frozen 2026-08-14 (S183) before any call was dispatched. ARM-elevation-resolution step 1.
1. The question, and why the third rung is the whole design
framework/v0.2 §7.4: "It says nothing about elevation: the run's Axis E voided on a dead body
and where published hands sit above plain English is still unmeasured." That is Q-e's open half.
The predecessor run (RS-20260813c) built the upper pole and could not read the interior. Its gate
G1 required the deliberately ennobled arm to read above the plain arm by ≥ +0.50 on a −1 / 0 / +1
scale, and got +1.0000 in 56 of 56 cells, zero variance, while the published hands it existed to
measure sat at +0.125 to +0.70. Method note (bnb): a calibration gate that passes at ceiling
has not calibrated anything in the range the primaries occupy.
On a source whose own register is uniformly low, how far above it do published translators pitch their English — and can this instrument resolve the interval at all?
The answer to the second half gates the first. A third rung (R26, light ennoblement) is built so
that the calibration gate is a one-step movement rather than a maximal one, and can therefore
fail.
2. Materials
Source. H. C. Andersen, «Flipperne» (1848), 756 Danish words in 26 paragraphs, copy-text
workshop/experiments/E-20260813g-register-quadrants/corpus/DA.txt, re-fetched from Danish
Wikisource this session and identical to the stored file apart from one paragraph break.
Single-witness and uncollated; two transcription slips recorded on the translation artifacts.
Seven English arms, all of the whole tale, all held in
workshop/experiments/E-20260813g-register-quadrants/corpus/:
| arm | hand | role |
|---|---|---|
A |
Gutenberg #1597, translator unattributed in the edition | published |
PAULL |
Mrs. H. B. Paull, 1888 | published, dependence-flagged |
BRAK |
H. L. Brækstad, 1900 | published |
R06 |
lead, single pass, no register rule | lower pole |
R26 |
lead, light ennoblement (N1–N7), written this session | middle rung |
R25 |
lead, ennoblement (E1–E8) | upper pole |
R08 |
lead, resistancy | source-ward arm, exploratory only |
2.1 The materials gate, run before this design was written, and it constrains it
tools/dependence_check.py over all 21 pairs. The ten pairs not involving R26 reproduce
RS-20260813c §2.1 exactly, which is a check on the tool as well as on the texts.
| pair | 7-grams | 12-grams | 15-grams | longest run | verdict |
|---|---|---|---|---|---|
A ~ BRAK |
28 | 0 | 0 | 11 | clean |
A ~ PAULL |
24 | 5 | 0 | 14 | DEPENDENT? |
PAULL ~ BRAK |
41 | 7 | 2 | 16 | DEPENDENT? |
A ~ R06 |
49 | 8 | 5 | 19 | DEPENDENT? |
BRAK ~ R06 |
38 | 5 | 1 | 15 | DEPENDENT? |
PAULL ~ R06 |
18 | 1 | 0 | 12 | DEPENDENT? |
A ~ R25 |
18 | 0 | 0 | 11 | clean |
PAULL ~ R25 |
9 | 0 | 0 | 10 | clean |
BRAK ~ R25 |
9 | 0 | 0 | 11 | clean |
R06 ~ R25 |
7 | 0 | 0 | 11 | clean |
A ~ R08 |
27 | 0 | 0 | 11 | clean |
BRAK ~ R08 |
10 | 0 | 0 | 10 | clean |
PAULL ~ R08 |
7 | 0 | 0 | 9 | clean |
R06 ~ R08 |
48 | 8 | 2 | 16 | DEPENDENT? |
R08 ~ R25 |
9 | 1 | 0 | 12 | DEPENDENT? |
R26 ~ R06 |
263 | 113 | 72 | 27 | DEPENDENT? |
R26 ~ R08 |
67 | 17 | 5 | 19 | DEPENDENT? |
R26 ~ BRAK |
72 | 20 | 3 | 15 | DEPENDENT? |
R26 ~ PAULL |
31 | 5 | 2 | 16 | DEPENDENT? |
R26 ~ A |
63 | 2 | 0 | 12 | DEPENDENT? |
R26 ~ R25 |
9 | 0 | 0 | 11 | clean |
Five consequences, all binding, all written before any code was seen.
R26is not an independent hand and no claim about published practice may rest on it. It is a constructed pole.Q3is the only prediction that putsR26beside a published hand and it is secondary and dependence-flagged, withR26~BRAKat 20 shared twelve-grams named at every use.R06is contaminated withA(run 19) andBRAK(run 15) and is used only as the lower pole, never as an independent hand.RS-20260813c§6's argument carries over: that contamination pullsR06up, which shrinks everyX − R06difference, soG1andG2are conservative gates — the contamination can make them fail and cannot make them pass spuriously. But a second bias runs the other way and the two are stated together, per critic finding 7(b): the crib is a lead-written word-for-word English rendering, and so isR06. A seat anchoring on the crib would push whichever arm most resembles it toward0, inflating everyX − R06difference — soG1/G2are conservative with respect to contamination and not with respect to crib anchoring, and the design does not claim otherwise. One handle on it is free and is registered here: the arm that most resembles a word-for-word crib isR08, which calques, notR06, so crib anchoring predictsR08sitting nearer0thanR06does. That comparison is reported.A~BRAKis still the only clean published pair.PAULLis dependence-flagged against both others, so the three published arms are not three independent observations.Q1andQ2are stated per hand on the clean pair, withPAULLreported alongside as a dependent sensitivity.R26~R06is the closest pair in the table by a wide margin, so a design using their contrast as its manipulation must show they differ at the grain the gate runs at — critic finding 6, and the paragraph-level check the design first carried does not entail it. Checked per candidate span before dispatch and recorded inpool.json: token-level similarity over the whole texts is 0.720 against 0.388 forR06~R25; of the 27 candidate spans, 1 is identical (S23a, «Forlovet!»), 3 differ only in terminal punctuation (S10a,S18a,S20a), and the remaining 23 differ lexically. So 4 of 27 spans carry no lexical manipulation at all, and if they are sampled they contribute cells whereG1measures nothing.G1is therefore computed twice and both figures are reported: over allmarkedsites, and over themarkedsites whereR06andR26differ lexically. The first is the registered gate; the second is what says whether a failure is the instrument's or the materials'.R26~R25is clean at 0 shared twelve-grams, so the two rungs above the baseline are not each other's copies.
3. Stage 1 — the site list, three annotators, source only
The repair this stage exists for: RS-20260813c §8 measured its own site list against two
independent readers and six of ten lead-marked sites were read as unmarked. The list was one
annotator's, and that annotator was the lead. RS-20260808e's three-annotator majority vote is the
procedure being copied.
3.1 The candidate pool is built mechanically, with no lead judgement in it. Frozen rule, applied to the copy-text before any annotation:
- Split each of the 26 Danish paragraphs into its quoted and unquoted material. The quoted runs of one paragraph merge into a single candidate (they are one speech turn interrupted by its tag); the unquoted remainder is a second candidate, kept only if it is ≥ 4 Danish words. The floor applies to narration only, and critic finding 8 is right that this needed a written reason: a one-to-three-word unquoted remainder is a syntactic remnant (a speech tag, «sagde Saxen»), which has no register content of its own, whereas a one-to-three-word speech turn — «Snærpe!», «Las!», «Forlovet!» — is a complete utterance and is exactly where this source's register lives. The finding's other half is accepted as a cap rather than as a floor: see §3.3.
- Drop any candidate longer than 40 Danish words — a whole-paragraph judgement of register is a
different measurement and the four dropped blocks (59, 42, 159 and 96 words) are named in
pool.json.
Realised pool: 27 candidates, 20 speech and 7 narration, frozen in
pool.json before dispatch.
3.2 The annotation. Three seats — P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
P3 x-ai/grok-4.5 — each receive the whole Danish tale and the 27 numbered spans, and return
one code per span:
Relative to the ordinary written Danish of an 1848 storybook, does the Danish at this span read BELOW that level, AT it, or ABOVE it? →
below/at/above.
No English appears anywhere in a stage-1 prompt. Output is a closed form — one code per numbered span, 27 codes — per note (bng).
3.3 The majority rule and the sample, both frozen here.
- A span is
markedif ≥ 2 of 3 annotators code itbelow;neutralif ≥ 2 of 3 code itat; otherwise it is excluded (no majority, or a majority ofabove). - From each class, order by
int(sha256("E-20260814d|" + span_id).hexdigest(), 16)ascending and take the first 10markedand the first 5neutral. Fewer than that available → take all. - At most 3 of the 15 sampled sites may be spans shorter than 4 Danish words (critic finding 8).
Once three such spans have been taken, later short spans are skipped in hash order and the next
eligible span of that class is taken instead. Six candidates are affected —
S12a,S14a,S15a,S23a(1 word),S20a(2),S18a(3) — and coding seven renderings of a single interjection is a materially harder judgement than the design's other cells. - The lead's role in stage 1 is confined to building the mechanical pool and extracting each arm's aligned text. The lead does not mark, vote, or break ties.
3.4 The dependence this creates, declared rather than discovered. The three annotators are the
same three seats that judge in stage 2. Calls are stateless and no stage-2 prompt contains a
stage-1 answer, so the dependence is at the level of a model's dispositions, not of memory — but a
seat that reads a span as low may be likelier to read an English rendering as above it, and that is
stated as a limit rather than argued away. The alternative — annotating with the reserve seats — was
rejected because their Danish is entirely unprobed and P5 is excluded outright by note (bne).
4. Stage 2 — the height run
One call per (site × seat), 3 seats, ≤ 15 sites: at most 45 calls. P5 is not seated, on note
(bne). Judgment is not parallelised: each call is stateless and sees one site.
Each site payload carries, for the site the annotators selected:
- the Danish span;
- a crib written by the lead: word-for-word rendering plus dictionary senses, and nothing else.
Machine-checked before dispatch against a banned register vocabulary — colloquial, homely,
blunt, everyday, literary, elevated, decorous, insult, shout, spoken, plain, formal, register,
low, high, ornate, dignified, slang, vulgar — because
RS-20260813c's critic finding 1 caught exactly this leak, and the prompt says in terms that the crib is not English prose and is not a yardstick of level; - the seven arms' renderings of that span, under labels
V1…V7, scrambled per (site, seat) bysha256(site_id | seat | "E-20260814d") mod 5040, so no arm sits behind a fixed label.
Each call asks two things, both closed-form:
- (a)
source_level— relative to the ordinary written Danish of an 1848 storybook, does the Danish at this site read below that level, at it, or above it? →-1 / 0 / +1. - (b)
heights— for each ofV1…V7independently: relative to the Danish at this site, does this English read below its level, at it, or above it? →-1 / 0 / +1.
Two instructions carried over from the predecessor's amended prompt, because they were accepted findings there: length is not level — judge the words chosen, not the number of them (critic finding 7), and no rendering is described as "independent" (finding 9).
Blinding assertions, split per note (bnm). The scaffolding is fixed text and carries the ban
absolutely: no prompt contains any of the 19 provenance strings Andersen, Paull, Brækstad,
Braekstad, Berman, ennobl, regime, foreigniz, resistancy, Venuti, vulgaris, Gutenberg, R06, R08,
R25, R26, 1888, 1900, 1597, matched as substrings (several are stems). 1848 is deliberately
NOT on this list — critic finding 1, and it was a BLOCKING error: the date is the period anchor the
question itself needs ("the ordinary written Danish of an 1848 storybook"), so the frozen
assertion was unsatisfiable the moment it was written, exactly as note (bnm) describes. 1848
identifies no translator and no arm; every English arm in the run postdates it. The crib is
lead-written data and carries the 19 register words ban — colloquial, homely, blunt, everyday,
literary, elevated, decorous, insult, shout, spoken, plain, formal, register, low, high, ornate,
dignified, slang, vulgar — matched on word boundaries, not as substrings, because low is
inside below and plain is inside explain (critic finding 10). The verifier counts any
breach rather than assuming none.
Caps, per note (bnl). Reasoning is never disabled on a judging seat. One probe call per
seat is dispatched on a real stage-2 prompt, reasoning_tokens is read off the response, and the
cap for the remaining dispatches is set at measured reasoning + 400 for the answer, floored at
1,200 and ceilinged at 3,000. The probe's own cap is 2,500. Stage 1's cap is 4,000 for all three
seats and is not probed, because the stage is only three calls.
Re-dispatch, per note (bnn). A body is usable only if it parses as JSON and carries its
stage's key (levels for stage 1, source_level and heights for stage 2) with the right arity.
Anything else is re-dispatched once at 2× the cap but never above 4,000 tokens, and the
stored record is marked usable: false so the analysis and the F3 count key off the same fact.
Stage 2's re-dispatches are capped at 6 in total (critic finding 4): beyond the sixth, an
unusable body is recorded usable: false without a retry, and the number not retried is
reported. Without that cap the design's own rules permit 42 extra calls and the declared ceiling is
not a ceiling.
5. Gates, run before any prediction is read
Aggregation, fixed here for every gate and prediction: a figure is the mean of seat-means
over seats surviving G4, computed over cells that returned a usable body. Minimum 2 surviving
seats.
G1— the resolution gate, and it is built so it can fail. Atmarkedsites, mean height(R26) − mean height(R06) ≥ +0.25. This is note (bnb)'s registered remedy: the gate separates the one-step arm from the baseline, not the maximal one. A failure means the instrument does not resolve a movement the size of the one published hands make, and that is a reportable outcome, not an accident.G2— the gross-movement check. Reported; gates nothing. Atmarkedsites,R25−R06≥ +0.50 — the predecessor'sG1verbatim, kept so that aG1failure beside aG2pass is interpretable as sees a maximal movement, not a one-step one rather than as sees no register at all.G3— the crib check. Not a model call. Every crib passes the banned-register-vocabulary screen before dispatch, or the site is re-cribbed until it does. Recorded as pass/fail per site.G4— inter-seat coherence, NOT agreement with anybody's map. Rebuilt per critic findings 2 and 3, both BLOCKING and both correct. With three seats, "the majority of the other seats" is a majority of two and does not exist when they disagree, so the statistic was undefined on exactly the sites that matter; and a fixed bar of 8 over an eligible set that can fall below 8 would have voided the run by arithmetic — note (bna)'s failure reintroduced one clause after the note was cited. The rebuilt gate: a seat is retained if itssource_levelcode agrees with at least one other seat's at ≥ 2/3 of the sites where at least one other seat returned a usablesource_level. If fewer than 6 sites are eligible,G4is indeterminate: it is reported as not having run, every seat returning ≥ 80% usable bodies is retained, and the result page says the coherence gate did not fire rather than pretending it passed. Agreement between the stage-2 seat majority and the stage-1 annotator majority is reported as a measurement, never used as a drop rule; it is this run's only check on whether a site list survives being re-asked in a different frame.
6. Predictions, registered before any call
Q1— the census. PRIMARY, on the clean pair only. Atmarkedsites, each ofAandBRAKhas mean height > +0.20. This putsRS-20260808e's ten-hand result to a sixth language pair, and it is the first time it will have been read on a gated interior.PAULLreported alongside as a dependent sensitivity, not part of the primary. Conditional onG1.Q2— is a published hand distinguishable from a deliberate ennobler? PRIMARY. Atmarkedsites,R25−A≥ +0.30 andR25−BRAK**≥ +0.30. **The outcome of interest is the null**: a Victorian hand arriving level with a rule set built out of Berman's complaint would say the complaint is a description rather than an exaggeration, which is the substantive questionARM-ennoblementwas constituted on and never answered. Conditional onG1`.Q3— where do published hands sit relative to the ORDINARY pole? SECONDARY, dependence-flagged at every use. Atmarkedsites, |A−R26| < 0.30 and |BRAK−R26| < 0.30. If it holds, a Victorian published hand and a modern one-step-up policy are the same height above this source, which is the practically useful form of the elevation question.R26~BRAKshares 20 twelve-grams andR26~Ashares 2; the flag travels with the number.Q4— policy discrimination atneutralsites. SECONDARY, n small and stated.R25E1 raises the neutral by rule;R26N1 explicitly leaves it alone. Registered: atneutralsites,R25−R06≥ +0.25 and |R26−R06| < 0.25. This asks whether the seats track two declared policies or a single global impression of fanciness, and it is the cheapest test of that this project has been able to build.Q5— the direction question. EXPLORATORY. Atmarkedsites,R25−R08≥ +0.30: markedness toward the source is not coded as elevation.RS-20260813b§7 found a target-only style question could not tell the two kinds of markedness apart; this asks the same question source-relative, which is the formR25's regime page says is answerable.R08is dependence-flagged againstR06andR26.
7. Failure criteria, registered
F1—G1fails →Q1,Q2andQ3are withheld.G2,Q4,Q5, the stage-1 site list and the full descriptive table are still reported, as instrument facts, with the resolution failure stated as the run's result.F2—G4leaves fewer than 2 seats → every prediction is withheld.F3— more than 10% of the bodies belonging to seats retained byG4are unusable after one re-dispatch → the run is reported as incomplete and every prediction is withheld. A seat dropped byG4cannot also void the run throughF3— this is note (bna)'s generalisation applied to the criterion that fired at S174, and it is registered here before dispatch rather than re-read afterwards.F4— stage 1 yields fewer than 6markedor fewer than 3neutralspans by majority → stage 2 is not dispatched and the run reports the site-list measurement only. The money not spent is recorded as not spent.F5— fewer than 3 annotators return a usable stage-1 body after one re-dispatch → the majority rule is undefined and stage 2 is not dispatched.- No override of
F1–F5is available to this session.RS-20260813b§6 refused an override that would have rescued its own headline andRS-20260813c§5 refused a second; those refusals are the precedent and they bind here.
8. What this run cannot establish
- Nothing about quality. Register height is a position, not a verdict. Tier D remains NOT PASSED;
every figure is
internal-judgment-onlyandprovisional. - Nothing about three independent published hands (§2.1(3)).
- Nothing attributing a
R25−R26gap to lexical height alone. The two regimes differ in height and in whether expansion and clarification are permitted;R26's own limitations section says the reduction is compound. - Nothing about Danish generally, and nothing about any period's norms:
R26's ceiling is ordinary now, and a Victorian hand's ordinary sits above a modern one's by an amount this design does not measure and cannot subtract. - Nothing about readers. Three language models coding height are not readers, certified or otherwise, and their Danish is unprobed at 1848 orthography.
- Nothing that survives
F1. If the interior cannot be resolved, the census numbers are not evidence and the run says so.
9. Money
Pre-flight priced from config/models.md as corrected 2026-08-14 (P2 $0.75/$3.75), from
max_tokens and not from an assumed answer length (note (abc)), with re-dispatch in the
arithmetic (note (abc)'s S079 firing).
| stage | calls | cap | worst case |
|---|---|---|---|
| stage 1 — annotation | 3 (+3 re-dispatch) | 4,000 | $0.16 |
| stage 2 — reasoning probe | 3 | 2,500 | $0.06 |
| stage 2 — main | 42 | ≤ 3,000 | $0.95 |
| stage 2 — re-dispatch, capped at 6 (critic finding 4) | ≤ 6 | ≤ 4,000 | $0.16 |
pre-run critic (P4, effort low, note (bnk)) |
1, actual $0.112383 | 6,000 | $0.25 |
| declared ceiling for this run | $1.50 |
Session ceiling $1.50. UTC-day headroom at design time: $2.763776 of $5.00 remaining
(2026-08-14, three prior sessions at $2.236224). A run that does not fit is split or deferred; F4
and F5 are both routes to spending materially less than the ceiling and that is a normal outcome.
Lead translation of «Flipperne» under R26 is $0 and is not ledgered (charter §3, A4), as are
the copy-text re-fetch, the dependence measurement, the mechanical pool and the verifier.
10. Verification
analysis/verify.py, importing nothing from tools/, recomputes every reported number from the
stored bodies and asserts:
- every stage-1 and stage-2 prompt is free of all 19 banned provenance strings;
- every crib is free of all 19 banned register words, and the count of breaches is reported rather than assumed zero (note (bnm));
- every arm span in every payload is a contiguous substring of that arm's stored corpus file, normalising whitespace only — the guard against a rendering being paraphrased into the payload;
- the label scramble is a bijection at every (site, seat) and agrees with the runner's record;
- every individual cell code matches the stored body, not merely the aggregate —
RS-20260813b§8's repair, where a flipped cell survived a majority-level check; - every reported mean recomputes from the per-cell codes;
- the site list recomputes from the stage-1 bodies under §3.3's majority rule and hash order;
- mutation sensitivity: flipping one stored code, and separately corrupting one payload span, each change at least one asserted figure.
11. Amendment before dispatch: the pre-run critic's findings, all ten accepted
Critic: moonshotai/kimi-k3 (P4), no role in this run's annotation or its jury (P1/P2/P3),
and not a seat in it at all. Ten findings, six BLOCKING, four NON-BLOCKING. All ten accepted; two of
them in an amended form with the reason written. The design above is the amended one, and nothing
had been dispatched. Cost $0.112383000, provider Chutes, 4,034 reasoning tokens of a 6,000
cap. Raw body: run/critic.json.
The verdict line did not arrive: the body hit finish_reason: length mid-finding-10, so the
critic's own summary verdict is missing and finding 10 is truncated after its first sentence. The
sentence is complete enough to act on and was acted on. Note (bnk) fired again, on the seat it
names, with its own remedy already applied — effort pinned low in the first dispatch and the cap
sized from the observed 6,000-token pass at S181 — and it still truncated, because ten findings is a
longer answer than eight. Note (bng) is the more exact diagnosis: an open-ended enumeration is the
one output shape whose cap cannot be guessed, and a critique is exactly that shape. The finding is
recorded rather than repaired: a second dispatch at a larger cap would have bought a verdict word
this session does not need, at the price of a second $0.11.
| # | finding | disposition |
|---|---|---|
| 1 | BLOCKING — the blinding ban listed 1848, and the design's own question text contains 1848. The verifier's first assertion would have failed on every prompt, before any model was called. |
Accepted, and it is the most serious finding. 1848 struck from the provenance list, which is now 19 strings; the reason is written into §4. The date is a period anchor the question needs and identifies no translator and no arm. This is note (bnm) exactly — a blindness assertion the design could not hold, unsatisfiable the day it was frozen — caught this time by a critic instead of by a verifier after the money was spent. |
| 2 | BLOCKING — G4's "majority of the other seats" is undefined with three seats. A majority of two requires both, and the statistic does not exist when they split — on the path to every prediction through F2. |
Accepted. G4 rebuilt as agrees with at least one other seat. §5. |
| 3 | BLOCKING — G4's fixed bar of 8 sits over a denominator that can be smaller than 8, so on the F4-minimum path (9 sites) a single dead body would void the run by arithmetic. |
Accepted, and it is note (bna) reintroduced one clause after the note was cited by name. The bar is now a proportion (≥ 2/3 of eligible sites) with a frozen minimum eligible count of 6, below which G4 is indeterminate and reported as not having run rather than failing everybody. |
| 4 | BLOCKING — the stage-2 re-dispatch tail is unbounded, so $1.42 was not a worst case. The design's own rules permit 42 extra calls at up to 6,000 tokens against $0.03 of slack. | Accepted. Re-dispatches capped at 6 and at 4,000 tokens, the number not retried is reported, and the declared ceiling is raised to $1.50 with the row priced. Note (abc)'s S079 firing said a worst case is max_tokens × attempts × slugs; the attempt count was missing from one row and the critic found it. |
| 5 | BLOCKING — the provenance list has 20 strings and §10 asserts 19 twice, so either one string goes unverified or the frozen text is wrong. | Accepted. With 1848 struck (finding 1) the provenance list is 19 and the register list is 19; §4 and §10 now name which count belongs to which list. |
| 6 | BLOCKING — the manipulation check is at paragraph grain and the gate runs at span grain. If the sampled marked sites are spans where R26 = R06, G1 measures nothing and fails for a materials reason the design would report as an instrument reason. |
Accepted in full and it is the finding that most changed the run. Every one of the 27 candidates was diffed before dispatch and the result is in pool.json: 1 identical, 3 punctuation-only, 23 lexically different. §2.1(4) now carries the numbers, and G1 is computed and reported twice — over all marked sites, and over those where the two arms differ lexically. |
| 7 | NON-BLOCKING — §2.1(2)'s "contamination is conservative" argument does not cover Q3, and the crib introduces a bias in the opposite direction that no section mentions. A literal crib resembles R06 more than any other arm; a seat anchoring on it pushes R06 toward 0 and inflates every X − R06 difference, so G1/G2 can pass spuriously after all. |
Accepted in full, both halves. §2.1(2) now states the crib-anchor bias and its direction and withdraws the unqualified conservatism claim; Q3 now says its threshold can be passed by contamination and not only failed by it. One amendment beyond what the critic asked: the arm that most resembles a word-for-word crib is R08, which calques, not R06 — so crib anchoring makes a prediction about R08, and that comparison is now reported as a handle on the bias rather than left as an unfalsifiable worry. |
| 8 | NON-BLOCKING — the ≥ 4-word floor applies only to narration, so up to 6 of 15 sites could be single interjections, and the "mechanical, judgement-free" rule embeds an unjustified asymmetry. | Accepted as a cap rather than as the critic's floor, with the reason written. Applying the floor to speech would drop «Snærpe!», «Las!» and «Forlovet!» — the sites where this source's register most obviously lives, and the ones the predecessor run coded. The asymmetry now carries its justification (a short unquoted remnant is a speech tag with no register content; a short speech turn is a complete utterance), and at most 3 of the 15 sampled sites may be shorter than 4 Danish words. |
| 9 | NON-BLOCKING — seats and predictions both use the labels P1–P5, so P5 names an excluded seat in §4 and a live prediction in §6. |
Accepted, in the form that costs no cross-references. The predictions are renamed Q1–Q5; the seats keep the names config/models.md gives them, since renaming a seat here would desynchronise this design from every other page in the project. |
| 10 | NON-BLOCKING (truncated) — the banned-word screen substring-matches, so low fires inside "below" and plain inside "explain". |
Accepted. The register screen now matches on word boundaries; the provenance screen stays substring, because several of its entries (ennobl, foreigniz, vulgaris) are deliberately stems. |
One thing that is not the critic's and is recorded so the amendment list is honest: the site
hash keys on span_id only, so none of these amendments changes which spans the sample will draw or
which label any arm sits behind.
12. Pre-dispatch amendment A1 — after stage 1, before any stage-2 call
Written and committed 2026-08-14 after stage 1 returned and before a single stage-2 call was
dispatched. Stage 1's three annotators returned 3 of 3 usable bodies and produced a majority site
list far thinner and far less agreed than the design anticipated: 7 marked, 13 neutral, 0
above, and 7 spans with no majority at all (every one of those seven split three ways,
above/below/at). F4 does not fire — its bar is 6 marked and 3 neutral by majority,
and both are met.
But two frozen rules then interact in a way nobody foresaw at freeze time. Six of the seven
marked spans are shorter than four Danish words, and §3.3's short-span cap (critic finding 8)
allows at most three such spans in the sample. The frozen sample is therefore 4 marked sites,
of which — by §2.1(4)'s own per-span table — S23a is lexically identical between R06 and
R26 and S20a differs only in terminal punctuation. Half the cells of the registered
resolution gate would carry no manipulation at all.
The amendment, and what it deliberately does not do.
- All 7 majority-
markedsites are dispatched, not 4 — 12 sites, 36 calls instead of 27. - No registered bar moves and no registered denominator changes.
G1,G2,Q1,Q2,Q3andQ5are computed on the frozen 4-sitemarkedsample, exactly as §5 and §6 say. The 7-site figure is reported as a stated sensitivity and may not replace the primary, whichever direction it goes. - This is registered here, in writing, before dispatch, because the failure mode it courts —
computing two figures and choosing between them afterwards — is the one this project's verification
discipline exists to stop. Adding cells cannot rescue a bar that is computed over a fixed set;
what it can do is make a
G1failure interpretable rather than ambiguous, which is the whole reasonG1was built to be failable. - The three sensitivity sites are
S18a,S14aandS12a, named now so the set cannot be chosen later.
Two stage-1 facts recorded here rather than in the result, because they bear on the dispatch.
P3 returned 6,516 reasoning tokens against a max_tokens of 4,000 and still delivered a
complete body — so on that provider the cap did not bound the reasoning at all, and the call cost
$0.045616 against a per-call worst case of $0.0272 built from the cap. Note (abc)'s worst
case is not a bound on a reasoning seat, which is note (bnk)'s point arriving from the other
side. And the three annotators' base rates for below over the same 27 spans are 7, 20 and 3,
which is the measurement §5's G4 was written to catch and the reason the stage-2 probe is worth
its three calls.