Repository path: wiki/archive/method-notes-S032-S243.md · rendered 2026-09-09
Page metadata (front matter)
| type | ledger |
|---|---|
| id | method-notes-archive-S032-S243 |
| status | active |
| created | 2026-07-26 |
| updated | 2026-09-03 |
| links | NEXT.md, wiki/tracks.md, workshop/experiments/README.md, wiki/findings/results/RS-20260731h-carryover.md |
Standing method notes
What this is for. These are the project's accumulated procedural lessons, each one paid for by a defect. They lived as a single unbroken paragraph at the bottom of NEXT.md, re-typed every session, and had grown to roughly fifty items — carried in full every session, unsorted, with no way to tell a live note from a dead one.
The wall was already losing items. Note (ii) — a self-test fixture must be shaped like the target material — was created in S025, fired in S026 and again in S027 ("third session running"), and had vanished from the list by S031. Nothing recorded its removal. It is restored below. That is the argument for this page: an append-only paragraph re-typed by hand loses things silently, and the loss is invisible precisely because the thing that would notice is the list itself.
Ids are permanent. The letter ids are cited verbatim in frozen result and design pages, so they are never renumbered, reused, or tidied. A discharged note keeps its id and is recorded as discharged — (n) is the example.
One id was issued twice, and it stays that way. (bha) names two notes: the audit-is-a-write note (S080, §Live below) and the front-matter/senses note (S081). Both are cited by that id in frozen pages, so renumbering either would break a citation, which is the one thing this rule exists to prevent. Decided at S084 rather than tidied: a citation of (bha) is resolved by its subject, not by its number — audits and execution-in-place resolve to the S080 note, senses: front matter to the S081 note — and both entries carry a pointer to the other. No third note may take the id.
Compaction rule
- A note fires when a session's work is changed by it. Record the session in
last fired. - A note discharged by a permanent structural fix (a checklist, a load-time refusal, a tool gate) moves to §Discharged with what discharged it. It is not deleted.
NEXT.mddoes not reproduce this list. It cites at most the two or three notes bearing on the next unit, by id.
Design and pre-registration
| id | note | first seen | last fired |
|---|---|---|---|
| (bsr) | A rate whose denominator moves with the predictor is not evidence, and a correlation registered on one cannot fail. Before registering any ratio against a predictor, ask what the ratio must do when the predictor moves and nothing else does. E-20260901b v1 registered "FILL falls as the metre's long-fraction rises, ρ ≤ −0.4" as its test of a translator's printed claim that English cannot supply stress for long-heavy measures. FILL is stresses-on-long ÷ long positions. With a roughly constant stress count per line, more long positions mechanically means a lower FILL, whatever the translator does — the prediction could not have failed, and the pre-run critic (P3-1) killed it before any figure existed. The measured ρ was −0.816, which would have read as strong confirmation of a claim it does not touch. Fires at: every design registering a proportion, rate, density or per-unit figure against a predictor that appears in its own denominator — coverage per opportunity, hits per site, cost per token, carriage per locus. Remedy, and it is a decomposition rather than a warning: split the quantity into the part the arithmetic forces and the part it does not. Here that was DEFICIT (the longs a line could not have filled however placed, which is the claim in the claimant's own terms) and SUPPLY (of the stress the line has, the share placed on longs, bounded and free to move either way). The forced part is reported and labelled forced; the free part carries the prediction. RS-20260901b-leaf-pattern §5; critic-response.md. |
S238 | S238 |
| (brp) | A gate built on "did the seat use the disclosed fact?" must not be scored on vocabulary the disclosure DISPLACES. Told nothing, a reader talks about rhythm; told about the source, it talks about the source — and a keyword list of unprompted sound-words then falls when engagement rises. E-20260825c's pre-run critic correctly killed a keyword list containing rhym, arabic, translat, original as pure prompt echo (both seats, independently), and the accepted remedy — restrict the list to vocabulary the prompt does not supply — produced a gate that ran the wrong way: unprompted sound-words fell 0.569 → 0.278 from the no-information arm to the disclosure arm, while the echo-inclusive rate rose 0.750 → 0.847, and the free-text reasons show why ("superior poetic cadence" becomes "preserves the rhymed prose (saj')"). Both lists are uninformative about engagement, in opposite directions, and neither is repairable by a better word list. Fires at: any design that checks whether an information manipulation was used, by coding what the subject wrote. Remedy: score engagement on a behavioural contrast the manipulation predicts and a placebo does not — here the specificity control (P3), where the same disclosure moved one segment each way on a pair neither arm of which rhymes, did the whole job the two keyword gates could not. A manipulation check that reads the subject's words is a check on register, not on uptake. RS-20260825c-worth-paying §6. |
S222 | S229 |
| (bro) | A two-way forced choice between two competent texts, offered with no context, measures presentation order — and the size of that is not a nuisance parameter, it is the first thing a preference design must measure. E-20260825c ran every cell in both orders and found the seats taking the first passage at 0.819 with nothing disclosed, against 0.739 on a control pair that was one text against itself with a single word changed: two real translations of the same passage were barely more discriminable than one text and its near-twin. Disclosure of one fact about the source dropped it to 0.625 and made the same arm win in both orders. A tie option was offered on all 408 main calls and used zero times, so the seats do not decline; they choose, and with nothing to go on they choose position. Fires at: every A/B preference, ranking or "which reads better" design — this project has run several, and any of them with one order per cell is worthless. Remedy, and it is three parts: (i) run both orders always; (ii) buy a near-duplicate control — the same text against itself with one word changed — which is what turns the position rate into a number with a scale; (iii) report as the primary a statistic position cannot produce, such as the same arm winning in both orders, and register it in advance rather than reaching for it after the gate fails, which is what this run had to do. RS-20260825c-worth-paying §2, §6. |
S222 | S222 |
| (brm) | When an arm's designed route fails the subject rule, the ROUTE has failed and not the question — look for a declaration, a figure or a record already in print before retiring the arm. ARM-declared-function was constituted with an instruction to close retired if the subject-rule sentence for its step 1 came out about this project's raters, and it did: build a rubric, blind two or three labellers, hold out a split is a sentence about the apparatus. But the arm's question — does a translator's declared account of what a device does predict what a reader gets? — needs a declaration, not a manufactured declaration. Four printed statements of translation policy were already on the shelf, one of them (Preston 1850) carrying a falsifiable prediction about readers with a stated moderator, and a sentence printed in 1850 cannot be a rater artefact. Fires at: any arm about to close retired because its design is method work. Remedy: before writing the retirement, ask what the design was going to manufacture and whether the record already contains it; if it does, re-plan the step, write the reason and what the re-plan costs, and keep the budget. |
S221 | S221 |
| (bqb) | A frozen design that states a drop-rule twice, in a gate table and in a missing-data paragraph, will state it two different ways — and the difference is load-bearing. Write every exclusion rule ONCE, and have the analysis import it rather than re-implement it. E-20260817e said in its gate table that a pair "missing any of its 6 cells is dropped from the primary" and, two rows below in the same section, that a dead body leaves two seats and the pair is dropped only if they split 1–1. Four bodies died. Under the paragraph the primary is 7 of 9 flips, p = 0.016; under the table row it is 4 of 6, p = 0.125 — significant and not significant, from one frozen design, with neither reading chosen after the numbers existed because both were in the text before dispatch. Reporting both is the only honest exit and it costs the headline. Fires at: every design with more than one place that says which data are used — gate tables, missing-data paragraphs, failure-criteria lists, and the analysis script, which is a fourth statement of the same rule in a different language. Remedy: state each inclusion and exclusion rule in exactly one numbered clause; every other mention cites the clause number instead of restating it; and analyse.py computes the denominator from that clause and nothing else. Where a contradiction is found after dispatch, compute both, print both, and lead with the one that withholds more — a design that contradicts itself may not be allowed to resolve in the direction that helps. RS-20260817e-mimetic-reading §6. |
S204 | S229 |
| (bqa) | Striking a clause out of a standing rule does not leave the rest of the rule intact — it removes a bound, and the next hand to apply the rule has to invent the replacement silently. Say in the same result what now bounds the rule, and enumerate the other clauses the strike kills. RS-20260816h struck exclusion (a) — rhyme produced solely by an identical enclitic pronoun is not admitted — on 9 of 9 bodies, and the strike is right. But ‑هُ is the commonest word-ending in Arabic prose, and with the exclusion gone and no threshold in its place, the rhyme clause fires at almost every pair of adjacent cola: an inventory that admits everything is the same instrument as one that admits nothing. The next rendering (R38, T-kalila-qird-R38-v1 §2) had to declare a threshold — three adjacent member-ends, or two where the members also match — which is this project's own decision presented as the handbook's, and exactly the kind of decision exclusion (a) itself was. And the strike killed a second clause nobody named: the repair also instructed that the union of the form rule and the sound rule be answered, and the sound rule carries no exclusion for the cognate figure, so exclusion (b) fell too — two loci in fifty-four exist only because of a consequence the striking result does not mention. Fires at: any result that removes a clause, a filter, an exclusion or a gate from a standing instrument — the removal reads as a subtraction and is usually a re-definition. Remedy: a result that strikes a clause states, in the same section, (i) what now bounds the rule where the struck clause bounded it, and (ii) which other clauses of the same instrument the strike makes inoperative. If neither can be answered, the strike is untested and the instrument is not yet repaired. Kin to (bpy): (bpy) is about naming a category on too little; this is about unnaming one and not noticing what the name was holding up. RS-20260816j-published-figure §3, workshop/regimes/R38-repaired-inventory.md. |
S203 | S203 |
| (bpy) | A rule inferred from a single locus is not a rule, and calling it one on a result page makes it look like evidence — replicate the class before you name it. RS-20260816c §5 read three unanimous disagreements between its seats and the project's figure inventory as "three rules and not noise", and the successor tested all three on a fresh chapter. Two did not survive. Disagreement 1 — a word repeated unchanged is not sound work — rested on O19, one locus, 0 of 3; on seven fresh loci of the same class the owe rate is 0.7143, and what actually varies is the distance between the two occurrences, a variable no one had looked at. Disagreement 3 — matched shape that does not rhyme is not admitted — rested on two loci; on five it is 0.7333, and the one locus that replicated was the one where the seats reported hearing no sound at all. Disagreement 2, which rested on three loci at 7 of 8, replicated at 3 of 3. What made the two failures look solid was unanimity, not size: 0 of 3 on one locus is three bodies agreeing about one place, and it reads on the page exactly like a class. Fires at: every result page that reports a per-item breakdown and then generalises a named category out of it — which is most of them, because the breakdown is where the interesting sentences come from. Remedy, and it is a writing rule as much as a design one: when a result page names a category from its own item-level data, print the number of loci the category rests on in the same sentence as the claim, and mark any category resting on one or two loci untested. A category worth naming is worth a successor's registered prediction; a category resting on one locus is a hypothesis with a name, and the name is what does the damage. RS-20260816h-target-set §2, §5. |
S202 | S202 |
| (bpu) | A design that prices "the same content, differently marked" must first show that the two markings ASSERT the same thing — at a depictive device they do not, and the propositional constant the design needs does not exist. E-20260816e built eleven English pairs at the mimetic sites of 「坊っちゃん」ch. 2, each differing only in whether the device is enacted (shambled off, rattled and clattered, grinning and grinning) or stated (walked off slowly and heavily, were very loud on them, grinning the whole time), rebuilt once to a pre-run critic's specification. The parity gate returned SAME on 6 of 11 against a bar of 10, and a second rater left 3 of 9 standing, while both caught every planted content error (4 of 4, 3 of 3). The five verbatim reasons are one reason: "A specifies a shuffling gait while B specifies a heavy and slow walk", "A specifies the speed of the blinking while B specifies repeated blinking". A depiction commits to particulars a statement leaves open, and a statement commits to particulars a depiction leaves open; neither paraphrases the other. Fires at: every design that would price a device's marking against fluency, carriage, or reader preference — which is most of this project's carriage work, RS-20260815b and RS-20260815e included, where the same defect appeared as "three of seven device classes are not form-only". Remedy: buy the parity judgment on the actual strings before the design is written, not as a gate inside it; and where parity cannot be had, stop asking which rendering is better marked and ask instead which reading of the source is right — the two arms are then supposed to differ, and no constant is needed. RS-20260816e-mimetic-carriage §3. |
S199 | S199 |
| (bpw) | When a design asks a reader WHERE something is, the registered truth must be the place the reader is being asked about — not the place the book prints its mark. The two can differ by a whole paragraph, and then the primary measures the edition. E-20260816f registered the night boundaries of the Nights at the printed divisions — after …the court broke up, and King Shahriyar went into his palace — because that is where the copy-text and Burton print their headings. Every reader instead named the sentence where the telling stops, a hundred words earlier, and the registered primary failed at P = 0.998 in the predicted direction while the design's own ±1 tolerance showed 56 of 56 boundaries found with 0 false alarms. Nothing was wrong with the data and nothing was wrong with the reading; the truth had been defined from the typography. Fires at: every localisation, span-marking or site-recovery design — which includes every census this project runs against a translator's own frozen site list. Remedy: before dispatch, write down both candidate locations for each target and register the primary on the one the instruction actually asks for, keeping the other as a declared secondary. Where they cannot be told apart in advance, register the tolerance that spans them as the primary and the exact score as the secondary, and say why. RS-20260816f-night-seam §1, §3. |
S200 | S200 |
| (bpt) | Before building a design on "the sites where the source does X", measure whether two independent readers can agree on which sites those are — determinacy is a SEPARATE judgment from the scale, and it does not inherit the scale's reliability. E-20260816d put 70 Dickens utterances to two seats that never translated, on a 1–7 footing scale with X = cannot be determined offered as a real answer. Where both gave a number they agreed within one point on 9 of 9; on whether a number was possible at all they agreed at 0.471, one answering X at 0.871 and the other at 0.343. Thirty-seven of seventy sites are a place one reader calls determinate and the other calls silent — and the disagreement is principled, one treating a politeness token as evidence about rank and the other not. The registered control (three sites carrying an explicit sir must be determinate) failed at one of three and withheld all three primaries. Fires at: any design whose population is defined by the absence of a property — silent sites, unmarked passages, plain segments — which is most compensation and supply designs. Remedy: buy the determinacy judgment from two independent readers as a cheap pilot before the expensive arm is designed, and if they do not agree, the instruction the design was for cannot be written and the finding is that. The same shape, on a different channel, one session earlier: RS-20260816c found "the source's sound figures" is not a determinate set, the lead's inventory and three readers agreeing on 13 of 21 loci. Two framework instructions have now died at the same place. RS-20260816d-lexical-channel §4. |
S198 | S198 |
| (bnv) | A census over hands that were never asked to do X establishes that they do not do X, and NOTHING about whether X is possible — where the conclusion is a practitioner instruction, add one hand that tries. framework/v0.2 §10 carried the sentence "there is nothing to recover and no craft that recovers it" on five hands across four language pairs: two published human translators, two 2026 language models prompted without any mention of honorifics, and one lead rendering. Every one of them was translating without being asked to carry the grading, so the census could not tell English has no device from translators decline to spend one. R27-footing-max was minted to ask, on 52 mechanically extracted sites of one Japanese span: 51 carried, 1 abandoned, 0 where no device came to mind. The second clause of the framework's sentence was false, and no amount of additional non-trying hands would ever have shown it. Fires at: every result whose finding is a loss — a mark, a device, a distinction that does not reach the target language — and whose write-up carries advice to a translator. Remedy: before writing the instruction, ask whether any hand in the evidence base was trying; if none was, either add one (it costs $0 when the lead is the hand) or write the finding as no observed hand carries it and stop there. Kin to (bnu) on the other side: (bnu) says a not-X arm must be tested for X affirmatively; this says a not-trying arm cannot answer a can-it question at all. T-genji-yomogiu-R27-v1; framework/v0.2 §10.5. |
S186 | S186 |
| (bnx) | A dispatch loop that writes its results only at the end can lose everything it paid for, and the killer need not be the API — a wall clock the script cannot see will do. E-20260814h's first grading dispatch held all 72 bodies in memory and wrote run.json once, at the end. It was killed at ten minutes by a harness timeout with $0.312882500 already billed and every body destroyed — 41% of the session's spend, on a run with no truncation, no dead seat, no cap problem and no design fault. This is note (bhf)'s failure mode (spend that buys nothing) arriving through a channel none of (bng)/(bnk)/(bnr)/(bne) covers: all four of those notes are about a single call returning nothing, and this is about every call returning something that is then thrown away. Fires at: every multi-call dispatch loop, and hardest on the ones that work — a loop that fails fast is cheap to lose. Remedy, and it is three lines of code: append each body to a bodies.jsonl the moment it returns, and make re-invocation skip any (cell, seat) already on disk. The re-run then costs only what was not already bought, and a timeout becomes an interruption instead of a loss. Corollary: run any dispatch loop in the background rather than under a foreground timeout, and never size a run by assuming the wall clock is longer than the run. RS-20260814h-footing-price §9. |
S187 | S187 |
| (bnw) | Where the cost of a device is CUMULATIVE, a site-by-site price underestimates it, and the per-site coding cannot see the error. R27 requires every carried site to be coded free (C-n) or priced (C-p), judged one sentence at a time. Fifteen of 51 sites came back free — each defensible on its own. The opening paragraph, which holds three of them and nine other sites besides, says ladyship ten times and pleased to six times in 159 words, and no reader would call it English. The sum of fifteen free carriages is not free. Fires at: any instrument that scores a text property site by site and reports the total as the property's cost — device censuses, divergence-site counts, per-locus carriage rates. Remedy: report a whole-passage figure beside the site table (here: 836 → 1,012 words, +21.1%, and 78 marked deference tokens where the baseline has 0), and say on the page that the column flatters the free code. The mechanism is density, not judgment: a mark that is an inflection in the source is an added word in the target, so the same marking rate is background in one language and foreground in the other. T-genji-yomogiu-R27-v1 §The price is cumulative. |
S186 | S186 |
| (bnp) | A register manipulation is visible where the source sits at its OWN level, not where it drops below it — so a design that samples only marked sites is sampling where its own effect is smallest. E-20260814d gave three seats a plain lead pass and a light-register lead pass of the same tale. At the sites three annotators had voted below the source's ordinary level, the two arms were coded +0.0833 apart and the gate failed; at the sites voted at that level, the same two arms, the same seats and the same scale gave +0.6667. The mechanism is concrete: six of seven marked spans were shorter than four source words — at «Las!» the arms are Rag! and Rags! — and a policy that raises by one step has nowhere to go inside a one-word shout, while over ordinary prose it has a clause at a time to act in. Fires at: every design whose site list is selected for source markedness, which on this project is every register design since RS-20260808e — whose 47 elevation sites are all low sites. Remedy: carry neutral sites as a first-class stratum with its own registered prediction, not as a construction check with n = 4; and where a gate is computed on marked sites, report the neutral figure beside it before concluding anything about the instrument's resolution. RS-20260814d-elevation-resolution §6. |
S183 | S183 |
| (bnq) | A majority vote over three annotators does not average out disagreement that is DISPOSITIONAL — check whether the splits share a direction before treating the majority as a measurement. E-20260814d put 27 source spans to three seats for a below/at/above call and got base rates of 7, 20 and 3 and pairwise agreement of 0.5185 / 0.5926 / 0.3704. Seven spans had no majority — and all seven split the same way, P1 above, P2 below, P3 at, seven times out of seven. Those are three fixed dispositions, not seven independent draws, so the majority rule has nothing to average: it silently deletes exactly the middle of the range (every no-majority span was a speech turn of 10–29 words, while all seven narration spans were unanimous and the three unanimous below spans were all single-word interjections). Fires at: every multi-annotator site list, census or vote in this project, including RS-20260808e's three-annotator majority which every low-register figure rests on. Remedy: report the per-annotator base rate and the direction of every non-majority split alongside the majority; where the splits share a direction, the majority is a statement about which dispositions happen to outnumber which, and the design says so rather than reporting n sites. RS-20260814d §4. |
S183 | S183 |
| (bnh) | A forced-choice question that GLOSSES the property under test in the words of the features the arms were built to differ on is not a measurement of that property — it is a test of whether the seat can apply a handed-over rubric. E-20260813h asked which of two renderings followed the original's own way of putting things, and glossed the phrase as "the way it builds its words, and the order in which it arranges its clauses" — a verbatim description of the two features the census had measured and the source-ward arm had been written to maximise. An independent pre-run critic called it BLOCKING before a cell was dispatched: a firing primary would have shown only that the seats could pattern-match a supplied definition, and the rival hypothesis never had a fair run. The gloss was struck from all three conditions and the bare construct name left standing; the primary then fired at 15 of 15 anyway, which is what makes the amendment worth having rather than a near miss. Rule: name the construct, never its operationalisation, in any prompt whose answer the operationalisation predicts — and if the construct is too vague to be answered without the gloss, that is a finding about the construct and not a licence to supply it. The critic's rejected fallback is part of the rule: keeping the gloss in a gate condition and striking it from the primary would make the gate a gate on a different question. Companion to (bhg)(ii), which forbids showing a blind assigner the strata; this forbids showing it the definition. RS-20260813h §1, design §12 finding 1. |
S179 | S179 |
| (bni) | Cells from a FIXED seat are not independent draws, so an exact binomial over pooled cells is the wrong arithmetic: register the seat as the unit of analysis and report the pooled figure as the clustered figure it is. E-20260813h registered its primary as ≥ 12 of 15 cells with an exact binomial of 0.0176 — 3 seats × 5 segments, treated as fifteen Bernoulli draws. Seats are chosen, not sampled, and five cells from one model share that model; five cells from one segment share a text. The pre-run critic named it BLOCKING and the primary was re-registered before dispatch as each seat ≥ 4 of its 5 segments (per-seat exact P = 0.1875, three-seat conjunction 0.0066 under independence, which three models sharing a training distribution do not have), with the 15-cell count kept as a described secondary. Rule: any design whose cells are crossed with a fixed panel registers the bar at the seat level, states that the inference is to those seats and not to a population, and never reports a pooled per-cell p-value without the word clustered beside it. The companion consequence, and the reason it is worth a note rather than a line: a strict per-seat bar can be failed by one seat while the pooled count looks decisive — which is exactly what happened to this run's CB probe, 13 of 15 pooled and reported as FAILED because one seat returned 3 of 5. This note prescribes future conduct and asserts nothing about past figures; whether any published per-cell binomial in this project is affected is unchecked, and checking it is method work that needs a blocked deliverable to be worth a session. RS-20260813h §3, design §12 finding 2. |
S179 | S179 |
| (bnc) | A run-length instrument can only price a declared exposure where the UNEXPOSED baseline is non-zero — measure that baseline before building the design, not after it fails. E-20260813d was built to price ARM-dakghar's span-A priming event: the lead had read a published English through unit [16] before translating, the boundary was known, and the arm's constitution said the known boundary is what makes it measurable rather than merely admitted. Span A was re-rendered whole by an unexposed pass and both renderings measured against the comparator. Result: zero shared twelve-grams between this lead and this published hand over the whole 1,147-word span, in either pass, exposed portion and unexposed portion alike, with longest runs of 5 and 6 tokens on the exposed stretch — inside the design's own declared coincidence band, so P1 was NOT READ. An exposure cannot raise a channel whose floor is already zero. The day before, on «Flipperne», the same instrument had a 19-token baseline to work against and measured something (RS-20260813c §4). Fires at: every design that prices contamination, priming, or influence by n-gram or run-length overlap against a named comparator. Remedy: run tools/dependence_check.py on the unexposed material against that comparator first, as a feasibility gate beside the selection gate the standing contamination rule already requires; if it returns zero twelve-grams and a run in the coincidence band, the design cannot answer and must not be built — say so and choose different material or a different measurand. Sibling of (bmw), which made the comparator's existence a selection gate; this makes its measurable distance one. RS-20260813d-first-span-again-dakghar §3. |
S175 | S175 |
| (bmx) | Where the question's answer is a token count, a dose ladder measures the tokens and nothing else — before laddering a manipulation, ask what the two extreme forms differ in BESIDES the dose. ARM-dose step 2 was written to ask whether a story's foreignness degrades gradually as more of its culture-bound items are Anglicised, and registered DA vs D1 on "which version is more clearly set outside the English-speaking world?". The pre-run critic's BLOCKING 1: "the model will trivially pick D1 because it literally contains foreign words, measuring only its ability to see non-English text." On that question "more domesticated" and "fewer foreign tokens" are the same variable, so no ladder built from D0…DA can separate them — which is also why the prior run's K7 and K8 both came back 0 of 30 and could not be read as a dose. Fires at: any design that ladders a manipulation and asks a question the manipulation's own by-product answers. Remedy, and it is what made the run worth dispatching: build the comparator that holds the by-product constant — here a sixth form NA, every site rendered in location-free English, so that DA and NA carry zero source words each and differ only in the kind of English that replaced them. That contrast is a measurement; the ladder was a count. Sibling of (bmv), one session older and the same disease: a primary made true by the construction of the materials rather than by the world. E-20260813a-world-dose/design.md §A2.1, critic.md finding 1. |
S172 | S172 |
| (bmv) | A prediction that the target carries a source's LEXICAL marker is a prediction that translators translate words — it is near-circular and must not be registered as a test. E-20260812i registered P1: at source places where the footing mark rides on a word, the English carries a device; where it rides on morphology alone, it does not. The pre-run critic's BLOCKING 1: "the English translator must render the word's propositional content, which inherently results in a 'carried' lexical item… P1 is true by construction of the translation task rather than by measurement." Observed afterwards at 15 of 15 and 10 of 10 in every hand of both texts — a ceiling, as predicted, and evidence of nothing. Fires at: any design that scores whether a translation "carries" a feature the source expresses in a word, phrase, epithet, title or formula. Remedy: score the direction that is a free choice instead — whether the hand SUPPLIES a device where the source had no word to force one (the REL label). That number was 22 of 30 where the near-tautology was 25 of 25, and it is the one the result rests on. Kin to (bkd): a critic can only save the predictions a design has not already made true by construction. RS-20260812i-footing-channel §4, E-20260812i/amendments.md §A1. |
S171 | S171 |
| (bmw) | Confirm the published comparator EXISTS, in reach, before choosing the work the translation limb will render. ARM-footing chose Chekhov's «Унтер Пришибеев» for its footing density, translated it whole in two regimes, froze both, wrote and critic-passed a design naming GAR as the primary hand — and only then discovered that no free English of that story exists: not in Garnett's thirteen volumes (the 233-story Gutenberg compilation), not in four other free collections. F4 fired and a second story had to be added mid-run to recover a published hand, at the cost of a comparator the lead had already read. Checking cost one grep and would have taken sixty seconds at selection. Fires at: every unit whose design compares the lead's hand against a published one. Remedy: the comparator's presence is a selection gate alongside the contamination measurement, run in the same minute and on the same footing — locate the actual text, count its words, and record where it was found, before a word of the source is translated. RS-20260812i-footing-channel §5. |
S171 | S171 |
| (bmi) | A build or gate script that PRINTS RAW COUNTS destroys the pre-registration of every prediction computable from them. Print checksums, word counts and structure — never the measured quantity. E-20260812's gate G1 required each container's raw punctuation forms to be enumerated before normalisation, so build_materials.py printed them; that table contains every primary mark total for every text, which is sufficient to compute all four registered predictions before the analysis script exists. The pre-run critic's BLOCKING 1: "Calling the subsequent arithmetic 'unseen' does not preserve confirmatory status." Three of the four predictions were retired and the census was reported as descriptive. The gate was right and its output was the leak — a verification step and a peek are the same act when the thing verified is the thing measured. Rule: any script that runs before the analysis writes its diagnostics to a file the designer does not open, or prints only quantities the predictions do not depend on; if a gate genuinely needs a human to look at the data, the predictions that data bears on are retired in the design before looking, in writing. Companion to (bkd): the critic can only catch what the design still has left to lose. RS-20260812-unlicensed-typography §Preamble; workshop/experiments/E-20260812-unlicensed-typography/critic.md finding 1. |
S163 | S163 |
| (bmj) | A paragraph-grain co-occurrence count is not a correspondence measurement, and naming it licensed writes a false sentence. E-20260812b froze a design in which a target emphasis mark counted as licensed if the aligned paragraph held any source mark — so a French italic on a book title would have scored as licensed because Poe italicised a different word in the same paragraph. The pre-run critic's BLOCKING 1 refused it, and the rebuild at span level was not a refinement: on the same corpus the paragraph proxy and the span measure are different quantities, and the design's own sentence — "the fraction of the marks that hand's reader sees which are licensed by a source mark in the same place" — was untrue of the number it named. Two rules follow. (i) A statistic whose name contains a spatial claim (in the same place, at that site, licensed) must be computed at the grain the name asserts, or renamed to the grain it was computed at — paragraph_cooccurrence, not licensed. (ii) Directional-bound language does not survive an unvalidated aligner: retention is an upper bound only if the alignment is right, and a length-based aligner that drifts can inflate as easily as deflate. RS-20260812's 0.9425 is not made false by this — it declared its grain and its bound — but the size of the gap this run found between the two grains is the reason to stop treating paragraph grain as a cheap approximation of site grain. |
S164 | S164 |
| (bmk) | Inventory the source's emphasis DEVICES before counting one of them. E-20260812b counted italics and would have scored Sasaki's _天邪鬼_ and _絞首台_ as unlicensed additions: they render Poe's PERVERSENESS and GALLOWS, small capitals, which the English witness represents by capitalisation and which no _-based counter can see. Neither French nor Japanese has a small-capital device, so both hands convert into their own emphasis channel — and a census restricted to one channel reads a conversion as an invention. Found by reading the unlicensed target spans rather than by any check the design had. Rule: before an expressive-mark census, list every device the source uses for that function — italics, small capitals, spaced type, 傍点, capitalisation, quotation — and say which of them the counter can see. Sibling of (bcg), which is about markup an extractor destroys; this is about a device the transcription never encoded as markup at all. |
S164 | S164 |
| (blk) | One non-lead arm cannot establish that a regime effect is ABSENT, and note (bkz)'s remedy is insufficient where it is read as a null. (bkz) prescribes at least one arm of the regime executed by a hand that is not the lead, and RS-20260809b supplied exactly one: +0.024 and +0.167, neither interval excluding zero, from which S141 published that the programme tax is "a property of the lead translating under rules, not of rules". E-20260809h put six independent hands from six labs on the same statistic and got +0.98 — larger than the lead's own +0.714 in the run that drew the conclusion — with five of six strictly positive, and ranked the lead 2 of 7 rather than an outlier. The published null was underpowered, not informative. Rule: (bkz)'s one arm licenses a POSITIVE finding about a regime and never a null; a null about what a regime does needs a plurality of non-lead hands, and the design says before dispatch how many it would take. The sibling caution is that more hands do not rescue the endpoint: P2 still missed at P = 0.0625 on a registered 0.05 bar, and the bar was not moved. RS-20260809h-rule-execution §3, §11. |
S146 | S146 |
| (bkc) | A divergence-SITE COUNT does not distinguish one hand from two: report how much text the sites swallow beside how many there are, or the count means nothing. E-20260807c aligned one 650-word chapter across four independent hands and against the lead's own blind self-rendering of a different work (RS-20260806d). Sites per 100 source words: one hand against itself blind, 20.06; two different hands, 23.2 to 29.1 — a factor of 1.16 to 1.45, and the registered 1.5x threshold FAILED at 6 of 6 cells. On the same alignments, agreement separates the two conditions cleanly: 0.814 within a hand against 0.317-0.650 between hands. The two statistics come from one SequenceMatcher and disagree because a count measures how many places the renderings part and agreement measures how much text is inside them: one hand parts from itself in nearly as many places, and the places are a fifth the size. This bites a published figure — ES-20260806-craft-report-koyhaa-kansaa §3.2 and RS-20260806d are built on 234 divergence sites, and a count of that kind cannot separate one translator from two. Rule: any design reporting divergence sites reports a size or agreement statistic on the same alignment, and no argument rests on the count alone. Companion to (bkb): the count was not weak, it was orthogonal to the question. RS-20260807c-two-hands §3. |
S127 | S154 — complied with rather than caught. RS-20260810x reports 160 divergence sites and, on the same alignment, pooled agreement 0.734 and 11.59 sites per 100 source words, set beside this note's own two reference conditions; no argument on that page rests on the count alone. |
| (bkd) | A pre-run critic's REMEDY is not covered by the critic's authority — check it against the design's own manipulation before accepting it. E-20260807d accepted all twelve findings of a NEEDS-REDESIGN pass, and one of the remedies destroyed a screen. ADVISORY 10 correctly said the parity screen's probe word register was undefined; its proposed replacement was does either passage imply that the speaker is socially above, below or level with the person addressed? — which is exactly and only what the treatment does. MARKED differs from BARE in nothing else, so the screen failed 12 of 12 by construction and F4 withheld the primary. The finding was right and the remedy was wrong, and the design accepted it in the same motion. Rule: for every accepted remedy, write one sentence saying what the design's treatment is and check the remedy does not test for it; a screen that must fail is not a screen. The same twelve bodies, read as a manipulation check rather than a screen, passed at ceiling — two blind seats named the difference as social standing at 11 of 12 pairs — so the information was there and the gate was mis-aimed. Companion to (bjm): a gate the run's own design forces into existence must still be aimed at something other than the run's own variable. RS-20260807d-marking-work §4.4. |
S128 | S128 |
| (bke) | A cross-sense specificity criterion stated as an absolute scale-point ceiling penalises a strong dose and rewards a weak one. State it as a RATIO to the on-target effect, and report the absolute form beside it. Tier D's §6.6 condition 3 is off-target drop ≤ 0.75 scale points, and it is the one number Tier D has ever failed (+1.12 at the heavy dose, S086; +0.50 and passing at the light one). E-20260807e's pre-run critic found the structural reason before any number existed: the same 0.75 is used as a detection FLOOR and a leakage CEILING, so a control that barely clears detection is likely to fail specificity, and the criterion is easiest to pass when the operator does almost nothing. The design registered the ratio form instead and measured both. The absolute form FIRED (0.833 > 0.75) and the ratio form did not (0.833 is 21% of a 3.944 on-target drop). Rule: any design importing Tier D's condition 3 states which form gates and reports the other; a bare absolute number is not comparable across doses. This does not reopen S086 or S113 — no published figure changes — and it is the criterion's operating characteristic, not a defect in any run that used it. RS-20260807e-sense-tradeoff §8. |
S129 | S129 |
| (bkm) | A prediction of NO DIFFERENCE that is tested by failing to reject one is not a result. Register an equivalence margin and report the interval — and expect the design to fail it. E-20260808c registered R1 as |Δ| < 0.75 and permutation P > 0.05, which is the absence-of-evidence error with a threshold on it; the pre-run critic named it as BLOCKING 1 before any number existed, and the design replaced it with a two-one-sided-tests criterion: the 90% permutation confidence interval must lie entirely inside ±0.75. The measured Δaccuracy between the two Venuti ladders was exactly 0.0000 with P = 1.000 — and the interval came back [−0.777, +0.666], so the registered criterion FAILED by 0.027 of a scale point on the cleanest possible point estimate. Under the original wording it would have been reported as a pass. Two rules. (i) Any design whose claim is that two things do not differ registers an equivalence margin, computes an interval, and states the margin's justification before dispatch; "P > 0.05" may not gate a null. (ii) Size the design to the margin, not to the effect it hopes not to find: seven segments carry roughly ±0.7 of precision here, so a ±0.75 margin was unreachable from the start and the arithmetic was available in advance. The run's null survived only because a second, independent ladder was registered as a co-primary and returned 0.000 with an interval of [−0.166, +0.166] — so the third rule is that a null worth publishing needs a second measurement, not a bigger p-value. RS-20260808c-sense-tradeoff-de §3. |
S134 | S134 |
| (bkn) | To decouple two properties that always travel together in your materials, pick materials by a criterion written down in advance — and the criterion that worked is: a source whose MARKED devices are the devices the target's UNMARKED register is made of. RS-20260807f needed to separate source form carried from markedness in the English and could not: on Schwob, carrying the litany marks the English, and its limit 2 concluded the project "does not have one on the shelf". It had not looked with a criterion. The one that worked: read wiki/goodness-senses.md's own unmarked / literary-contemporary anchor A-mchugh-presence as a shopping list — invisible free indirect discourse, say-only dialogue tags, texture by named particulars, structural one-sentence paragraphs — and find a source built out of those. Bang 1886 has all four, and carrying his sixteen devices cost +0.095 of naturalness where Schwob's cost −1.444, with the same scale moving −2.524 on the control arm. The rule: when a design's validity turns on separating two properties, name the property that must NOT move, find the target-language description of its zero point that the project already owns, and select the source against that description before reading the passage for translation. RS-20260808d-carriage-decoupled §8. |
S135 | S135 |
| (bkp) | A content-parity control must compare the arm against the SOURCE. A control that compares your arm against ANOTHER TRANSLATION cannot tell your arm damaging something from the other translation adding something — and published translations add freely. E-20260808e built a parity gate on the question do A (the published English) and B (the low-instructed arm) state the same content?, with a 25% non-equivalence bar. It came back at 16 of 41 = 0.3902 and withheld every primary in the run — and the reason was substantially the published side: these hands expand, Eh bien ! la mère, et c'te santé, toujours bonne ? becoming "Well, Mother, you are always pretty well and hearty, I am glad to see." The same call's second, source-relative question — does B state or omit something against the SOURCE? — came back at 9 of 41 = 0.2195 and would have cleared the same bar. The gate as registered is the gate that fired and nothing was rescued. Rule: a parity or equivalence gate names the SOURCE as its comparand, and any translation-against-translation figure it also produces is descriptive. Corollary, from the same call's third question (amendment A6): an arm instructed to move a scale reaches for things that answer to nothing in the source — the low arm added an oath, intensifiers, a tag and three epithets at 13 of 41 sites — so count additions explicitly whenever an arm is told to move in a direction. RS-20260808e §6.1–6.2. |
S136 | S136 |
| (bkr) | A content-parity gate cannot license a register claim about register-marked sites, because at those sites nothing passes it — including the published hands the claim is about. ARM-low-pole was blocked in both its sessions by a content-parity control, under two different phrasings, the second being the repair of the first: E-20260808e's translation-against-translation question failed at 0.3902, and E-20260808f's source-relative replacement — stricter in kind, applied to five arms, with a calibration on the published hand that the first never had — failed too, at err(LOW-A) 0.7021 against err(PUB) 0.6809, one site in forty-seven. The calibration is the finding: put through the identical question, the four published hands of 1890–1918 are flagged as omitting or misstating at 0.6809 and as supplying at 0.7660, the highest supply rate of any arm in the run, above a 2026 model told to write low (Craig's "To pull a good oar", "our urchin"; Morri's "Say, you big bluff … ha ha"). And the two questions proved inseparable in the instrument: 31 of PUB's 32 err flags also carry an add flag, 13 of the err strings beginning "adds …". Rule: where a design needs to know that an arm did not damage content, at sites the source itself has marked, a propositional-equivalence bar is the wrong gate — it measures translation, not damage. What to use instead, unresolved and stated as such: a gate on the specific relation the claim needs (a named participant, event or number preserved), not on global equivalence. RS-20260808f-placeless §2.2, §2.1. |
S137 | S137 |
| (bkq) | After applying a critic's amendments, re-read the gate set for a gate that has become IDENTICAL to the prediction it gates. E-20260808e froze G2 (a reachability floor) and P3 (the claim that the floor is reached in narration) as different quantities. Two accepted amendments — A1, restating every primary as a within-cell contrast, and the remedy to BLOCKING 7, restating G2 at N sites against the same comparand — made them the same arithmetic statement, so G2 could no longer gate anything: a design in which the gate and the claim are one expression has, at that point, no gate. Neither the lead nor the critic saw it, because the critic had already been run and each amendment was checked only against the finding that produced it. Rule: amendments are applied as a SET and the gate list is re-read once, end to end, after the last one — asking of each gate what it is a function of and whether that is still different from what it protects. RS-20260808e §11 limit 3. |
S136 | S136 |
| (bko) | An edit is not "unlicensed" because you did not intend to license it. Write down what the SOURCE does at that site, per edit, in the materials file — and the risk scales with how close the two languages are. Twice now a manipulation meant to add markedness answering to nothing in the source has been caught adding a calque instead: RS-20260807f's pre-run critic found that all six odd prepositional government edits copied the French preposition at the site (a third of that run's factor, replaced before dispatch), and E-20260808d would have done the same with fronted-adverbial inversion, because Danish is V2 and English inversion is its ordinary syntax. Neither was visible from the English alone. The rule: every unlicensed-markedness edit carries, in the materials file beside the edit, a written statement of what the source does at that site, and the statement must read PLAIN; and the operator TYPE is checked against the source language's ordinary syntax before any site is chosen, not after. The check is cheap and mechanical and it is now a verify.py assertion. RS-20260808d §2, materials/arms.py. |
S135 | S135 |
| (bha) | The front-matter vocabulary cannot distinguish a page that JUDGES on a sense from a page whose question is ABOUT that sense and which judges nothing. A pre-run critic found senses: [naturalness, style-correspondence] on a design whose own text says it makes no quality claim anywhere, and read the field as smuggling one back in. It is not: CLAUDE.md rule 3 requires the field on evaluative pages and says nothing about the others, and the ids the critic proposed instead (lexical-dependence, carryover) are not in wiki/goodness-senses.md, so the prescribed remedy would have breached the rule it was meant to serve. The gap is real and the remedy is not a new sense id: a reader has to reach the design's last section to learn that nothing was judged. Until the schema has a way to say it, a page that carries senses: and makes no evaluative claim must say so in its opening paragraph, where the front matter is read. E-20260801e-lead-carryover §13 A5. ⚠ Id collision: (bha) also names the S080 audit-is-a-write note in §Live; resolve a citation by its subject (S084). |
S081 | S081 |
| (bhf) | [FIRED S162, and this time the LEAD broke rule (iii) itself. moonshotai/kimi-k3 was chosen as pre-run critic BECAUSE it is not one of the three judging seats, and returned null content with 2,497 of a 2,500 cap on hidden reasoning. The lead then raised the cap to 6,000 — the one thing this note forbids — and got null content with 5,997 of 6,000. $0.053388 + $0.100358 = $0.153746 for no body, 15.8% of the session's spend, the second half of it avoidable by reading this note. Same slug as S106, S123, S127, S128. The pass that returned came from openai/gpt-5.6-terra, which IS a judging seat in that design, so the run has no critic–jury independence guarantee and says so on its own design page rather than in the result. RS-20260811h §7.2.] [FIRED S126, FIFTEENTH time, and it falsifies this note's own S123 remedy. nvidia/nemotron-3-ultra-550b-a55b — the slug (bhf) itself recommended at S123 as the rule-(iii) replacement, on the strength of a 15-finding critique for $0.0165 — returned finish_reason: length with zero content characters as pre-run critic, 16,000 of 16,000 completion tokens on hidden reasoning, $0.05976540. Findings were salvaged from the reasoning field and were degraded (repetitive, broken numbering). Rule (iii) applied a second time in the same session: qwen/qwen3.7-max, non-panel, returned a clean six-finding NEEDS-AMENDMENT for $0.04985647, and its BLOCKING 1 caught a real content confound in the session's own gate that the first pass had missed. So no slug is a standing safe critic seat, and one critic pass is not adversarial coverage when the seat can return nothing: budget for two, on different labs.] [FIRED S123, THIRTEENTH AND FOURTEENTH times, in one session, on two slugs, in two roles, and BOTH were foreseeable from this note. moonshotai/kimi-k3 as pre-run critic: finish_reason: length, zero content characters, 16,000 of 16,000 completion tokens on hidden reasoning, $0.2414856 — the same slug that had already failed this exact role at S106, chosen because the design did not read this note before naming its seats. Rule (iii) applied and nvidia/nemotron-3-ultra-550b-a55b returned a 15-finding critique for $0.0165, one fifteenth. Then google/gemini-3.6-flash as a PROFILING seat: length with the JSON truncated mid-object on 4 of 4 source calls at a cap of 1,200, of which 1,104–1,152 were reasoning tokens, $0.0416715 wasted. There the remedy was note (bhq)'s — the cap was raised for that seat alone, sized from the measured appetite, and 12 of 12 subsequent calls returned stop. So (iii) and (bhq) are not in conflict: change the seat when it is a role you can re-cast, raise the cap when the seat is a registered member of the panel. 28% of the session's spend bought nothing.] [FIRED S116, TWELFTH time and in a NEW MODE: z-ai/glm-5.2 at effort high returned NO BODY AT ALL in 600 seconds as pre-run critic — not a length body, not empty content, no response — and was killed by the wrapper. The key-usage delta across that killed request is exactly 0.000000000: a request killed before its response returns bills nothing, which is worth knowing before anyone raises a ceiling to rescue one. Rule (iii) applied as written and the SEAT was changed: deepseek/deepseek-v4-pro returned NEEDS-AMENDMENT in 636 s for $0.02631576. The same run also informed P2's ceiling from this note before dispatch — 20,000 rather than S111's 12,000 — and P2 returned clean on both orderings.] A seat's reasoning-token appetite is a property to MEASURE before it is given a load-bearing role, and max_tokens is the wrong instrument for measuring it. E-20260802-voice-warrant gave moonshotai/kimi-k3 a registered control — the independent home-assignment its agreed-subset rule depended on — and the seat returned zero characters of visible content on three completed dispatches: 5,997 of 6,000 reasoning tokens, then 5,997 of 6,000 again, then 19,997 of 20,000, finish_reason: length every time. $0.5030868 on a seat that produced no data — more than every other call in the run put together. Note (abc) says build the worst case from the cap; the missing half is that on a reasoning seat the cap bounds the bill and bounds nothing useful, so raising it is not a diagnosis. Two rules. (i) Before a registered control depends on one seat, spend one cheap probe on that seat with that task shape; a seat that has never returned content on the shape is not the seat a failure criterion should hang on. (ii) Write the fallback into the design before dispatch — here amendment A4 declared the two-way subset in advance, which is the only reason abandoning the seat after the third failure was a decision and not a rescue. Companion to (abc), (b) and (bdl); (bdl) says change the parameter not the seat, and this is the case where changing the parameter three times taught nothing. [fired S113, ninth and tenth times, on two more distinct slugs: z-ai/glm-5.2 burned 11,595 then 32,436 reasoning tokens and returned empty content at caps of 12,000 and 32,000 ($0.114 for nothing), and moonshotai/kimi-k3 burned 7,997 at a cap of 8,000 ($0.130). The remedy that worked both times was NOT a raised cap but a CHANGED SEAT — a labelling task went to mistralai/mistral-medium-3-5 and answered first time for $0.0064. Rule (iii), new: after one dead body from hidden reasoning, change the seat rather than the ceiling; a seat that spent its whole allowance thinking will spend a bigger one thinking.] Prior firing record, preserved: S111 — fires again on google/gemini-3.6-flash: two length bodies with zero content at the run's own 12,000 ceiling on one ordering while the SAME seat returned a complete body at the SAME ceiling on the other, $0.174 wasted, a quarter of the run. A $0.000456 stage-0 probe had reported all five seats "alive" on a three-character reply — liveness is not appetite, and a probe that does not use the real payload shape measures neither. · S087 · S106 — fifth firing, same slug (moonshotai/kimi-k3), second on capacity: $0.2128698 for zero content characters as pre-run critic; replaced by z-ai/glm-5.2, which returned NEEDS-REDESIGN for $0.0377, one sixth of the failure. · S108 — SIXTH AND SEVENTH firings in one session, and the worst yet in aggregate. moonshotai/kimi-k3 again: $0.2432 for two zero-content bodies as a fork classifier at cap 6,000, and $0.1564 for a length body as a ratification voter at cap 8,000. google/gemini-3.6-flash then produced the same shape twice at cap 6,000 and qwen/qwen3.7-max four times, so it is no longer one slug's property. Total wasted on zero-content and length bodies: $0.553. What finally worked was a seat with no reasoning appetite (nvidia/nemotron-3-ultra-550b-a55b, cap 20,000, $0.0157) — i.e. rule (i), the cheap probe, is still the unpaid remedy.** |
S087 | S128 — SEVENTEENTH firing and the first time rule (iii) FAILED: the recognition seat moonshotai/kimi-k3 returned zero content characters at 2,500 and again at 12,000, the seat was changed per rule (iii) to z-ai/glm-5.2, and the replacement returned zero content characters at 12,000 too. Three dead bodies, two models, one short JSON answer; the measure was non-gating and was left resting on the single seat that did answer. So rule (iii) is not a repair, it is a second draw from the same distribution — and a run should decide in advance which of its measures may end on one body. Prior: S127 — sixteenth firing, on moonshotai/kimi-k3 as a JUDGING seat (zero content at cap on one block, 199 characters on another, $0.252 wasted) while two other seats cleared the same cap on the same payload. Rule (iii) applied to that seat and note (bhq)'s cap raise applied to a different seat in the same run, on the distinction (bhf) itself draws: change the seat that has not performed the role cleanly, raise the cap for the seat that has. |
| (bhg) | A stratum label the lead wrote is the lead's opinion until an independent seat reproduces it, and a census figure of the form "sense X is the home of n of N decisions" is ONE CODER'S. E-20260802-voice-warrant declared 24 translation decisions to voice / accuracy / style-correspondence, then put the same 24 items and the seven sense definitions verbatim — no design, no hypothesis, no strata — to an independent non-Anthropic seat in a stateless call. Agreement: 9 of 24 (37.5%), with naturalness, a sense the strata excluded, taken on six items. The design's registered floor was 16 and the primary was withheld. Two further measurements travel with it and neither is optional: re-cut on the independent seat's own labels, neither axis separated any pair — so the effect visible on the lead's strata is not visible on anyone else's; and the same seat, shown the design (whose §3.2 tabulates the strata), matched the lead on 15 of 24 and assigned voice to 11 items instead of 6, so the framing moved a classifier by six items. Rules: (i) any design that strata-fies by lead-declared sense carries an independent reassignment and a registered floor, and the floor must be able to withhold the primary; (ii) never put the strata in a document a blind assigner is shown — the pre-run critic caught exactly this here, as a BLOCKING finding, and it was a defect not a hypothetical; (iii) every published "home of n of N" figure states whose homing it is. Sibling of (bfq) and (bdq): reachability and reliability are properties of the individual reader, and so is the label. |
S087 | S087 |
| (bhh) | A linear congruential generator's LOW bits have period n, so lcg % k is not a shuffle and not a resample. E-20260802-voice-warrant's bootstrap CI, written to satisfy the pre-run critic's finding 3, drew indices with s % len(vals) from an LCG mod 231 and returned a degenerate interval — [+0.292, +0.292], the point estimate twice, over 10,000 resamples. It looks like a confidence interval and is a constant. The same weakness sat in the permutation shuffle, where it is less visible because a bad shuffle still produces plausible p-values. Fix: take high bits ((s >> 15) % k). Rule: for any deterministic resampling written to stdlib — and this project writes them often, because Math.random is barred and seeds must be written into the design — use the high bits, and sanity-check that a computed interval is not a point. The general form: a statistic added to satisfy a critic still needs its own smoke test**; this one was added, printed, and would have been quoted on a result page unread. |
S087 | S087 |
| (bhd) | A control's decision rule needs a POWER figure before it is adopted, not only a false-positive figure — and this project has now paid for that twice and been repaid once. S040 found that Tier D's sham band had never had its power computed: 5-of-6 on 6 units catches a genuine 70% edit-presence bias 0.4202 of the time, so a sham that stayed silent said almost nothing. S083 ran the repaired form — 9-of-15, 0.8689 power at the same branch probability (0.004193 against 0.004639) — and its silence is therefore a measurement. Rule: every firing rule a design registers reports (i) its exact null probability and (ii) its power against a stated alternative, both computed by enumeration, before the rule is adopted. A rule that reports only (i) is a rule whose failure to fire means nothing. RS-20260801f-tierD-stages12 §1. |
S083 | S083 |
| (bgi) | An overlap statistic between two translations is not register-neutral: how much two INDEPENDENT translators converge depends on the register they are aiming at. Three seats at three labs rendered one 285-word Hungarian passage from the source alone, twice each, under two frozen register catalogues. Under the unmarked-contemporary target the three pairwise longest common runs are 16, 16 and 14 tokens (28, 31 and 24 shared 7-grams); under the period-idiomatic target, on the identical source and seats, 5, 9 and 8 (0, 10 and 3). Roughly double, with no seat having seen any other's output. The consequence is not about any one figure but about reference distributions: a distribution built from published pairs silently fixes a register, and a threshold read off it is a threshold at that register. It clears no run and reopens no verdict. RS-20260731h-carryover §5. |
S076 | S076 |
| (bgf) | A contamination floor certified by the lead is the lead marking its own homework; a floor built from texts the lead did not write cannot be gamed. A recall probe the lead writes to show it does not recall a comparator can pass by avoiding the comparator's wording, and avoidance is invisible in the statistic. The deterministic alternative is usually already in the materials: compare the comparator with other texts of the same kind, period and provenance that no lead text touched. At S075 the same 1907 anthology supplied fifteen — different translators, five source languages, one publisher, one year — and they share with the comparator a longest run of median 3, maximum 5, and zero 7-grams in 15 of 15 cells. That null killed the "period English resembles period English" explanation for free, and it also caught a measurement artefact: the publisher's identical per-story imprint line put a spurious 11-token run into the null until it was stripped. Strip the boilerplate before measuring a null over an anthology. | S075 | S075 |
| (bel) | A selection gate may not be run on the outcome variable of the study it gates. CLAUDE.md's standing rule is that the lead's contamination on candidate material is measured before a locus is selected. Where the study's own statistic is that measurement — as in ARM-revision step 2(b), which compares a draft's and a revision's longest shared run with a published comparator — selecting the locus on a low run conditions the sample on the dependent variable, and the resulting figure is drawn from loci pre-chosen for the value being compared. The gate is then correctly not run, and the reason is written on the artifact rather than the omission confessed. What must still be preserved is the thing the gate was protecting: the translator does not open the comparator. workshop/translations/niewola-tatarska/R04-v1/translation.md §Contamination. |
S062 | S062 |
| (bgj) | A declared reserve is a ROLE ASSIGNMENT, and it must be checked against every role already assigned in the design — including the design's own critic. E-20260731h declared qwen/qwen3.7-max as the fall-through seat for one subject and had already used it as the pre-run critic. The fall-through fired, and three translations were produced by a model that had read the frozen design and knew what was being measured — the S053 role collision, arriving through a table the critic was never shown. The critic reviews design.md; the reserve chain lives in run.py. Two remedies, both cheap: state the reserve for every seat in the design page, where the critic can see it, and check the reserve list against the critic and rater lists before dispatch. $0.1501180769 voided, 26% of that session. |
S076 | S076 · S077 — closed by construction rather than by attention: run.py carries an assert that the critic slug and the critic reserve appear in no rater reserve table |
| (bei) | When two rules share a clause, a set-overlap statistic between them is inflated by the shared clause and can be near-unfailable. E-20260730-grain-clause registered "Jaccard between the sites where each rule's test 3 fires, ≥ 0.80" as its warrant-preservation measure. All three rules under test share test 3's first limb verbatim, so every established-borrowing site sits in every rule's firing set and every pairwise Jaccard is dragged toward 1 whatever the second limb does — the criterion could barely fail. The pre-run critic found it and named the fix: measure flips relative to the baseline rule, not overlap — the set of sites where the variant's decision differs from the baseline's. Flips are zero when a rewrite preserves a clause's extension, they grow as it departs, and the shared clause contributes nothing to them. Whenever a comparison is between near-identical objects, state which part of them the statistic can see. |
S061 | S061 |
| (e) | Write revision triggers naming the observation that would matter, not the outcome you expect. | S017 | S017 |
| (o) | A threshold must be reachable by the thing it is applied to. | S020 | S033 |
| (abd) | A rule stated as an "iff" and a veto stated below it are two rules, and they will contradict each other. A condition that can void a claim belongs inside the iff, as a numbered condition, or there is an outcome with no assigned reading. | S033 | S033 |
| (bcz) | Repeat the control, not only the treatment. A design that measures its baseline once has no estimate of the baseline's own noise and cannot bound the effect it reports. S049 sent one condition twice, byte-identical, at temperature 0, and the two measurements — 0.575 and 0.750 — straddled the single measurement of the control (0.758). The reported treatment effect and the within-condition spread were the same size, 0.175, so the sign of the effect was not established. The repeat had been included to bound per-rater instability and was never aggregated to the condition level, which is why two experiments and two independent pre-run critic passes missed it. Aggregate every repeat at the level the claim is made at, and repeat every arm the claim compares. | S049 | S049 |
| (bdc) | A gate is prior; the result it gates is posterior — so "the gate would have blocked a run that then worked" shows conservatism, not invalidity. S050's frozen design predicted that a repaired stage-1 gate would be shown "not merely stricter" but "invalid", on exactly that inference. The independent pre-run critic refuted it before a number was computed, and the design was amended to claim an operating characteristic — a false-negative count on the cases available — and to rest the wrong-statistic argument on the statistic's own properties instead. The corrected version is stronger, and the original would have been quoted for years. Any argument that runs from a downstream success back to a prior rule's validity is this mistake. | S050 | S050 |
| (bde) | Removing a between-group inflation from a pooled statistic does not remove the same inflation one level down. RS-20260727b §4 showed a pooled jury gate passing on variance that was 56.8% between jurors, and repaired it to a per-juror statistic. The per-juror statistic has the identical defect between senses: a juror who marks one sense two points above another scores dispersion for doing so. Recomputed with each sense's own mean removed, the only cell in the project's record that passes the gate stops passing (0.866 → 0.617), and its drop is five times any other cell's. When a decomposition fixes a statistic, apply it to every grouping the statistic pools over, not only the one that was noticed. |
S050 | S050 |
| (bda) | An empty stratum is UNEVALUABLE, not a failure — and the disposition has to be written before the data, or the temptation is irresistible. S049 pre-registered its headline on a stratum the materials came back empty; the design was amended before dispatch to record that the pre-committed "prediction fails" disposition does not fire, that the arm could not close on it, and that no re-briefing of the extractor would be done to manufacture the stratum. Re-specifying a population after seeing that the first one was inconvenient is fitting materials to the hypothesis, and it is exactly the move the ratifying vote had already refused this line of evidence for. | S049 | S049 · S081 — three seat dispatches failed and none was re-briefed; the design's own floor rule for a short-handed passage was applied instead of a repair |
| (abe) | Two rules of the same shape at different n are not "the same rule". Check proportion, sidedness, exact null probability and power before calling a comparison symmetric; a design that leans on the symmetry is leaning on four claims, not one. | S033 | S033 |
| (abj) | Sort an evidence base by evidence class, not by topic — a failed gate does not damage it uniformly. Tier D's failure destroys everything resting on jury scores and leaves close readings and machine measurements untouched. Sorted by topic that is invisible; sorted by class it is one column, and the blast radius is countable. | S035 | S035 |
| (abk) | An option-set, a diagnostic question and a taxonomy are not recommendations. They read like guidance and decide nothing. The test is whether a claim would have changed what the translator actually wrote — apply it to each candidate before a release, because a framework accretes this material without noticing. | S035 | S035 |
| (r) | Matchers, tokenizers and sentence splitters must be frozen in the design page verbatim — and so must the tie and direction rules of every statistic. | S021 | S030 |
| (kk) | When an instrument has one free parameter you cannot justify, do not choose it — bracket it. | S026 | S031 |
| (qq) | A summary statistic over a few units can conceal a sign reversal; report per-unit values and register per-unit predictions. | S028 | S029 |
| (tt) | If you mean the ordering, register the ordering, not the sign of a correlation. | S029 | S031 |
| (vv) | Assert the count of every unit you expect to extract, and fail loudly. | S030 | S031 (caught an error in a frozen spec, A17) |
| (yy) | Commit the answer key before the attempt, not just the design. | S030 | S031 |
| (ww) | Design the rank statistic so its subject and its reference are the same size. | S030 | S030 |
| (uu) | A rank statistic saturates when subject and reference differ in length by orders of magnitude. | S029 | S029 |
| (nn) | When a ratio's denominator is unstable, stop dividing and measure the distribution. | S027 | S028 |
| (oo) | Adding a text to a cell changes every other text's centrality; state the cell size. | S027 | S027 |
| (t) | A landmark must declare what kind of fact it is. | S023 | S024 |
| (y) | An accept set encodes the auditor's vocabulary. | S023 | S023 |
| (bdm) | When a category label will not reproduce between readers, test whether it is asking two questions before concluding the phenomenon is fuzzy. RS-20260729 found the drift window's four-class scheme reproducing at 0.654 and read it as evidence that "no overlap at all" is not a line two readers can share. It is partly that and partly something cheaper: the scheme runs semantic distance and register down one axis, so duguð → doughty drew total from one reader ("does not mean retainers") and marked from the other ("still means valiant but is archaic"), and both were right about different questions. Splitting the label into two graded axes at S054 raised three-rater reliability from α = 0.51 (nominal, four-class) to α = 0.78 (semantic) and α = 0.89 (currency). The diagnostic is one question — "could a reader answer this label two ways and be right both times?" — and it costs nothing to ask at design time. | S054 | S054 |
| (bds) | Two failure criteria whose antecedents overlap will both fire, and one of them will rest on a measurement the other has disabled. E-20260729d registered F3 (does this candidate enter the inventory?) and F4 (is the project's standing zero a property of translation decisions?) with antecedents sharing a term, plus F1, a reliability floor over the pass F4's antecedent counts come from. All three fired. The adjudication is available — F3's antecedent needs only an occurrence plus a second, reliable pass; F4's needs a count from the failed pass — but it is an adjudication made after seeing the numbers, which is exactly what pre-registration exists to avoid. Registering criteria is not enough: their antecedents must be checked for overlap, and any criterion resting on a quantity another criterion can void must say in advance what happens when it is voided. | S056 | S056 |
| (bew) | When one agent writes a coding rule, applies it and reports the statistic, the rule's definition and its application come apart silently — and the cheapest check is to re-derive its clearest cases from the definition's own text. R07 §5 defines D as a rule names the feature and only ONE live option satisfies it. Across three runs a D was in fact recorded whenever a rule excluded an option. Three sites were forced into E-20260730e as a positive control because the lead judged them its clearest cases; at two of the three, F4 excludes one option and two English renderings remain, so the correct code is P. Three independent readers returned P there. The control failed because the control was wrong, and the readers were right. The defect is invisible to any amount of care in the individual judgments, because each one felt decided. Rule: for any coding scheme the project both defines and applies, take three cases the coder is most confident of and re-derive them from the definition's wording before reporting a rate over them. | S064 | S064 |
| (bey) | Read a long table through a truncated view and you get a wrong denominator; parse tables, do not read them. E-20260730e's pre-registration quoted T-osso-di-morto-R07-v1 at 21 logged sites; it has 44, and its own tally section says so correctly. The first 21 rows were taken for the whole table. The same parse then found T-postmaster-R07-v1's stated split (P 8 / S 2) wrong at P 7 / S 3. With T-osso-di-morto's own note that its hand tally of 7/28/9 was wrong, three of five R07 tally lines were wrong on first writing and every one was caught by parsing rather than rereading — and every error was in the P/S split, never in a D count. Rule: any count over a table in this repository is computed by a script that reads the file, and the script's row count is asserted against the table's own stated N. | S064 | S064 |
| (bib) | "Tier D does not bind this run, because nothing here judges quality" is a true premise that does not reach an operational rule change — and a design that leans on it will have its motion refused. D-20260803-15 proposed amending naturalness on descriptive markedness coding by an uncalibrated panel, arguing correctly that Tier D governs quality verdicts and this run made none. The ratifying vote accepted the premise and refused the motion: the proposed clause would govern how future evaluations score markedness, which is operational and consequential whether or not it is itself a quality claim. Two further grounds it named, both usable in advance: a difference between two prompted rating procedures is not a demonstrated property of the target language or of readers — say which you measured, in the sentence, every time — and an effect whose magnitude nearly doubles with item order supports no threshold-based decision, however clean the pre-registration. Rule: any motion to change a sense's operational wording on model-rater evidence must pre-register validation capable of distinguishing model prompt-following from the property claimed, and must pre-register what order sensitivity would void the threshold. Registering the threshold is not enough if the estimate is not stable under the design's own robustness arm. | S099 | S099 |
| (bhr) | [FIRED S162, on an amendment the lead adopted from its own critic. E-20260811h accepted A10 — a minimum of 24 decided judgments before any equivalence verdict — and in the same session cut three control contrasts to one order per passage, giving them a maximum of 24 judgments. Two equivalence verdicts were therefore unreachable by construction and returned INCONCLUSIVE at 21 and 17. The arithmetic needed no seat, no call and no result. An amendment is a failure criterion too, and adopting one does not exempt it from this note. RS-20260811h §5.] Every failure criterion computable from the frozen materials must be computed BEFORE dispatch, not at analysis time. E-20260802f registered FC3, a matching check on two arms the lead had already written — and the pre-run critic computed it and found it failed at 38.3% against a 30% ceiling, which would have withheld the run's second primary before a single datum existed. Nothing about that check needed a seat, a call or a result; it needed arithmetic over two frozen strings. A criterion that can only fire after the money is spent is a criterion the design chose not to test. The rule: at freeze time, partition the failure criteria into those that depend on returned data and those that do not, run the second set, and put their values in the design. This is not the same as note (abc) (estimate from the cap) or note (bfc) (declare a reserve); it is about the design's own logic**, not its cost. | S092 | S092 |
| (bio) | A mechanical check on a frozen specification catches literal strings and nothing else, so the pre-run critic must be asked for the paraphrase BY NAME — and when it finds one, the specification does not get rewritten, the CLAIM does. R11 requires that a purpose specification not be written in the vocabulary of wiki/goodness-senses.md, and records that at S084 a paraphrase got through the check and would have manufactured a trade at every site. S107 repeated the failure with the check passing. materials/check_specs.py cleared both frozen specs against 35 strings; the critic, asked for the paraphrase by name in the design's own brief, found each spec restating one half of the definition the run existed to test. By then both arms were rendered against those specs. The repair is not to rewrite a freeze — that is worse than the defect — but to re-register, before any measurement call, what the run is allowed to conclude: S107 downgraded its primary from the halves are separable to the halves are separable by a translator pointed at these two readers, promoted the null to the strong outcome, and raised its own motion bar. Rule: on any run whose specification could contain its own answer, (i) write the paraphrase question into the critic brief, (ii) run the critic BEFORE the material is produced if the specification governs production, and (iii) if it fires late, amend the claim and not the freeze. Companion to (rr) and (big). | S107 | S107 |
| (biq) | A blind is not a property of the materials; it is a property of the materials AT A GRAIN, and a design that reports "the seats recognised the translators" without naming the unit has not measured what it thinks. RS-20260804c put whole loci of «Певцы» to three seats and all three named Garnett and all three named Hapgood, and the run concluded — in the strongest sentence in it — that "a documented human ranking of two published translations of this period cannot be tested blind on this instrument, because the instrument is not blind." E-20260804i put the same story, the same pair, the same three seats one to three sentences at a time, gave them a false attribution and asked them to name the translators independently of it: P1 followed the false label at 24 of 24 cells and the true text at 0. The blind that failed at locus level holds at site level, and an arm step written to go and find new materials did not need to. Rule: every recognition probe reports the unit it was run on, and no blind claim generalises across grains without being re-measured at the grain the design actually uses. The corollary is cheap and worth having: when a blind fails, shrink the unit before abandoning the materials. Companion to (bfq) and (bdq) — reachability is a property of the reader — with the addition that it is a property of the reader and the window. | S108 | S108 |
| (bir) | An amendment that changes a design's shape silently invalidates the failure criteria written for the old shape, and the pre-run critic will not catch it because the criteria were clean when it read them. E-20260804i registered FC2, a degeneracy check excluding any seat that returns the identical ranking at ≥ 11 of 12 sites. It was written for a four-arm stage. The same critic pass that approved it produced BLOCKING finding F4, which abolished the four-arm stage in favour of two-arm blocks — and with two arms, "the same arm first at 11 of 12 sites" is not degeneracy but exactly what a real dimensional difference looks like. Applied literally after the run, FC2 would have deleted the whole of stage 2's accuracy data because both its seats were consistent. The run reported both readings and declared the defect rather than quietly dropping the criterion, which is the only honest move left once the numbers exist. Rule: after accepting an amendment that changes arms, stages, or units, re-read every registered criterion against the NEW shape and re-freeze it before dispatch — an amendment is a redesign, and a criterion inherited across it is inherited untested. Sibling of (bds), which is the same failure between two criteria rather than between a criterion and an amendment. | S108 | S108 |
| (bip) | A no-effect control site must be chosen so that the artefact could show up in it — a control whose ceiling is below the effect's floor cannot fail. E-20260804h registered a localization control on two no-fork sites, and the pre-run critic showed arithmetically that it could not fail: one of the two was twenty words long and nearly unflattenable, capping the no-fork rate near 0.50 while the fork rate could reach 1.00, so the control would have passed with the artefact present. Four sites were added, giving six and six, and the criterion was restated as a margin rather than a comparison. It then returned 0.778 against 0.278 and the added sites were where the run was most informative — they split in both directions, which is what showed the translator's hand was not the explanation. Rule: before registering a control-vs-treatment comparison, compute the control arm's attainable range under the artefact you fear; if the artefact cannot move it past the bar, the control is decoration. Companion to (rr): a critic that does this arithmetic is worth more than one that reads the argument. | S107 | S107 |
| (bin) | Byte-identity is not neutrality: a repair that holds material fixed in order to vary one thing must first read what it is holding fixed. E-20260804g was designed to re-run six sites with only the yardstick changed, and proposed to hand the yardstick seat the same frozen literal glosses the graders get. The pre-run critic went and read them: five of the six state the relation outright — little-dear-earth, your-servant, daughters rather than sons ... feminine, you[plural] — and at one site the literal gloss and the filed rendering are the same English sentence, so the design would have shown its "independent" seat the very translation it promised on its own face that seat had never seen. Rule: when a design's argument is we changed only X, enumerate what the held-fixed material actually contains before dispatch, and say which of it could produce the result on its own. The held-fixed inputs of the new role are not covered by the freeze that protects the old comparison, and inheriting them unread is how a control leaks. Companion to (rr): the critic is the mechanism that caught it, and it caught it by reading the materials rather than the design. | S106 | S106 |
| (bll) | A manipulation that produces a MIXTURE of two arms is tested against a composition-matched arm, never against an interpolation between them. E-20260809i registered its primary as a dip below the linear interpolation of an all-domesticated and an all-foreign arm at the drifting arm's own foreign-token proportion. The pre-run critic's BLOCKING 1 showed the null is unearned: jury scores are bounded ordinal judgements, and any convex response to that proportion produces the registered dip with no drift effect present at all — and can do it on one sense and not another, so the cross-sense interaction the design called immune is not immune either. The remedy was cheap and had to be built rather than argued: an item-uniform arm carrying exactly the same number of foreign tokens in every block (12/16, 13/19, 14/19), drawn from as nearly as possible the same items, so the contrast is a direct one between two texts of equal composition. Rule: where an arm is a mixture, build the matched arm; a chord residual is an inference about a curve nobody has measured. The general form: an additive or linear null asserted over a scale whose shape has never been measured is an assumption presented as a control. RS-20260809i-terminology-drift §6. | S147 | S147 |
| (blw) | A registered count is untestable unless it says how a case with two causes is counted, and a rate computed from a list drawn after the data are open is selection on the outcome. RS-20260810x registered P-B — sites where a rule bound after S132 makes the new English differ — inside a frozen translation log, before the comparator was opened, which is the hard part of pre-registration and was done right. It then adopted the precedent's exclusive-cause precedence after the run, under which the count is 5 and the 6–15 band fails; under P-B's literal wording it is 9 and the band passes. Both readings are legitimate and the prediction cannot be scored, so it is recorded INVALID AS REGISTERED rather than as a pass or a fail — choosing between them after seeing the answer is the thing pre-registration exists to prevent. The same page's first draft reported 23 of 24 named decision sites agree; the independent critic's BLOCKING 5 showed the 24 were named after both texts were open and were memorable because they matched. Replaced by a frame fixed inside the freeze — the 22 decisions the translator's own log names — on which the figure is 12 agree, 10 differ. Two rules. (i) Any registered count states its counting rule, including how an item with more than one cause is assigned, before dispatch. (ii) Any agreement or convergence rate is computed over a frame that existed before the comparison did, and a hand-picked list of matches is published as examples with no denominator. Sibling of (bkm), which is the same failure on a null; companion to (bkc), which is about what a count of divergences can mean at all. RS-20260810x-legend-again §6, §7, §12. | S154 | S154 |
| (bnj) | In a census of what published translators did with a source device, a locus that is not in a translator's own source text is not a locus he dropped — code it separately, never fold it into the loss count, and where the copy-text is not the text the hands worked from, say so on the verdict. E-20260814 censused eight rhymed-prose loci over Lane 1839 and Burton 1885. Locus 8 turned out to be absent from both, and the reason is neither translator's: both of their Arabic texts carry a vizier-and-letter episode that the copy-text compresses into eleven words, so the join the locus names does not exist in either. Lane abridged and Burton expanded, but neither can explain material present in both of them and missing from the copy-text. Had absent been scored as lost, Lane's figure would have read 8 of 8 lost and Burton's 5 of 8 — both wrong, and wrong in the direction that flatters the finding. The design registered the four-value scale and a power floor (≥ 6 of 8 loci locatable in both hands) before either text was opened, which is why the ambiguity is reportable rather than fatal. The rule generalises past unstable recensions to any census over hands whose source texts are not demonstrably identical to the copy-text — which is most of them, and had not come up before only because the project's first eighteen source languages all had single authorial texts. RS-20260814-saj-carriage §6a, limit 3. | S180 | S180 |
| (bnl) | An effort pin is a request, not a control: disable reasoning outright where a seat judges nothing, and where it judges, size the cap from MEASURED reasoning tokens before dispatching the stage. E-20260813f proved both halves in one session. Half one: reasoning: {"effort": "low"} was ignored by qwen on Alibaba (note (bnk)'s second firing), while reasoning: {"enabled": false} on the same seat and prompt returned a complete body with 0 reasoning tokens for $0.001081 — a 40× cost fall and the failure mode gone. z-ai/glm-5.2 behaved identically; google/gemini-3.6-flash rejects enabled: false with HTTP 400 and needs {"effort": "minimal"}, so the disable idiom is per-provider and must be probed, not assumed. Half two: the design's §8 caps of 500–700 were built from assumed answer length, and one probe per stage measured what the judges actually spend before answering — P1 0–403 reasoning tokens, P3 326–430, P2 up to 1,043 on one stage. Every judge stage would have truncated. Fires at: every stage of every run on a seat that reasons, which on this panel is all of them. Remedy, and the line it draws: disable reasoning on seats that write documents or code gates — they judge nothing and the project already wanted minimal effort there — and never on a seat whose output is a judgment, because whether a judge thinks before answering is part of the instrument; for those, spend one probe per stage on the real prompt and set the cap from the number it returns. RS-20260814c-plain-yardstick §9 and E-20260813f §12 A6. | S182 | S182 |
| (bnm) | A prompt that embeds model-generated text cannot be asserted free of a word that the generating prompt requires — assert the ban on the scaffolding, measure the contamination in the block, and never claim a blind the design cannot hold. E-20260813f §10.3 asserted that no GS prompt contains original, Japanese, source, effect, comparable or translation — the whole point of GS being a target-only style question. But a GS prompt embeds a YP document, and YP's own instruction two stages earlier requires the word: "what the passage does to a reader of Japanese". The assertion was unsatisfiable the day it was frozen. Measured: 24 of 90 GS prompts carry a banned word inside the document block, across 5 segments, so in a quarter of its cells the seat could see that the two passages are translations from Japanese. Fires at: every design with a blindness or leakage assertion over a prompt containing generated material — every yardstick design this project has run. Remedy: split the assertion in two. The scaffolding is fixed text and the ban holds over it absolutely; the generated block is data, and the verifier counts the leak and the result page reports it as a limit. And read the generating prompt against the ban at freeze time — here the two prompts sat four pages apart in one document and contradicted each other. RS-20260814c-plain-yardstick §9. | S182 | S182 |
| (bnn) | A dispatcher's usability test must require the reply to PARSE and carry its stage's key; "non-empty content" passes bodies that answered a different question. E-20260813f's runner implemented F6's truncation guard as empty, or unbalanced braces. yp__S14 returned 1,200 tokens of bare prose with no JSON anywhere and finish_reason: "stop" — a parse failure F6 explicitly covers — and the test called it usable. It was caught only because a downstream count came back 14 where 15 was expected. Fires at: every runner in this project, all of which key deadness off content being non-empty. Remedy: one table from stage to expected key, checked inside the dispatch loop so the re-dispatch ladder actually runs; and mark the stored record usable: false so the analysis and the F5 count key off the same fact rather than re-deriving it. RS-20260814c-plain-yardstick §9. | S182 | S182 |
| (bnk) | On a reasoning seat, max_tokens caps the answer PLUS the hidden thinking, and the thinking will take all of it — cap the reasoning, not the total, or budget for buying nothing. E-20260814b's pre-run critic (P4 moonshotai/kimi-k3) was dispatched at max_tokens 4,000 and then 14,000; both returned finish_reason: length with empty content after spending 3,997 and ~14,000 tokens on reasoning, and a third attempt was billed and lost to a network error. $0.526018 — 47.4% of the run's total spend — bought no output token at all. The fourth attempt, identical except for reasoning: {"effort": "low"} and max_tokens 6,000, returned a full NEEDS-AMENDMENT verdict with eight findings in 91 seconds for $0.083293. Fires at: every dispatch to a seat that reasons before answering — on this panel P4 above all — and every pre-flight estimate that prices such a call. Remedy: send reasoning: {"effort": "low"} (or an explicit reasoning cap) on any call whose value is the written answer rather than the depth of the thinking, and treat note (abc)'s worst case as a floor rather than a ceiling on a reasoning seat: (abc) prices the cap, and this note says the cap can be consumed without producing anything. RS-20260814b-honorific-hands §8. | S181 | S221 |
| (bod) | "Sentence 24" is not a reference unless the design says whose sentences, and any property a selection rule claims must be ASSERTED BY THE SCRIPT rather than described in prose. E-20260816 numbered the Arabic source's sentences in its frozen figure inventory and the English rendering's sentences in its decoy-selection rule, called both sentence N, and the two sequences collide at 24, 35 and 45. The pre-run critic read the collision as a BLOCKING defect — three decoy loci allegedly containing enumerated source figures — and it was wrong on the facts (those decoys render Arabic sentences 19–20, 31 and 41, against figures at 24, 35 and 45) and right that nothing in the design could settle it. A reader who cannot check a claim has to disbelieve it, and a critic who disbelieves the wrong claim spends its findings there instead of on the design. Fires at: every design that indexes two texts — which is every design in this project with a source and a rendering, and the collision has simply not been noticed before. Remedy, two parts. (i) Name the text in every index (Ar-24, En-24), never the bare number. (ii) Where the prose asserts a structural property of the materials — no decoy span overlaps a figure span, the arms differ only here — put it in the cell builder as a check that fails the run, so it is verified rather than believed; E-20260816's cells.py now asserts pairwise locus disjointness and the class multiset. Corollary worth keeping: a critic finding that is factually wrong BECAUSE the design was ambiguous is a finding about the design, and is recorded as accepted-as-defect rather than refused — refusing it teaches the next session nothing. RS-20260816-answering-figure §4; E-20260816 §11 finding 1. | S191 | S191 |
Controls, gates and critique
| id | note | first seen | last fired |
|---|---|---|---|
| (bss) | When the question is whether a hand tracked an ARRANGEMENT rather than a COUNT, the null is a permutation of the arrangement with the inventory held fixed — not a comparison against some other real arrangement. E-20260901b v1 tested whether Leaf's English fitted his own Persian measure by scoring it against other Persian measures of the same syllable count. Both critic seats rejected it independently and for the same reason: rival real schemes differ in how many short positions they have and where, target-language stress is position-structured, and beating an alien row shows nothing. The replacement holds the subject's own data fixed (the line's stress vector) and the scheme's inventory fixed (its length and its number of shorts), and permutes only the order. Any bias that is a property of the item is then held constant by construction, and no fitted model is needed. It also has a closed form worth knowing: choosing k short positions out of N is hypergeometric, so the per-item null mean is just the item's own overall rate, and only the clustered significance test needs simulating. The measured separation was 0.0472 against 0.4982 with the shuffled-stress control at +0.007. Fires at: any claim that a form, order, template or arrangement was followed — metre, rhyme scheme, colon length, clause order, information structure. Remedy: permute the target structure within its own inventory, report the closed-form null where one exists, cluster the significance test on the unit that repeats (here the metre, not the poem), and add a target-language-typical template as a second control so a generic effect cannot masquerade as fidelity. RS-20260901b-leaf-pattern §4. |
S238 | S238 |
| (bsp) | A calibration floor built from hand-made contrasts can exclude your BEST coder — gate on real items, and treat a synthetic floor as a floor and never as a filter. E-20260901 registered two gates on three blind seats: 24 hand-made word-order pairs (bar 21) and 40 real study items the lead had coded blind (bar 0.75). P2 scored the run's highest real-item agreement, 36 of 40, and was excluded on the synthetic set at 20 of 24 — so one seat of three voted, the design required two, and both word-order primaries were withheld after 88 coding calls and $2.96 had already been spent on them. Both pre-run critics had said the synthetic pairs would prove little about the real task; what neither predicted, and what this run measured, is that failing them proves little either. One of the 24 items was itself wrong — all three seats rejected the key unanimously, and on inspection they were right — though excluding it does not change the outcome at the same proportional bar. Fires at: any design that gates a bought coding stage on items the lead constructed rather than on items drawn from the corpus. Remedy: make the real-item key the gate and the synthetic set a diagnostic reported beside it; build the synthetic set only to catch a seat that cannot do the task at all, and set its bar where a competent coder cannot fail it — or drop it. And never repair the floor after seeing which seats it excluded: the repair belongs to the next design, with the outcome of this one already fixed. Kin to (bsj) — a control built inside the genre can invert — of which this is the calibration-side twin. RS-20260901-inversion-habit §4. |
S237 | S237 |
| (brs) | A forced-choice preference task between two near-identical passages must have its ORDER-CONSISTENCY measured before its numbers are believed: on minimal pairs these seats agree with themselves across a swap of presentation order barely above chance. E-20260826b ran 168 forced choices over seven windows, four line-end arms and three seats, both orders, 0 dead. Of the 84 (seat × locus × edge) cells, only 54.8% returned the same arm when the two texts were swapped, against a chance rate of 50% — QR 67.9%, P2 50.0%, P1 46.4%, below chance — and every seat preferred whichever passage was shown first (0.625 · 0.714 · 0.768). The registered order criterion fired on three of four edges and withheld two of the three registered quantities. Fires at: any A-vs-B preference or which-reads-better task on passages differing in a few words — the shape that produced RS-20260824c's extent figure and RS-20260822b's. Remedy: run both orders and report the pooled difference, which a symmetric position effect attenuates rather than biases (this is why RS-20260824c is not impeached); and compute the per-cell order-consistency rate as a gate, not merely an order gap on the pooled rate — a 0.257 pooled gap and a 54.8% consistency rate are the same data read at two grains, and only the second shows how little of the answer is about the text. A pooled order gap alone cannot distinguish an instrument that is attenuated from one that is empty. THIS IS THE SECOND OCCURRENCE AND THE FIRST ONE WROTE NO NOTE: RS-20260825c-worth-paying §2–3 (S222, the day before) found the same thing in stronger form — "told nothing, the seats are not choosing between the translations at all, they are choosing the first passage", first-position rate 0.819 — and repaired it with a post-hoc content-decided statistic instead of a note, so the lesson did not reach E-20260826b, which paid for it again. And S222 carries the remedy this note should propagate: in its I2 condition, where the seats were told what the original does, the same seats preferred the same arm in both orders. So the bias is a property of an underdetermined task, not of the seats: give the judgment something to be against — the source's behaviour, a stated criterion — before concluding that a preference instrument is unusable. RS-20260826b-radif §5, diagnostics.json; RS-20260825c-worth-paying §2–3. |
S222 | S229 |
| (brq) | A length, evenness or spacing statistic computed on each hand's OWN printed marks is a statistic about typography until it is shown to survive a segmentation applied identically to every hand — and it may not survive. E-20260826's pre-run critic named this as BLOCKING before any measurement: Preston 1850's units are printed lines, Chenery 1867's are em-dash spans, and nothing showed the two encode the same thing. Three segmentations were run instead of one, and the answer split. Preston's clause-length ceiling and his low dispersion hold under his own lineation and under a comma-inclusive rule applied to all hands, and reverse under a strong-punctuation rule — his 95th percentile goes from below Chenery's to 54 and 61 syllables against 42 and 46. A design measuring only OWN would have reported that he obeys his own printed rule throughout, and that is false. Fires at: any census of clause length, period length, line length, paragraph size, sentence count or spacing over published hands — the anchor shelf is full of candidates. Remedy: compute at least one segmentation that is nobody's own marking and report the agreement; where a result holds only under a hand's own marks, say in the sentence that reports it that it is a property of the printed page. RS-20260826-balanced-period §5, critic-response.md B2. |
S223 | S223 |
| (brn) | Determinism is a property of the TASK, not of the seats — a near-deterministic result on one task shape does not licence skipping the repeat check on the next. S220 measured these same three seats at temperature 1 returning one distinct answer vector in six draws in 8 of 12 cells on a reference-resolution task, and every P value on that page was bounded accordingly. One session later, on a graded aesthetic judgement over the same seats and the same temperature, E-20260825b stage R returned 0 of 15 repeat cells identical to their first draw. A task with a retrievable right answer collapses the seat's output distribution; a task asking for a preference does not. Fires at: any design tempted to inherit a determinism finding, in either direction, from a run of a different shape. Remedy: budget a stratified repeat stage in every run that will report a P value, and read S220's finding as being about reference resolution and nothing else. |
S221 | S229 |
| (brg) | When the registered prediction is an ABSENCE, a control of the form the treatment rate must exceed the background rate disqualifies exactly the hands that confirm it — a control has to be able to be passed by a true hypothesis. E-20260824c predicted that published translators produce no rhyme at a source's rhyme loci, and its F1 said in the same breath that a hand whose locus rate does not exceed its own chance rate is not credited with carrying anything. A hand that never rhymes cannot exceed chance; proving the headline would have thrown the data out. The P2 seat found it in a pre-run pass and this design had not. The remedy applied was to demote the comparison to a descriptive background rate that gates nothing, and to say on the result page that the study therefore has no control — which is true and is worth more than a control-shaped number. Fires at: every design whose primary is that something is not done, not carried, not registered — a large and growing share of this project's, since a null is a first-class result here. Remedy, at design time: write down what the control returns if the hypothesis is true, before freezing. If the answer is it fails, it is not a control. And where no control can exist — as here, al-Ḥarīrī's unrhymed prose in this maqāma being six words long — say so on the page instead of substituting a chance rate and calling it one. RS-20260824b-hariri-hands §11, critic-response.md B5. |
S218 | S218 |
| (brh) | A FORM code read off the page — indented, line-broken, capitalised at the line-head — cannot see a translator who lineates prose and does not rhyme it, and the same hand then scores 0.000 or 1.000 on a definition. Preston 1850 sets al-Ḥarīrī's rhymed prose out one source colon to a line, evenly balanced and unrhymed, and says in his introduction that it is "a species of composition which occupies a middle place between prose and verse". On the typographic rule his GENRE-CHANGED is 65 of 65; on his own printed declaration it is 0 of 65. E-20260824c coded it COLOMETRIC and reported both. This is not a curiosity: framework/v0.2 §7.27 item 2 publishes a rate of the form hands move rhymed prose into verse at 4 of 128, and that figure is safe only because none of the four Gulistan hands lineated. Fires at: any census that counts a genre change, a verse setting, a stanza, or a paragraph, off the printed page. Remedy, in the design and before any book is opened: state the third value — lines that are not verse — and state which reading the headline uses; where a translator's own preface names his form, quote it rather than inferring from the layout. RS-20260824b-hariri-hands §6. |
S218 | S218 |
| (brk) | A mechanical rhyme grader that accepts ANY dictionary pronunciation of a word can certify a chime a reader will not get — where a bearer is a heteronym, the design must fix which pronunciation the carrier forces, and the tool cannot do it. E-20260824c v1 chimed the written decree is read with said, and tools/rhyme_pairs.py graded the pair STRICT: CMUdict carries both /rɛd/ and /riːd/ for read, and the rule relates two bearers if any pair of their pronunciations relates. That rule is correct for the question it was built for — is a rhyme reachable from this list of synonyms — and wrong for the question a carrier asks — is a rhyme there, since a reader reading a passive present verb hears /riːd/ and gets nothing. The round-1 critic caught it; the tool, the hand and the exhaustive 70-variant pre-dispatch grade all passed it, and a mutation test in verify.py confirms the pair still grades STRICT today, so this is a property of the instrument and not a slip. It would have attenuated the primary at one locus in seven and inflated the extent statement in the direction the lead had predicted. Fires at: every design whose carriers are graded by tools/rhyme_pairs.py rather than merely screened by it. Remedy:** before dispatch, list every rhyme-bearer that CMUdict gives more than one rime for, and either rebuild the carrier on an unambiguous word or state in the design which pronunciation the syntax forces and why. RS-20260824c-run-placement §6. |
S219 | S219 |
| (bqy) | When the instrument reads TEXT and the question is about SOUND, the design must carry a dissociation probe with a registered veto — spelling and pronunciation are near-collinear in English and the confound is invisible without one. E-20260822b measured, before dispatch, that the mean shared final-letter run between its two words was 2.63 at STRICT, 1.60 at NEAR, 0.13 at NONE: the independent variable and its orthographic shadow were almost the same variable, and nothing in the design would have shown it. The probe that fixed this is small and decisive — 6 full rhymes spelled unlike (blade / weighed, near / austere) against 6 non-rhymes spelled alike (sword / word, beard / heard) — and it returned 0.778 against 0.722, with one seat of three at 0.833 orthographic / 0.333 phonetic. The veto (if the spelled-alike arm outscores the real-rhyme arm, the phonetic reading of every figure is withdrawn) must be registered before dispatch, or the probe becomes a thing to argue with afterwards. Fires at: any design about rhyme, assonance, alliteration, metre, or any other property of a text's sound, run on a seat that receives characters. Building the probe is also how you learn the confound's size: of 30 classic English orthographic traps, tools/rhyme_pairs.py scores 25 as NEAR, so the permissive class and the spelled-alike class very largely coincide. RS-20260822b-echo-threshold §4, critic round 2 BLOCKING 1. |
S212 | S212 |
| (bqz) | A fluency or meaning screen inherits the register of its prompt, not of its material: ask for "ordinary English" about elevated prose and it will reject the prose. E-20260822b's screen judged 76 of 138 perfectly grammatical items not fluent, with reasons that name register and not grammar — "Archaic and literary phrasing, not modern or ordinary English" — because the material is the Gulistan's English and the prompt asked for ordinary. Applying the registered exclusion rule would have dropped 22 of 25 loci and left the primary uncomputable. What saved the run was the planted breaks, and they are what makes a screen an instrument rather than an opinion: the same screen missed 1 of 4 ("the footing must be laid before the therefore" passed as fluent), which is what licensed disclosing it as failed instead of applying it. Remedy: word the screen for the register the material actually has ("a well-formed sentence of English", not "ordinary"), always plant breaks, and register in advance what happens when the screen fails its own control — because deciding that afterwards is deciding it with the result in view. RS-20260822b-echo-threshold §5. |
S212 | S212 |
| (bre) | Temperature 0 is not determinism, and this project has been buying one call per cell on the assumption that it is. E-20260824 dispatched stage D twice by accident (note (brf)), so 368 of its 390 cells were called between two and five times at temperature 0 — and 28 of those 368 changed their answer between calls, 0.076. By seat: QR 18, P1 9, P2 1, so the instability is concentrated in the seat that was already the odd one out and one seat is very nearly deterministic. Fires at: any design whose cell count is chosen for power, any registered criterion that turns on a small number of cells, and any figure computed from a handful of (cell, seat) draws — a per-seat Δ over 14 loci is 28 draws, and 7% of them are coin flips. Remedy, and it is not 'buy three calls everywhere': say in the design which figures could move if a cell flipped, and where a criterion sits within one or two cells of its bar, buy the replication for those cells only. RS-20260824-eye-or-ear §11a; the primary there is unmoved by the choice of rule (+0.071 first-attempt against +0.077 averaged) and the per-seat figures are not (P1 +0.143 against +0.101). |
S217 | S217 |
| (brc) | A comparator census over published hands that do not share a base text describes BOOKS, not translators — decide which of the two the design is about before it is frozen, and say so in its own title. E-20260823b coded ten narrative-level loci across four English Nights and found that at eight of the twenty cells where a hand differs from the copy-text, the difference is a recension reading and not a decision: the tale King Yunan tells his vizier is the falcon in the copy-text, the parrot in Būlāq and in Galland, and both in Calcutta II, so two loci are permanently N/A for two hands; and the request formula three English books print at the unmarked seam is in all three of their Arabics and not in the copy-text. Even a supplied display heading is ambiguous here, because the second Arabic witness supplies tale headings the copy-text does not (D29). Fires at: every design that reads two or more published hands against one source — this project's commonest study-limb shape, run at spans A, B, C, F and G of ARM-alf-layla and in ARM-persian-hands. Remedy, and it costs nothing at design time: state in the design whether the claim is what a reader of each book gets (always available) or what each translator did (available only where the hands share a base text, or where the base text is itself reachable and checked). E-20260823b declared a conditional check of Galland's French and it decided two loci — the crocodile's absence is Galland's, the deleted nights are Forster's — which is the whole of what could be attributed out of forty cells. |
S215 | S215 |
| (brd) | A stoplist calibrated for short expressions lets pronouns into the pool the moment you point the same tool at whole clauses — and it will report you/do and I/my as rhymes a translator could use. tools/rhyme_pairs.py was built 2026-08-22 to adjudicate lists of synonym expressions, two or three words long, and its STOPWORDS is 31 items. E-20260823c pointed it at whole-clause literal renderings, where the candidate pool is every content word on the page, and its first pass reported the widening gain at AFFIX loci as +0.231 — of whose three loci two were pronoun pairs, you/do and I/my, both scoring STRICT on the rule as written. With a declared closed English function-word list applied symmetrically to every class the figure falls to +0.077. The registered primary was untouched (it was zero on both), which is the only reason this is a note and not an erratum. Fires at: any reuse of a lexical-adjudication tool on a longer unit than the one it was calibrated on — the stoplist, the tokeniser and the bearer rule are all calibrated to a unit length, and none of them announces it. Remedy, and it is cheap: before reusing such a tool at a new unit length, print the pairs it finds, not only the rate. Two minutes of looking at you/do is what caught this; no aggregate would have. |
S216 | S216 |
| (brl) | Indirection is not what buys headroom — a RETRIEVABLE right answer is what removes it. Note (brb) prescribed indirect measurement after a direct identification task saturated; E-20260825 built the indirect design (sixteen marked referring expressions, a closed roster, the words level, frame, nesting and speaker never used) on a question with a definite fact of the matter — who is speaking to whom across an unmarked four-level nesting — and it saturated just as completely: 1,120 item-judgments, 3 errors, fifteen of sixteen loci at 1.000 in all four arms, with the sharpest cell in the work at 18 of 18 in the unmarked arm. The measure that did have headroom (RS-20260824c, and RS-20260822b before it) asked which of two variants read better — a preference, with nothing to retrieve. Fires at: any design that probes whether a property is registered by asking a seat a comprehension question. Remedy:** before buying, ask whether the question has a right answer a competent reader could look up in the text in front of them; if it does, expect 1.000 and either take the passage away before asking, or ask for a preference instead. |
S220 | S220 |
| (brb) | A direct identification task on model seats is a ceiling, not a measurement: if you ask a seat to NAME the property, it will do mechanically what your own tool does, and the design has no headroom left to measure anything with. E-20260823 asked three seats to label the rhyme scheme of numbered English lines — the natural, obvious way to ask whether a rhyme is registered — and got 144 of 144 adjacent full rhymes and 78 of 78 full rhymes at distance 2, at a false-alarm rate of 0.0055 and 0.000. The registered primary, a matched minimal pair differing only in the distance between the rhyme partners, came back 1.000 against 1.000 and was withheld by its own saturation criterion, after three critic rounds had been spent making that pair clean. The contrast that works is RS-20260822b's: it never asked a seat to identify a chime, it asked which of two variants read better and inferred registration from the difference — an INDIRECT measure, which leaves the seat somewhere to fall short. Fires at: any design measuring whether a formal property is registered, noticed, or heard. Remedy: register a saturation criterion (this design had one, which is why the session lost a primary and not its honesty), and before buying, ask what the seat would have to fail at for the number to be informative — if the answer is the thing my own tool does deterministically, the design is a ceiling and should be rebuilt as an indirect measure. |
S214 | S214 |
| (bra) | Grade a claimed carriage between the words that render the source's rhyme-bearers, never between the ends of the English clauses — the two measures disagree, and the disagreement is not noise. At S11 of E-20260822c («ای مردان بکوشید یا جامهٔ زنان بپوشید», where Sa'di rhymes two verbs) three published hands and the lead all put an audible English chime at the clause-ends — men / women — and not one of them was answering the figure: those words render the nouns, and the renderings of the two verbs are exert / wear, Exert / wear, bear yourselves / put on, exert / put on. A clause-end measure scores that locus four-for-four; the bearer measure scores it zero. The chime is a fact about the English words for men and women. The rule that follows: the alignment from each source rhyme-bearer to the English word rendering it is written down and committed before any pair is graded, and only the aligned pair is graded — no cell-wide maximum, and no pair credited to two loci. Multi-bearer loci need a frozen aggregation rule too (adjacent pairs in source order). Fires at: any count of whether a hand carried a sound figure. The lead's own rendering committed the same error under R43 and it is recorded rather than corrected, which is how it was found. RS-20260822c-persian-hands §4; critic rounds 1–3, three BLOCKING findings on this one point. |
S213 | S213 |
| (bqw) | Where the quantity you are measuring is a COUNT of how often a relation holds between candidate expressions, decide the relation mechanically under a written rule — do not ask a seat. RS-20260821c §7 says in its own limits that its echo instrument counts a near-echo and that a strict recount moves every figure; E-20260822 had to count chiming pairs across 2,173 candidate pairs, where a judge that disagrees with itself about what chiming is would have made the answer a fact about the judge. tools/rhyme_pairs.py was built as a declared, timeboxed gate — a pronunciation dictionary, a rule fixed before any list was drawn, STRICT and NEAR reported separately and never summed, and fixtures written before it scored anything, including the paradigm case both ways round (ends / friends must be STRICT and consequences / ends must be NONE). It then decided 2,173 pairs at $0, deterministically and replayably, and printing the pairs behind the permissive class is what showed the class to be empty: spread / tend, scattered / suspended, creatures / entities produce a 0.708 reach that no reader hears. Fires at: any design whose primary is a rate of how often does relation R hold, for any R a rule can state — rhyme, length, shared morphology, string overlap, matched shape. Remedy: build the rule, fixture it against cases whose answers are not in dispute, and — this is the part that earned its place — print the instances behind every reported rate, because a permissive rule looks like a finding until you read what it admitted. This is the fourth firing of the near-echo over-count, after (bcd), (bgf) and (bgi), and the first where the offending pairs are on the page rather than inferred from raters disagreeing. RS-20260822-synonym-reach §4, tools/tests/test_rhyme_pairs.py. |
S211 | S211 |
| (bqv) | Run the source-side gate over the CONTROL stratum as well as the treatment stratum, and when the gate reclassifies a control, that is data about your own classification and belongs on the page. The strata of a design are almost always cut by the party holding the hypothesis, from the source, before any measurement exists — which is exactly the position RS-20260821-matched-shape found fatal on the treatment side. E-20260822 cut 36 Persian loci into rhymed and non-rhymed and shuffled all 36 into one indistinguishable gate stream. The gate returned 24 of 24 rhymed and 11 of 12 controls, and the dissent was not noise: C11, «نخلبندی دانم … شاهدی فروشم», was written into the control pool by the lead and the seats are right that both verbs carry the same first-person ‑am. The lead's classification was wrong and three independent readers of the source caught it before the primary was computed. Because the design had registered the gate over both strata, the fix was a declared one-line sensitivity — remove C11 and the control rate falls from 0.042 to 0.000, the conclusion unmoved — rather than an argument after the fact. Fires at: every design with a treatment stratum and a control stratum cut by the same hand, which is most of them. Remedy: put the control items through the same source-side gate, in the same shuffled stream, and register in advance both the bar for the controls and what a reclassified control does to the analysis. RS-20260822-synonym-reach §3. |
S211 | S211 |
| (bqt) | A predictor that asks two blind raters for “the ordinary word” and then asks whether their answers match will fire on IDENTITY, and identity is not the relation you are measuring — exclude it in the design, before the data exist. E-20260821c measured, for each rhymed Arabic locus, whether the ordinary English words for its members chime. Two seats glossed 39 members one at a time with the siblings masked, which is the part that worked. What did not: at five of fifteen loci a seat gave the same English word twice — الزمان and الأوان both time, الأطباء and الحكماء both doctors, اليونانية and الرومية both Greek, جليسًا and أنيسًا both companion — and every such list was coded RHYME by every judging seat, correctly, because two identical words do echo. The registered PR1 failed on those five loci and on almost nothing else. A repetition is a fact about the lexicon (English does not distinguish the members) and it is not a chime a translator could build on. Fires at: any design whose predictor is do two independently produced renderings resemble each other, on any dimension — sound, length, register, shape. Remedy: in the design, before dispatch, state that an item whose members are identical after normalisation is excluded, or classed on its own row; and put a planted identity pair in the controls so the instrument's behaviour on it is on the record rather than inferred afterwards. Kin to (bqb) — state the exclusion once, in the design, and have the analysis import it. RS-20260821c-echo-availability §6. |
S210 | S210 |
| (bqr) | Where a figure's identity is POSITIONAL — parallelism, isocolon, anaphora, matched shape — no rewording removes it, so a subtractive design must buy an INDEPENDENT judgement that the replacement no longer carries the figure, before the run, and must register a floor on the surviving set. E-20260821b had what looks like the ideal subtractive material: at every locus where the hand built a matched shape, R39/R40 rule 5 had made him write down the plain wording he refused, at the same span, at the same sitting as the rendering and long before any experiment existed — 22 usable minimal pairs across Arabic→English and Chinese→English. A blind seat shown only the replacement's members and their sentence said they still echo each other in form at 16 of 22, naming the property each time (the places of + noun, parallel 'what he + verb' structure, adj + preposition + complement). A second seat, which caught 4 of 4 planted content errors, said 10 of 22 replacements do not say the same thing. Two loci survived both screens against a floor of ten declared in the design, and the primary was WITHHELD before a single body of the main run was dispatched. The mechanism is not carelessness in writing the replacements: the members stay in matched position because the content requires them to, and English then supplies a frame whether or not one is wanted — so the hand cannot write an unmatched alternative even when that is exactly what he is trying to do. Fires at: every subtractive or write-it-plainly design over a figure that lives in the relation between members held in matched position. Remedy: buy the screen before the run, register the floor in the design, and expect to withhold. Kin to (bqm), which prescribes a second writer: where the property is positional a second writer reproduces it too, so what is needed is a second judge. RS-20260821b-matched-heard §3. |
S209 | S209 |
| (bqq) | An agreement bar is meaningless until the constant-answer baseline is computed and printed beside it — a threshold below what a rater who says the commonest label to everything would score is not a bar. E-20260821-matched-shape v1 registered lead–seat agreement ≥ 60% on 40 items as a prediction. The pre-run critic's MAJOR 8: the lead's own coding is PLAIN at 34 of 40, so a procedure that answers PLAIN to every item scores 85% — the bar sat twenty-five points below the trivial floor and could have been cleared while being wrong on every one of the six cells the design was about. v2 replaced it with a prediction stated on the non-PLAIN cells only, a full 3×3 confusion matrix, and the trivial baseline printed on the result page. Fires at: every design that reports agreement, concordance, reproduction or replication as a proportion of items, which is most designs that check one coding against another; and it bites hardest where the label distribution is skewed, which is exactly when an agreement figure looks most impressive. Remedy, before dispatch: compute the majority-label rate of the reference coding, print it, and state the bar over the cells where the two codings can actually differ — or use a chance-corrected statistic and say which. Kin to (bna), which is about the denominator a bar is written over; this is about the floor the bar has to clear. RS-20260821-matched-shape §2, E-20260821-matched-shape critic-response finding 8. |
S208 | S208 |
| (bqp) | One round of critique per design version: findings are addressed on the page, the next round is bought only if the last one killed a numbered primary, and a round is never bought to widen a primary. Fixed at S206 before that session's third critic pass, cited by name in E-20260820c §5 and in E-20260821-matched-shape §5, and filed here at S208 — two sessions after it began binding designs, because a rule of conduct that lives only in the designs that cite it is not in the place a future session reads. The defect it prevents is critic regress: each pass on a re-written design finds new findings, the design keeps improving in ways nobody registered, and the spend and the session both go to the apparatus. Fires at: every design that takes a pre-run adversarial pass. Remedy: state the ceiling in the design before the first pass, and when a pass returns NEEDS REDESIGN, answer every finding in writing, revise, and dispatch — a second pass is bought only where the first killed a numbered primary and the revision might not have fixed it. S208 fires: the pass returned NEEDS REDESIGN with 3 BLOCKING findings, all eleven were accepted and answered in critic-response.md, no second round was bought, and the reason is the rule's own second clause — every remedy narrowed the primary. |
S206 | S212 |
| (bqn) | A within-item paired design does NOT protect against between-run drift in the shared "before" cell — repeated seats and repeated items can move the reference cell's flag for reasons that have nothing to do with the manipulation, so a paired design that compares against an earlier run must include a per-item continuity check on the reference cell, not only aggregate continuity. E-20260820c reused RS-20260817e's nine Botchan sites, its three seats (P2 P3 QR), and its two English strings byte-identical, and manipulated only the source. The aggregate mimetic-present cell reproduced S204's shape (5 of 7 at "adds nothing"), but per item it disagreed at 3 of 7 for reasons unrelated to the mimetic: two seats' flags moved to ADDS because they read differently onto material in the English frame ("in her hands", "bath was ready") that was there under S204 too, and at one site they read no to the adds flag because they had shifted attention to an omitted clause. Aggregate continuity gates ratify a stable panel and hide item-level drift. Fires at: every paired design that reuses seats and items across sessions to compare a new arm with an earlier run's cell. Remedy: at design freeze, transcribe the earlier run's per-item majority into the analyser, register a per-item continuity gate at ≥ (n-1)/n, and make its failure void the between-run inference (not only the aggregate replication claim). RS-20260820c-mimetic-subtraction §4, and the P3' gate the pre-run critic (MAJOR 8) forced into the design. Kin to (bly) — which prohibits crossing sessions for a new comparison — this is the same defect surfacing when the reused arm is the baseline of the new comparison. |
S207 | S207 |
| (bqm) | A subtractive control is not specified by the RULE for writing the replacement — it is specified by the replacement itself, so write it down, and where the claim turns on it, buy a second hand and compare the two before dispatch. E-20260820b gave an independent hand the translator's own rule verbatim — replace the marked stretch with the plainest English rendering of the same events; change nothing outside the stretch; add and remove no event — on the same twelve stretches. It did not remove the marked property. It re-expressed it: in great mouthfuls → greedily, craning his neck and staring → watching closely, at an unhurried walk → slowly, and a rhetorical question → another rhetorical question. Δmanner came out 0.000 at all three manner loci, and every probe 0.000 at the rhetorical question, on eighteen bodies with two versions — while the same probe moved +1.000 elsewhere in the same bodies, so it was not a dead instrument. Two corollaries, both bought at a cost. (i) A mechanical screen for the replacement is identical to the original catches identical strings and cannot catch a different string doing the same work; a Δ of zero then means the arms did not differ, not that a property survived removal. (ii) Going plainer is not going unmarked: at two loci the seats read the replacement as more attitudinal than the figure — "greedily tells how AND JUDGES the eating" against "describes only the way he ate" — because an adverb evaluates where a depiction shows. Fires at: every design with a write it plainly / delete the device / say the same thing without it arm, which includes §7.16's subtractive test, RS-20260816f's deletion control and E-20260817e's. Kin to (bpu), which is the same defect from the parity side. Remedy: freeze the replacement text, not the rule; have a second hand write it independently; and treat disagreement between the two as a finding about the design's reading of the passage rather than a nuisance to average away. RS-20260820b-device-function-2 §2, §7 limit 5. S207 fires: applied verbatim to the substitutive rebuild of §7.19; two hands wrote plain-Japanese substitutes, agreed at 8 of 9, and the one disagreement (M08) was exactly a case where the two substitutes preserved different properties — the remedy the note prescribes. |
S206 | S207 |
| (bpz) | A closed answer vocabulary built from the design's own hypothesis forces every item into the hypothesis — give the jury at least one option the design does not predict, and expect it to carry a result. E-20260816h v1 offered five labels, all five of them classes the design was testing, and the pre-run critic's MAJOR 7 and 8 named the consequence: an item with no sound relation at all had nowhere to go except a sound label, so the negative stratum could not speak. v2 added PARALLEL — parallel in grammar or in sense, but not in sound — and OTHER. Six loci took PARALLEL, and not one was called owed by any seat: the five hard negatives at 0.0000 and one figure the inventory had classed as matched shape, also 0.0000. The option the critic added is the option that produced the page's cleanest separation, and without it those six verdicts would have been distributed across four sound labels and read as noise. Fires at: every judgement task with a closed vocabulary, and hardest where the vocabulary was derived from the very classification under test — which is exactly when it feels most natural to write. Remedy: before freezing a label list, ask what an item outside the hypothesis would have to answer, and add that option; where the list also drives a level clause or a negative stratum, the out-of-hypothesis option is what makes the stratum measurable rather than merely present. Kin to (bnu) — a control domain defined as the absence of the treatment is not a control domain — this is the same defect in the answer space rather than in the item space. RS-20260816h-target-set §1, §9. |
S202 | S202 |
| (bnt) | A blind rater judging a stylistic property needs a control passage written in the material's own register with that property ABSENT — a flat-modern negative control does not test the failure mode the material actually has. E-20260814g asked three seats whether an English passage uses conspicuous sound-patterning, over Lane 1839 and Burton 1885 — both archaic, ornate, Orientalist prose. Its two controls were a heavily alliterative passage and a flat modern one, and the pre-run critic's finding 7 was that neither tests the obvious false positive: ornate-sounding English that carries no sound device. A third control ARCH was written — "Thereupon the merchant, being minded to depart, made ready his beasts and his servants…" — dispatched blind among the census, and came back N N N. Had it come back Y, every rate on the page would have been uninterpretable and the run wasted. Fires at: any design putting a blind rater on a stylistic property of prose whose register is itself marked — archaic, ornate, dialectal, technical. Remedy: build the negative control in the material's register, not in the rater's default one, and make its failure fail the instrument. Kin to (bkt), whose lesson was that a positive control matched on the wrong dimension makes a null unreadable; this is the same defect on the negative side. RS-20260814g-supplied-sound §4. |
S185 | S185 |
| (bnu) | A control domain defined as the absence of the treatment is not a control domain — audit it affirmatively for the treatment, with the same criteria, before the run. E-20260814g's PLAIN domain was originally "span B prose segments that do not contain one of the pre-frozen PATTERNED loci". The pre-run critic's BLOCKING 2 was that this establishes nothing about whether the Arabic there is plain: a segment can be patterned in a way no locus list happens to name. The audit was written — the same three criteria the PATTERNED domain is built from, applied to every candidate — and it excluded one of the eight selected segments (الصيد والقنص, a near-synonym doublet), replaced by the next eligible segment under the same mechanical rule. Fires at: every contrast where one arm is the treatment and the other is everything else. Remedy: the not-X arm gets an explicit, scripted, outcome-independent test for X, and the exclusions it makes are recorded on the page. RS-20260814g-supplied-sound §2; E-20260814g §10 finding 2. |
S185 | S185 |
| (bnf) | A positive control must be checked against the CODING INSTRUCTION that will be applied to it, not only against the hypothesis — an instruction can define a control away, and then the control tests nothing. E-20260813e froze C4, nineteen sites of lexical honorifics (拝領, 仰せ, 殿様), as its positive control at a bar of 0.85: these are marks whose content the English must render, so they must be carried. The same design's coding instruction said a word the English must use in order to say what the sentence says is not a device. The two clauses contradict each other, C4 came back at 0.082, and the run reached its analysis with no working evidence that its coders could see a device at all — which would have forced every number to be withheld. The pre-run critic, which caught three blocking findings including one that improved the design, did not catch this. Fires at: every design whose positive control and whose verdict definition are written in different places, which is most of them once a critic pass has rewritten one of the two. Remedy, cheap and mechanical: before freezing, take each control site and read the verdict definition against it aloud — if the definition's own words settle the verdict without looking at the data, the control is not a control. And the recovery that worked here is worth keeping as a standing fallback: a mutation control built after the fact — insert the device the verdict is defined as, at a handful of sites, re-code blind, and require the coders to flip them. It caught 6 of 6 on both seats with one false alarm, and it is what made the run reportable. RS-20260813e-slot-typology-ja §5. |
S176 | S176 |
| (bnb) | A calibration gate that passes at ceiling has not calibrated anything in the range the primaries occupy — build the gate so it can fail short of ceiling, or it licenses nothing. E-20260813c's gate G1 required the deliberately ennobled arm to read above the plain arm by ≥ 0.50 on a −1/0/+1 register scale. Realised: +1.0000, with the plain arm coded 0 and the ennobled arm coded +1 in every one of 56 cells — two seats × 14 sites × two arms, zero variance. The gate passed as wide as the scale allows, and the quantities it was gating — three published hands at +0.125 to +0.70 — all sit in the interior, where the gate demonstrated nothing. A pass at ceiling establishes that the instrument detects a large deliberate manipulation and is silent about its resolution among small natural differences, which is normally the measurement being made. Kin to Tier D's own dose logic, which already knows this: RS-20260802 reports detection at ceiling at 8 sites and reads the light dose precisely because a ceiling cell cannot discriminate. Fires at: every design whose manipulation check uses a deliberately maximal arm. Remedy, registered before dispatch: include a half-dose arm — the same rule set applied at reduced strength — and require the gate to separate it from the baseline, not only the full-strength arm; and where a full-dose gate does pass at ceiling, say on the result page that it licenses detection and not resolution. RS-20260813c-ennoblement-direction §7. |
S174 | S174 |
| (bna) | State an agreement-gated bar over the AGREED denominator, or seat disagreement is scored as manipulation failure. E-20260813b's manipulation check required "the plain document is matched to R06 at ≥ 11 of 15", where the match is a two-seat agreement. Realised: of the 14 segments that had a document, the seats agreed on 10 and all 10 went to R06, none to R08 — a unanimous direction — while 4 splits and 1 missing document counted against the bar exactly as a wrong-direction match would have. The bar failed at 10 of 15 and voided the primary. Fires at: any gate whose statistic is k-seat agreement but whose threshold is written over all units. Remedy: write the threshold as a proportion of the units where the gate's own statistic is defined, and register the minimum agreement rate as a separate clause — agreement on ≥ X of N, and of those, ≥ Y% in the predicted direction. The two failures are different failures and a single fraction cannot tell them apart. Not a licence to move a bar after seeing it miss: E-20260813b refused exactly that override (RS-20260813b §6), and this note is what the next design writes before dispatch. RS-20260813b-affect-yardstick §6.2. |
S173 | S174 — FIRED AGAIN, IN A NEW PLACE, IN A DESIGN THAT HAD APPLIED THIS NOTE'S REMEDY ELSEWHERE IN THE SAME DOCUMENT. E-20260813c cited this note by name and wrote its seat-coherence gate G3 over sites that returned a usable body — correct. It then wrote its completeness criterion F3 over all 42 dispatched bodies, while the primaries are computed over the seats surviving G3. All five dead bodies belonged to P5, which G3 dropped anyway, so the two seats carrying every primary returned 28 of 28; F3 nevertheless fired at 11.90% and voided the run. The generalisation this second firing forces: the denominator rule is not about agreement gates, it is about every threshold in a design — each one is written over the units its own statistic is defined on, and a design that fixes one bar must sweep the others in the same pass. No override was taken. RS-20260813c-ennoblement-direction §5.1. |
| (bmy) | A statistical routine copied forward from run to run has never been tested outside the (k, n) region the earlier runs happened to land in — recompute it a second way over the whole region the NEW run can reach, before dispatch. clopper_pearson was written for E-20260811h, copied verbatim into E-20260812e, and copied again into E-20260813a. Checking the pre-run critic's arithmetic meant computing the same intervals by bisection on the exact binomial tail, and the two disagreed: clopper_pearson(1, 60, 0.10) returned [1.000000, 1.000000] against a true [0.000855, 0.076640], and clopper_pearson(0, 60, 0.10) returned an upper bound of 1.000000 against 0.048703. The cause is a missing symmetry guard — the incomplete-beta continued fraction converges only for x < (a+1)/(a+b+2) — so the routine is wrong for k ≤ 1 whenever n is above about 50. No published figure is false: every k ≤ 1 interval in RS-20260811h and RS-20260812e is at n = 30, where it is correct, and all five were re-derived and agree. But RS-20260812e's K7 and K8 came back at 0 of 30 and this run's primaries run at n = 60, so the next step of the same programme walked straight at the broken region. Fires at: every run that inherits an analysis helper. Remedy: the verifier's independent implementation is not a formality to be run after the analysis — sweep it against the analysis routine over every (k, n) the design can produce, at design time. Here that is 272 pairs, 0 disagreements after the repair, 4 before. E-20260813a-world-dose/design.md §A2.5. |
S172 | S172 |
| (bme) | A stemming or normalisation rule imported from a previous run is exercised against the actual vocabulary of the NEW materials before it is frozen — and a rule stated as language-independent usually is not. RS-20260810w §5 introduced fix_stem — the modal 6-character prefix over words of length ≥ 5 — as a post-hoc repair for seats answering at different granularities, and called it mechanical, identical for every hand and every language. E-20260811d's pre-run critic found it is not: German Mantel and Mäntelchen reduce to mantel and mäntel, two stems for one lexeme, for a reason about the umlaut and not about the translator. The repair — NFD plus combining-mark deletion before the prefix, applied identically to every hand — was registered before dispatch and moved exactly one number, by 0.0666, and it changed the German hand from second-least fixed of four to joint first. Rule: an imported normalisation is run over the new run's own answer space (or, before answers exist, over the target passages) and the classes it merges and splits are inspected, before it is registered. Generalises (blz) from grounding rules to stemming rules; the same failure, one step earlier in the pipeline. RS-20260811d-title-network §4.2. |
S159 | S159 |
| (bmf) | A control item can be deleted by the very hand it is controlling — pick control lemmas the hands demonstrably render, and read a control miss before calling it seat incompetence. E-20260811d's F1 gate scored two seats at 9 of 10 on Field 1916, and both misses were the same site: «швейцарскую», which Field does not render — his sentence is "All his colleagues hastened to see his splendid new one" and the porter's lodge is not in it. Two seats reported a real deletion and the competence gate scored it as their failure. No cell was voided here (the bar was 8 of 10), but the same pattern voided a cell at E-20260810w, where Field's was the one voided cell and the result page had to carry the confound for a session. Rule: a control lemma is checked for presence in every hand's passage before the gate is registered — a one-line substring test on materials that already exist — and any control miss is looked at before it is counted, because a hand that deletes is the hand a census most wants to measure. RS-20260811d-title-network §6. |
S159 | S159 |
| (bly) | A new arm is never compared with a baseline generated in an earlier call, and the price of doing so is now measured: 0.167, larger than the effect under test. E-20260810z was frozen comparing a new source-first +I arm against the −I arm of the previous session; the pre-run critic's BLOCKING 1 said that conflates the manipulation with independent-pass variability and demanded a contemporaneous −I arm at the same clause, model, temperature and format. It was added, and on one judge × hand block the two −I passes — identical instruction, identical items, identical temperature, six days apart — returned loc of 0.200 and 0.0333. Against the old baseline the effect would have read +0.10; against the contemporaneous one, +0.2667. Rule: every arm a difference is read across is generated in the same dispatch stage, and a frozen arm from an earlier run may anchor a REPRODUCTION but may never be one side of a new comparison. The corollary for reuse: reusing frozen output is cheap and legitimate, and it is legitimate for exactly one thing — measuring the same quantity again. RS-20260810z-idiom-reach §2, P1d. |
S155 | S155 · S156 — honoured: the lead's frozen arms are REPORTED beside the gate corpus and are not one side of the primary comparison |
| (blz) | A text-grounding rule is a filter on the markers it will meet, and it must be tested against every marker class in the run before it is registered. E-20260810z amendment A6 required a judge's named marker to be present in the English shown, matched whole-word over tokens of three characters or more. That rule is correct for every arm the primary used and wrong for exactly one: the respelling arm's markers are an', o', 'em, which the three-character floor deletes, so a judge quoting the marker verbatim scored ungrounded and the arm's rate fell 0.55 → 0.167. The defect was invisible until the data arrived and would have been visible in five minutes against the previous run's stored loc strings, which were on disk. Rule: a grounding, matching or normalisation rule is exercised against the actual answer strings of the run it is copied from before it is frozen — the corpus that tests it already exists whenever the instrument is being reused. Second-order: the run then had to resolve which reading a gate was read on after seeing both, which is note (blw)'s failure arriving by a new route. RS-20260810z-idiom-reach §7. |
S155 | S155 |
| (bli) | A sense that measures markedness cannot serve as an execution or uptake measure unless a matched-oddity negative control runs in the same jury, on the same sense, in the same run. E-20260809h declared perceived-source-carriage as its execution measure and, on the pre-run critic's BLOCKING 1, replaced an arbitrary bar with a within-run control: CLUNKY, the lead's clean translation with eight word-order manglings proved by the verifier to be strict word-multiset permutations — not one word added, removed or duplicated. CLUNKY scored +1.4583 of carriage above the translation it permutes, higher than five of the six hands' foreignizing arms and above their median of +0.9583, so the registered gate fired and P1 was withheld. Without that control the primary would have PASSED: the six carriage gains are all positive, mean +0.8299, exact sign-flip P = 0.03125, the minimum attainable at n = 6 — the run would have published a written programme is executable by hands that did not write it and would have been reporting that odd English is odd. The mechanical corroboration built in the same amendment ran ρ = −0.4287 against the jury measure, so the two disagree about which hand executed most. Rule: before a sense is allowed to certify that a manipulation happened, put a manipulation-free oddity through it and require the arms to beat that, not zero — and where the sense has never been through Tier D, say so beside the number. RS-20260809h-rule-execution §2. |
S146 | S146 |
| (bji) | A control that cannot fail as evidence about the SUBJECT may still be the only thing that measures the INSTRUMENT — re-purpose it, do not kill it. E-20260805c's pre-run critic found that DECOY — a lead-written paraphrase required to mark nothing — could not fail, because it was built not to mark; the prediction "DECOY recovers at ≤ 1 site" was therefore no test of R1, and the arm was withdrawn before dispatch (amendment A2, S111). The critic was right about the prediction and wrong about the arm. With DECOY gone, every remaining arm preserved the passage's content, the unbriefed paraphrase recovered at 8 of 8, and the run concluded that the procedure might score content rather than marking — a conclusion that stood for five sessions. E-20260805h (S116) dispatched five of those same DECOY spans unchanged, as STRIP: they scored 0.222 against the filed translation's 0.861, and the conclusion inverted. An arm whose floor is known is worthless as evidence about the subject and is exactly what calibrates the ruler, because a measured floor is what turns a high score into a signal rather than a ceiling effect. Rule: when a critic shows a control cannot fail, ask what it would measure if the question were about the instrument instead, and keep it under that heading — withdrawing it removes the run's only lower bound. Companion to (rr): this is a case where an accepted BLOCKING finding cost more than it saved, and the cost was invisible until the arm was run. RS-20260805h §5. |
S116 | S116 |
| (bjj) | Put a byte-identical repeat of one arm in any design whose finding is a DIFFERENCE between arms; it is nearly free and this project ran the same procedure four times without one. E-20260805h graded each site's filed translation twice, under two letters, in the same body — the REPEAT arm — and the two verdicts disagreed in 0 of 36 pairs. That single number is what licenses reading a 0.639 gap between two other arms as signal rather than as within-body noise, and no prior run of the E-20260802e procedure (S091, S101, S106, S111) had it, so every difference those runs reported was interpreted against an unmeasured precision. The cost is one extra item per site on a payload already being dispatched. Rule: a design that reports a difference between arms carries a repeat of one arm, and reports the disagreement count beside the difference. Distinct from ordering controls, which measure position effects rather than repeatability. RS-20260805h §8. |
S116 | S116 |
| (bjk) | A paraphrase control carries the paraphraser's framing, and an unstripped control is not a control. E-20260806's floor arm N was generated by a seat told only "say the same thing a different way"; it wrapped its output in an assistant's preamble and closing note, and the runner's strip removed only the terminator. Both of N's two "different person" verdicts cite that framing and not a narrator — "Passage 2 explicitly presents itself as a rephrased version", "rewritten in modern language by an AI assistant" — and they fired two pre-registered gates (F3, F7) and withheld the primary. A declared post-hoc probe on the stripped text returned 0.000 at 8 of 8 cells where the unstripped arm had returned 1.625. Rule: any machine-generated control arm is stripped to its body before use, by a check that asserts the arm's first and last sentences are part of the text and not about it, and the check is in the runner rather than in the eye. Note (bix)'s third firing, and the first in which the paraphrase's advantage was an artefact rather than a finding. RS-20260806-same-man §2. |
S117 | S117 |
| (bjl) | You cannot ask readers is this the same person about a content-matched pair: the matching answers the question. R18 P5 and R19 Q5 require propositional parity, because without it a paired rendering is confounded by content — and parity is exactly what source-blind readers use to answer. Four seats, four arms, eight cells each, 27 of 32 cells at the scale's minimum, including for a rendering that changed 70.1% of the baseline's tokens and moved its measured register 2.33 → 6.00; a seat's own reason: "The passages share the same events, grievances, social setting… variations of the same narrator rather than different people." The control that makes the contrast valid determines the outcome, and no number of seats, languages or renderings repairs it. Rule: a design asking about narratorial identity over matched renderings uses a fit question — one rendering described against an independently authored source-side description — never a difference question over a pair. This retires the design wiki/goodness-senses.md §voice's reachability note specified from S097 to S117. RS-20260806-same-man §3. |
S117 | S117 |
| (bii) | A blind is not blind until recognition is measured, and a null memorisation probe does not license one. E-20260804c showed three non-Anthropic seats two Victorian-era English translations, source absent, arms lettered under a per-seat permutation, and asked who wrote them. Three of three named Constance Garnett and three of three named Isabel Hapgood — including the translator who is not canonical and whom no one reads. The same seats, given a held-out paragraph of the same story to translate cold, reproduced Hapgood more closely than Garnett at two of three seats: so the design's assumed mechanism, verbatim possession of the famous text, was measured and absent while recognition was total. Two rules follow. (i) Any design whose validity rests on a blind between published translations must carry a recognition probe and treat its outcome as a condition on the result, not as a footnote — prestige-confounding is the reason charter §5 demoted Tier P, and this is what it looks like when measured. (ii) Recognition and memorisation come apart, so a clean dependence_check does not establish that a seat cannot identify a text, and the contamination instrument this project already owns cannot be used to certify a blind. RS-20260804c-peer-record §3–§4. |
S103 | S103 |
| (bij) | When the lead's own artifact is an arm in a comparison, the lead must not adjudicate the text of a rival arm — replace the judgement with a rule, and let the rule overrule you. At S103 the lead hand-adjudicated eight single-token OCR divergences between two scans of Hapgood, whose translation was competing against the lead's own in the same run. Every call was defensible and the provenance was still wrong, as the pre-run critic said (BLOCKING, finding 6). The repair is cheap and mechanical: decide each divergence by frequency of the minimal differing token across both whole scanned volumes, write the decisions to a file, and have the extractor apply that file and nothing else. The rule disagreed with the lead once (it preferred a systematic scanning artifact, factory -hand, 3 to 1 — so frequency adjudication favours systematic OCR damage and that limit belongs in the script's header) and tied once, on a token that sat inside a stripped footnote and so never needed deciding. A rule that can overrule its author is the point; one that cannot is a restatement of the judgement. workshop/experiments/E-20260804c-peer-record/materials/adjudicate.py. |
S103 | S103 |
| (bhc) | A tool that reports an EXAMPLE of what it measured turns its own stored output into a blindness hazard. tools/dependence_check.py prints and stores max_run_text, the actual shared string, in every dependence.json a contamination gate writes. A later session reading a stored gate file to find out whether a figure exists is therefore shown archived translation text — which is what happened at S081, before the design that needed the blind had been written, on one of three pairs. Rules, in the order they cost least: (i) decide what you must be blind to before surveying the archive, because the survey is where the breach happens; (ii) when a breach occurs, declare it in its own commit before the affected artifact is written, name the exact tokens seen, and register a check on whether the seen text appears in the new one; (iii) state the breach's DIRECTION — here it pushed toward the reported result, not away, and it was bounded at 5 of the 32 shared 7-grams that would have been needed to reverse the sign; (iv) a blind seat that reproduces the same string from the source alone converts the leak into a forcing result, and two of the three leaked strings were disposed of that way. Do not adjust the artifact to avoid a string you have seen — that is fitting the work to the hypothesis, and it is worse than the breach. RS-20260801e-lead-carryover §5. |
S081 | S081 |
| (bgg) | A verifier's check count does not measure whether it read the response body, and neither does its mutation count. Eighteen frozen verifiers open a stored .raw. Truncate the accepted body they verify to 40% and mark it finish_reason: length — the shape (bdt) names — and four exit 0 reporting a full clean check count: 34/34, 50 checks / 0 failures, 207/207, and 6 of 6 mutations behaved as expected. The last one is the note: a verifier's own mutation tests can all pass while it is blind to the body it is verifying, because its mutations perturb the analysis and not the input. All four are silent under a non-truncating content perturbation as well, so they are not reading the content at all. Every result page in this project quotes a check count as its warrant. The count warrants less than it looks like. The repair is an assertion on finish_reason and on content length against the stored figure; the demonstration is a mutation of the body, and a grep cannot substitute for it. |
S075 | S075 |
| (bgh) | A crash is not a catch, and separating the two needs a second mutation. When a verifier is re-run against a corrupted body, a traceback and a failed check are not the same evidence: a script that raises on a malformed tail would not have raised on a plausibly truncated body, which is exactly the case (bdt) warns about. The separator is a second mutation that perturbs the content without truncating it. Run both: at S075, 7 verifiers crashed on truncation, of which 5 also crashed on a clean perturbation (genuine content sensitivity, expressed as a traceback rather than as a check), 1 caught it and 1 was silent. Without the control all seven would have been scored as sensitive. The pre-run critic named the confound; the control is the answer to it. | S075 | S075 |
| (bgn) | A failure criterion written as the null comes within ε of the instrument does not fire when the null WINS, and reading it literally lets a worse outcome through than the criterion was built to block. E-20260801's F4 voided a reachability claim if the prose-only leak null came within 0.05 of the panel. The null finished 0.1176 AHEAD (0.5294 against 0.4118), which the wording does not cover. A one-sided proximity criterion needs the dominance case written into it: fires if the null is within ε OR ahead. The session applied the stronger reading and recorded the defect rather than taking the literal pass. |
S077 | S077 |
| (bgo) | Quoting a defect back at yourself does not prevent it. E-20260801's design quotes, in its own §3.3 and in its arm page's standing constraints, the two prior occasions on which a synthetic control in this line of work was mis-built — S064's, and S067's replacement for S064's, mis-built the same way. It then mis-built one anyway, and the pre-run critic caught it as BLOCKING: a control expected to derive P derives S under the instrument's own definition of bearing, which would have voided the whole run through the design's own F1. Three occurrences, three caught only from outside. The operative remedy is not a stronger warning in the design; it is that the critic is pointed at the control block first, in the prompt, which is what produced the catch here. |
S077 | S077 · S078 — first affirmative application, and it worked twice over: the critic's control verdicts were item-specific rather than blanket, and its ADVISORY finding 4 PREDICTED the control that then failed at 3 of 3 seats |
| (bfg) | A whole-text POLICY read as a site choice is now a THREE-TIME defect, and it survives frontier models that agree with each other on everything else. E-20260730i condition A planted two synthetic entries describing a policy for the whole text with three renderings named as examples. Both raters counted them as site choices — 4 of 6 controls each, the same two failed, and the two raters otherwise agreed at 0.8095 exactly. RS-20260730h §3 found the same class in four of the 39 published R07 D codes, and S067's pre-run critic found it in that session's own control. Any instrument that counts decisions must carry a policy-versus-site control, and a rater's agreement with another rater is no evidence at all that both drew the line in the right place. |
S068 | S068 |
| (bfn) | An arm page can carry a stale premise, and nothing in the mechanism checks an arm page against the results it cites. ARM-anchor-second-read named three anchors to audit; two of the three had already been audited at S015, string by string with a blind arm and a decoy-controlled adjudication arm, by an experiment the arm's own evidence table cites two lines above the list that contradicts it. The list was written from the arm's intent, the table from its evidence, and they were never reconciled — for the six sessions the arm sat unworked and led the staleness block. wiki/tracks.md says an arm page exists "so its first session does not buy them again", which is exactly the guarantee that failed. The cheap remedy is the one that worked: read the results an arm cites BEFORE taking its step list, not after. Sibling of (bdk) — a page's prose and a page's tally are two readings and nothing makes them check each other — one level up, between a plan and its own evidence. |
S069 | S069 |
| (bfp) | A rate is a classification decision with a number attached, and three sessions running have now found one reported without its rule. S067: a coverage rate that was measuring the translator's own option list. S068: a DECIDES zero that was measuring the length of that list. S069: A-sasaki-kuroneko's pronoun repayment rate of 33%, which is 5 of 11 sites if only the imported pronoun counts as repayment and 7 of 11 if lexical noun repetition does — and C1, the claim the rate exists to support, says repayment is transcoding into lexis, so the second reading is the one its own theory licenses. The page states the alternative in prose ("a bare noun … or nothing at all") and never counts it, so the reader cannot see the number turns on it. Neither figure was wrong; printing one without the rule was. The first two instances were in the framework and this one is in the evidence base, which is what makes it a note rather than a repetition: the defect is not confined to the object the project was building, it is in the thing it was building from. Rule: a rate carries the classification rule that produced it, in the same sentence, or it is not a rate. |
S069 | S069 |
| (bfl) | Splitting one judgment into two separately-elicited axes does not make the axes independent; the same rater still divides a fixed quantity. RS-20260729g asked, in separate calls with counterfactual clauses written to break the dependence, whether a site is a matter of form and whether it is a matter of referent: pooled Spearman(F, R) = −0.358, per-rater −0.41 / −0.32 / −0.23. Every statistic built on (F − R) or min(F, R) is therefore computed on a stretched variable. The obvious repair — different raters per axis — costs the within-rater median split such designs are built on, so there is no free fix and this is a standing constraint on any multi-axis rating instrument here, not a task. Discharges the S059 backlog row at age 10 by retirement, with the measured figure kept as a prior a future design must address or be defective on arrival. |
S069 | S069 |
| (bfm) | A sense-agreement figure is not comparable across source languages, because a sense that points at morphology has nothing to point at in a language without it. RS-20260729g §3: style-correspondence reproduces at α 0.863 on Korean sites and 0.539 on classical-Chinese ones, same three raters, one shuffled list, neither text named; the categorical instrument moves the same way (0.72–0.81 against 0.55). The same fact explains why S049's pre-registered person-reference stratum came back empty on the classical-Chinese text. The comparison is confounded four ways — language, period, genre, and which model chose the sites under which brief — so this is a constraint on reading such figures, not a finding about Chinese. Do not pool or compare sense-agreement across source languages without saying what the sense has to point at in each. Absorbs the S059 backlog row at age 10. |
S069 | S069 |
| (bfi) | A census page declares its own under-count in the log; it does not report a completeness figure. The completeness control the project owed from S058 was attempted and the instrument failed — E-20260729f's determinate-list control returned 0.667 on a task with a right answer — so a census denominator cannot currently be certified complete by any panel this project reaches. What is owed instead is one sentence on the artifact: decisions made without noticing are not in the list, so the denominator is a floor and every rate over it is an over-estimate. Discharges the S058 backlog row that carried the un-buildable control. |
S068 | S068 |
| (ber) | A third option offered and never taken is not a neutral extra; it is evidence the instrument's response space is narrower than the design believes. E-20260730d offered UNCLEAR on 266 cells and both raters used it zero times, so a three-way judgement behaved as a two-way one. Second independent instance: RS-20260729g-graded-senses offered both and composite 240 times across two source languages and no rater took either once. Two designs, two different third options, one outcome. Report the take-up rate of every option a design offers, and treat a zero as a finding about the instrument rather than as tidy data. NARROWED 2026-07-31 (S072, RS-20260731d-sense-axes): the zero was partly a fact about the SLOT, not about the raters. Both prior instances offered the third option inside the label slot, competing with the answer. E-20260731d asked the same question — is one of these two descriptions enough, or does it need both? — as its own separate field, and take-up went from 0 of 240 and 0 of 266 to 20 of 120 (16.7%), landing on seam items only: 0 of 12 on each control pole and 0 of 18 on the foils. One rater in the same run also wrote BOTH into the two-way label slot on four items — exactly its four separate-field BOTH items — i.e. manufactured the option the slot withheld. The note's own cases stand and its rule is unchanged; what is added is a repair: before concluding that a third option is unused, offer it once in a field of its own. And the narrowing comes with its own limit — take-up was P1 15, P2 4, P3 1, α(needs) = 0.048, and no item drew it from all three raters, so the option became reachable without becoming agreed. Sibling of (bfq): reachability is a property of the rater. |
S063 | S072 |
| (bfw) | A definition that names two contrasts needs a positive control with one arm per contrast, because it can turn out that only one of them is doing the work. E-20260731d built its control on wiki/goodness-senses.md's own words — style-correspondence is local and formal, voice is global and cumulative — and scored both clauses separately on eight sites built to be unambiguous. The label half passed 8 of 8 with all three raters unanimous on every item; the scope half passed 5 of 8 and failed the registered threshold. Formal-versus-personal separates the two poles by +2.917 of 4 (p = 0.0001), scope by −1.000 (p = 0.027) — same direction, one fifth the size, on sites whose label nobody disputed. A single conjoint pass/fail control would have reported only control failed and the run would have carried nothing; splitting it by clause turned a void run into a located defect in the page's wording. Rule: enumerate the properties a definition claims, give each its own control arm and its own threshold, and report the arms separately even when the conjunction fails. Sibling of (ben), which says the control must be built before the threshold it tests. |
S072 | S072 |
| (bfx) | A leak null that does not condemn can miss condemning by almost nothing, and the verdict without the margin is the misleading half. E-20260731d registered that its run would be void if a mechanical classifier over the lead's own item prose matched the raters as well as the raters matched each other. Raters↔raters 0.7574; a frozen hand lexicon↔raters 0.7222. The rule says not void, and the true statement is that 0.035 is the entire margin by which the run is about the object rather than about the lead's wording, which every figure downstream inherits. The run's other null was worse in the opposite way — a nearest-definition cosine scored 26 of 40 items neither and agreed with the raters at 0.127, a degenerate null whose reassuring number is an artifact of its own distribution. Rule: quote a null's margin and its realised distribution beside its verdict; a null that cannot reach the raters' level, and a null that nearly does, both read as passed. Sibling of (bdq) — a statistic computed where nothing can vary measures the silence. |
S072 | S072 |
| (bfy) | A critic asked to endorse the lead's control items sight-unseen will endorse them, and that is now two consecutive sessions in which the blanket endorsement was the run's weak point. S072's critic endorsed all 8 control items and the control then failed. S073's critic endorsed all 8 and two of the four CONSTRAINED items were not constrained: «Walek», a personal name, drew 3 / 2 / 2 renderings because transliterating a name is a decision, and «Pietnaście», a numeral, drew 3 / 1 / 2 because seats pulled the following measure noun in. F1 passed at 2 of 3 seats on that item list and at 3 of 3 on the two items whose class is not in doubt (post-hoc). The endorsement is cheap and it is not a check — the critic is reasoning from the same intuitions as the lead about what "one obvious rendering" means, and neither of them ran the item. Rule: a control item is qualified by DATA, not by a second opinion — either by a prior run in which it behaved as declared, or by building more control items than the criterion needs and reporting which ones behaved. Sibling of (bfw), which says give each clause of a definition its own control arm. |
S073 | S073 · S078 — the remedy this note PRESCRIBES was applied and FAILED: six items qualified by a prior run's data returned 3/6 and 4/6 from the same two seats that scored 6/6. See (bgr) |
| (bfz) | Completion length moves on a byte-identical request and it is free to report, so report it. workshop/experiments/README.md's cross-day drift gate asks for it and S073 is its first affirmative application. Across an unchanged request in the same session, P1's completion went 1512 → 1623 and P2's 2532 → 3171 while both returned the same 32 item lines; P1's mean statistic moved 0.000 and P2's moved 0.375, which was larger than any effect that design was built to detect. Without the repeat the run would have published "the framework raises the option count in two of three seats", and it would have been wrong to. Rule: a design reporting a panel figure runs one condition twice in the same session and reports the delta beside the headline — and reports completion length per seat, which is already in every stored usage block. |
S073 | S073 · S078 — the repeat moved k by 0.333–0.667 in EVERY arm, all one direction, against a four-arm range of 0.714 |
| (bga) | A FIFTH dispatch failure shape, and it is note (bfb)'s without the parameter that (bfb) blames. x-ai/grok-4.5, finish_reason: stop, content: null, the answer's first item only inside message.reasoning, 1,001 completion tokens — with no reasoning: {"effort": ...} set at all, which is the remedy (bfb) prescribes. A plain retry of the identical payload to the identical slug returned all 32 items at 3,184 tokens. So the shape is not caused by the effort parameter and is not fixed by changing it; it is a truncation of the reasoning stream that reports itself as stop. Both this and the session's other wasted dispatch (a critic call returning finish_reason: error after 4,019 reasoning tokens, note (b)'s nineteenth firing) billed nothing — the key-usage delta matched the per-request sum to 1e-9 — which is S070's outcome and not S069's, so (bfo) stays withdrawn. Rule: read message.reasoning on every null-content body, count the items in it, and retry once before falling through — a partial answer in the wrong channel is not a seat failure. |
S073 | S073 |
| (bes) | Stripping authorship from a translation does not blind a rater to the WORK, and on a canonical source the work is what carries the confound. RS-20260730d §5 asked both raters, after their census answers, whether they recognised the passage. Both named T-postmaster-R07-v1 — written that session, from the Bengali, 11-token longest run with the only reachable published English — as "The Postmaster by Rabindranath Tagore". Neither recognised any of the four CC-licensed non-canonical stories. A-yosano-yomogiu, A-beowulf-ingeld, A-chekhov-pari and the Tier D held-out arm all stripped authorship from a lead translation of a canonical work and called the result blind. Nothing shows recognition changed a judgment there; nothing shows it did not. Ask the recognition question in every rating pass and publish the rate — it costs two output lines — and prefer openly-licensed non-canonical material where blinding has to hold. Sharpens ARM-evidence-audit's standing constraint from an assertion into a measurement with a mechanism, and gives D-20260725-06 a second independent reason. |
S063 | S063 · S067 |
| (bem) | A positive control's verdict is not portable across sessions, because its criterion is a threshold on a quantity the instrument re-draws every day. RS-20260730c-revision-close §§2–3: S057's F3 needed E(surface) to move 12.87 → 20, a gap of 7.13; the same instrument's own cross-day movement on 216 byte-identical items is up to 8.83. A control rebuilt in a later session cannot be credited with a movement smaller than the instrument's own drift, whatever the control's verdict says. The operative rule: an attribution claim must compare the effect against the measured drift, not against a tolerance chosen in advance — and a tolerance larger than the effect is a gate that cannot fail. Found by an independent pre-run critic, as BLOCKING finding 1, before the calls went out. |
S062 | S062 |
| (ben) | A control the designer builds after seeing the threshold it failed is selection-confounded and no amendment repairs it; the fix is to have it built by an agent that cannot see the threshold. E-20260730c critic finding 2, BLOCKING, accepted at the cost of eight already-frozen control items. The blind builder was given the corpus's span statistics and the required operation and not the axes, the thresholds or any prior score; the registered rule the first N well-formed items in the order returned was fixed in the script before the call, so the lead selected nothing. Realised span match: median 2.0/2.0 and mean 2.800/3.200 against the corpus's 2.0/2.0 and 2.717/3.195 — closer than the lead's hand-built set managed. The residual confound is declared rather than dissolved: the builder's prompt was written by a lead that had seen the scores. |
S062 | S062 |
| (bej) | A one-clause repair to a rule cannot be tested against the unrepaired rule alone, because the repair's mechanism is already known to work. RS-20260729d had measured that a fully mechanical rule out-agrees a warranted one 0.933 to 0.452, so any rewrite that removes a judgment from a clause was guaranteed to raise agreement, and a two-condition re-run would have re-measured the known thing and called it a repair. The design's answer is a matched groundless rewrite of the same clause — same lookup burden, same length to within one word, same position, criterion connected to nothing in the evidence base — so that the treatment is compared against mechanisation without aptness rather than against no mechanisation. Where an intervention's mechanism is already known to move the outcome, the control must hold that mechanism fixed. And the honest discount, registered before the run: the control was written by the session that wanted it to succeed, while the treatment's wording was fixed by an earlier session. |
S061 | S061 |
| (l) | A control has to be built and measured, not asserted — nobody finds an archaizing sham by reading the design. | S020 | S040 (the sham/targeted edit mismatch of RS-20260727b-tierD-rules §3 was found by constructing a sham by hand; two critic passes and a 71-check verifier had read the design and not seen it) |
| (bdb) | A fall-through chain needs an acceptance test, not an emptiness test. S049's extractor runner fell through on empty content and would have accepted a body truncated mid-object at finish_reason: "error" — 4,776 characters of valid-looking JSON with no closing bracket. Accept only on finish_reason == "stop" AND a structurally complete payload; a model that returns half an answer is a failure that looks like a success, which is worse than one that returns nothing. |
S049 | S049 |
| (bdd) | Before tuning a threshold, check whether the population contains anything the threshold could ever exclude. S050 was asked to settle a number for a gate certifying that a juror can express a 0.75-point margin. On all six juror-run cells the project has, every juror clears that margin when there is damage to find — so any threshold above the lowest observed value fails a demonstrably capable juror, and any threshold at or below it passes everybody. A gate with no incapable case is not a strict gate; it is a decision-shaped object, and the honest discharge is unrepairable, with the consequence stated, not a number. The cheapest refutation is one incapable juror, and the note says so on the page it came from. | S050 | S050 |
| (p) | A qualification instrument needs a case it should pass as well as one it should reject. S041 is the cleanest instance yet and it cost nothing: the arm-identifiability metric returned A = 0.833 on the arm it was aimed at and exactly 0.500 on the temperature control, where 0.500 is the value a working instrument must return — six longer, six shorter, on the same code path in the same run. The null was measured on real data, not asserted, which is the difference between an instrument and a number. |
S021 | S046 (A4 was registered in advance as run the same computation on the degenerate case, and the case it was pointed at was this session's own registered criterion — which failed; see (bct)) |
| (u) | A positive control must differ from the failing case on the dimension the rule measures. | S023 | S023 |
| (z) | A self-test gate catches type errors, not world errors. | S023 | S026 |
| (abg) | Freeze the artefact, then measure it — and download the comparator without reading it. A contamination declaration made after the comparator has been read is not a measurement; the blind is the whole instrument, and it costs one script. S041: Takahashi 1927 was downloaded to a scratch path before translating and never opened; both lead artifacts and both logs were committed first; the measurement ran afterwards and returned 8 tokens against a null-control floor of 5. | S033 | S041 |
| (ii) | A self-test fixture must be shaped like the target material. A single-sentence fixture cannot fail a landmark that only fails when earlier distractors exist. (Restored S032 after silent loss — see above.) S038: all eleven fixtures of dependence_check_cjk.py are Japanese, and one of them — two unrelated sentences still sharing 「た」 — is the fixture that states why a bare character-run length needs a reference floor. |
S025 | S038 |
| (rr) | ⚠ STREAK BROKEN AT S093 — the first session since S029 to run an experiment with NO independent pre-run critic pass. E-20260802g-programme-divergence was dispatched without one, declared at the head of its design, in its result's limits, in wiki/tracks.md and in NEXT.md's drift check. The reason was attention and budget spent on a ratification gate that overran; that is an explanation, not a defence. The note's rule is unchanged and is restated: every design gets an independent pre-run critic pass before any datum exists, and a session that skips it says so where a reader will see it. The thirty-eight-session record and what each pass bought is below, unaltered. [fired S066, TWENTY-FOURTH consecutive session: NEEDS-AMENDMENT, nine findings, three BLOCKING, all nine accepted. Its finding 7 replaced the primary statistic with a permutation null and showed that BOTH registered predictions would have passed under any no-effort null — the session would have reported chance as a finding. Its finding 2 added a second, unforced translation limb, which is where the session's actual result came from.] Put the critic on the statistic definitions, not just the argument. Six sessions running this produced the single most valuable finding, and six times it was the same shape: a sentence in the design asserting what its own statistic or rule controlled for, which was false. S034 is the first time the critic missed one — it proved the targeted and held-out rules were not "the same rule at the same threshold" and left the sham's band, one section away, uncomputed. Pointing a critic at some of the definitions is not the note; the note is to compute the null probability of every branch of every pre-registered rule. |
S028 | S020 |
| (bdg) | Priming does not only change what you notice; it changes what you write down — so a declared bias direction on a self-reported denominator is incomplete until you have asked whether the priming moves the denominator's size. S051's design conceded that the lead had read the candidate list before writing the translator's log it was about to measure coverage on, argued the bias ran upward (a primed lead spots inventory-shaped decisions), and concluded that a null was therefore safe to read. The pre-run critic's answer: a primed lead may also record extra fine-grained decisions no candidate reaches, enlarging the denominator and pushing the proportion down. Two uncontrolled effects in opposite directions, so neither a rise nor a null is interpretable, and the design's central reading was false before a number existed. The general form: when the denominator is something you generated, "which way does the bias run" has two answers — one about the numerator's recognition and one about the denominator's construction — and naming only the first is the error. | S051 | S051 |
| (bdh) | Predict how your own published figure will do under an independent check, and register the prediction, because being wrong in the reassuring direction is a result you will otherwise discard as unremarkable. S051 predicted that RS-20260726e's decision-to-claim mapping — a figure that page itself called "the weakest link" — would fail a second-reader check badly, on the basis of S048's independent classifier agreeing with lead coding at 0.483. It came back at κ 0.905 and 0.715, one disagreement of 21 from one reader. The registered prediction technically HELD (its threshold was more than 2 disagreements, and one reader gave 3), and its reasoning was refuted by an order of magnitude. Without the registration the session would have reported a passing check and moved on; with it, the project learns that its coverage instrument is more reliable than its own pessimism, which changes what later sessions may lean on. |
S051 | S051 |
| (abh) | Compute the null probability of every branch of a pre-registered rule, including the branches that mean "fail". §6.5's sham band was called three-way and carried verbatim across two runs; its lower branch fires with probability 0.534 on a null jury against the upper branch's 0.0046. A rule whose failure branch is a coin flip is not a control, and S020's sham sat one unit above it. Reachability in both directions (note (dd)) is necessary and not sufficient — the branch has to be improbable, not merely reachable. S040 sharpens it into a second half, and the second half is the one that bites: compute the branch probabilities AND the branch POWER. The 0.534 branch made the band fire when nothing was wrong; the 5-of-6 threshold makes it stay silent when something is — power 0.42 against a jury with a 70% edit-presence bias. Every review this project ran asked whether the rule fires wrongly; none asked whether it fires at all. See (bci). | S034 | S040 |
| (abi) | State a scale-usage or dispersion gate per juror, not on the pool. S034's §10 gate passed at pooled SD 0.783 against a 0.75 threshold while P5 used two adjacent integers across 48 scores (SD 0.143). A pooled gate that one juror can carry certifies nothing about the others. S040 measured why, and the number is the note: 56.8% of the pooled sum of squares was BETWEEN jurors. Pooled SD 0.783 falls to 0.515 once the juror means are removed — below the threshold it passed. A pooled dispersion statistic counts juror disagreement about the mean as evidence that each juror uses the scale, which is by construction the one thing it is not. | S034 | S040 |
| (bcs) | A cost statistic the arithmetic forces is not a cost statistic — subtract the floor before reporting it. S046 designed a length-matching pass and registered number of edit sites as the price of matching. An independent critic pointed out before the run that if the word count has to move, at least one site must change, so the statistic fires with probability 1 and measures nothing; the registered prediction built on it was near-vacuous. Replaced by **tokens_changed − |
Δ | , the tokens moved beyond what the arithmetic demanded — a pure trim scores 0, and the run scored 4. Before registering a count as the cost of a constraint, compute what the constraint forces on its own, and register the excess.** The general form: any statistic whose null is unreachable given the procedure is a description of the procedure, not a measurement of it. |
| (bct) | An interval-width criterion inverts at small n, and the tell is that the naive interval becomes WIDER than the clustered one. S046 registered fails iff the bootstrap interval is wider than 0.30 as its stratification criterion, after a critic correctly rejected a binomial interval on clustered votes. On six-item strata it fired as predicted (3 of 5 and 4 of 5 senses). On the two-item stratum it PASSED, with intervals narrower than the six-item ones (0.250 against 0.361) — because a bootstrap over two numbers can only resample two numbers and reports the spread of a two-point set as a population's. The diagnostic: the anticonservative binomial interval, which wrongly assumes 24 independent votes, was wider there (0.523). When the estimator that ignores clustering gives a wider interval than the one that models it, the clustered estimator has run out of clusters. Sibling of (o) and distinct from it: the threshold was reachable, and it was reachable for the wrong reason. Use a hard minimum item count, not a width — S046's specification sets 5, and the n computed from the item SDs for a usable interval is 22–30 per stratum against the 6 the project had. | S046 | S046 |
| (bci) | A control fails in two directions and this project has only ever checked one. A gate can fire when nothing is wrong (α) and stay silent when something is (power), and every critique instrument the project has built — pre-run critic passes, post-run verifiers, standing dispositions — is pointed at the first. S040 computed the second for the Tier D sham for the first time: the repaired 5-of-6 branch catches a genuine 70% edit-presence bias 42% of the time, so six units is not a control at any threshold. Before freezing a gate, tabulate its power against the effect it exists to catch, and state the n it would need. The answer is often "more units than the budget assumed", which is a finding about affordability and belongs in the design rather than in a later post-mortem. | S040 | S040 |
| (bcj) | A control and its treatment must be matched on every axis the subject can see without reading. Tier D's sham was word-count neutral by constraint (0 of 16 sites change length; 6 of 16 change no word at all, only order) while the targeted operator it bounds adds a clause at 10 of 24 sites — +44 words, ≈4% per item. A false-alarm floor measured on edits that cannot change length does not bound the false-alarm rate of edits that do. The design already knew length was a cue: it listed the 8–11% length gap as a threat for the held-out arm, where the materials imposed it, and missed the 4% in the arm where the design itself created it. A threat named for one arm is not thereby handled in the others. S041 measured the general form and it cuts both ways: on the S010 regime comparison the treatment arm is separable without reading (the revision came back shorter in 10 of 12 pairs, A = 0.833) and the jury did not detectably use it, while the control arm is exactly unseparable (A = 0.500) and the jury tracked length on all five senses. Matching is necessary and is not sufficient, and a matched control can still be dominated by the axis it was matched on. |
S040 | S041 |
| (bck) | Whether the arms of a comparison are separable and whether the judge uses the separation are two properties, and one instrument sees only one of them. A correlation between preference and the size of a difference detects a graded effect and is structurally unable to detect a sign-only one: where the difference has the same sign on every item and the judge responds to the sign, preference is constant, so Pearson's r is undefined (0/0) — not zero — while the arm is identifiable with certainty. The converse also occurs. S041 found both cells in one dataset: MAIN identifiable at 0.833 with the correlation silent, TEMP unidentifiable at 0.500 with the correlation firing on 3 of 5 senses. So run both — a sign count AND a correlation — on every paired design. The sign count costs one function call, and half the failure space is invisible without it. Two qualifications, both established rather than assumed: the blindness is a knife-edge at exact constancy (move one item and r returns to 0.585), and where the effect really is graded the correlation is the right instrument and works. | S041 | S041 |
| (bcl) | Declaring a criterion "soft" or "a judgment" describes it; it does not bound it — and every ratio computed over an unbounded criterion is a ratio over one reader's list. Criterion (iii) of the «Jeli» free-indirect inventories ("an idiom or evaluative epithet attributable to the community rather than to a neutral narrator") was declared soft in advance at S039, which felt like sufficient hedging. At S042 an independent critic, asked to check completeness, produced eight further qualifying sites in one span without trying to be exhaustive. So RS-20260727-jeli-fid's "P1 cost accuracy at 2 of 9 Class C sites" has no denominator, and the class was withdrawn rather than narrowed — narrowing moves the boundary without drawing one. Before a criterion is allowed to yield a count, ask whether a reader who is not its author could apply it and get the same set; if not, report named instances and no ratio. Cheap to check at any time: one critic call, three sessions late. Fired again at S043, which was designed around it: the whole unit asks whether a rule defines a set, measured as whether readers who are not its author land in the same place. S053 gives the note a number on the friendliest possible case. A site-inclusion rule was written out in full, with four numbered clauses and a list of worked exclusions, frozen in git, and handed verbatim to an independent reader along with the same 67 lines. The two enumerations overlap at Jaccard 0.44 — the independent reader missed 8 of the lead's 23 primary sites, including three of the four that carry the headline arm, and supplied 11 the lead had not thought of. So an explicitly written criterion buys under half a shared set, and a criterion that was never written down — which is what A-beowulf-ingeld §4.3's was — has no claim to do better. Fired again at S054, from the inside and on the friendliest case yet. The lead wrote an 80-site reflex census of a prose passage, with the inclusion rule stated in full, and then translated the passage itself — and the census had silently merged līefan "permit" with lǣfan "leave, bequeath" into one lemma, so the ¶1 site was never entered. The translator then took the reflex at exactly that site and got it wrong ("nor left it to other men" for nē … ne lēfdon), and only the revision pass caught it. A census written by the person about to translate is not a control on that person's translating: the sites it misses are the sites they will not be watching. |
S042 | S054 · S071 — fired on the translator's own diff count: log B10 said 21 of 28 paragraphs differed and the verifier found 25 |
| (bcm) | A repair prescribed by a result page must have its yield measured, and the case that motivated the repair cannot count as the repair's yield. RS-20260727-jeli-fid found one missed site, declared its inventory "a floor, not a census", and prescribed widening the criterion. S042 widened it and measured: over 4,301 words of frozen prose the widened criterion found 0 uncontested sites no previous reading had found, and 1 disputed one. Counting the motivating case as a find is circular; counting a retrospective find as evidence of prospective performance is a second error the same page nearly made. The prescription was still right to carry out — the null is a result, and it cost one session — but "the criterion was incomplete" and "widening it buys something" are different claims and the first does not establish the second. |
S042 | S042 |
| (bcq) | An elicitation whose safe answer is "I don't know" has a floor at zero, and a floor at zero is not a measurement. S045 asked three models to write out Constance Garnett's published English of eight Turgenev passages, with UNKNOWN offered for anything not recalled. UNKNOWN came back at 24 of 24 cells, so every recall score was 0 — and the same three models, asked instead merely to identify the passages, named Turgenev for 8, 7 and 8 of 8 and the individual 250-word prose poem for 6, 4 and 5 of 8. The floor was a refusal to guess, not an absence of knowledge, and the pre-run critic had predicted exactly that before dispatch. Design a knowledge probe so that a subject who knows nothing scores at chance and cannot decline — forced choice against a distractor, not free recall. The general form: the null of an instrument must be a value the instrument can reach for a reason other than the subject opting out. | S045 | S045 |
| (bcr) | A permissive threshold and a strict one on the same data are one measurement, not two, and the disagreement between them is the finding. S045 kept a previous session's absolute rule (≥7 of a run's tokens) as primary and registered a proportional one (≥⅔) as secondary, before the numbers existed, with the reason written down: the absolute rule is permissive on long runs, and permissive meant biased toward the verdict that would have overturned this project's own standing rule. The eight loci came back 6 forced on the permissive rule and 4 on the strict one — and the headline locus split, forced on one and not on the other, which is precisely what its prose description had claimed. Where a threshold is being carried across materials it was not written for, register both and report both; choosing the rule that favours the position you already hold is the degree of freedom this prevents. | S045 | S045 |
| (xx) | A bug in a shared tool outlives the session that wrote it, and a second implementation cannot see it — so a verifier imports nothing from tools/ and runs the tool's fixtures inside itself. S045: the 150-check verifier reimplemented the tokeniser, the metric, both verdict rules and the confound flag from the design text — and the check that earned its keep was not a statistic at all but a count, 15 probe calls recorded, which caught two API costs the runner's cache branch had silently dropped. A verifier that only recomputes numbers cannot see a number that was never recorded. |
S030 | S045 |
| (g) | Machine-verify every quotation, and the tallies over them, and your own corrections. S039: 61 guillemeted Italian quotations plus all 21 pre-registered inventory rows checked at their stated paragraphs against the stored source; 21 of 21 and all quotations pass. S038: 28 Japanese and 23 English quotations in A-sasaki-kuroneko checked against the stored files; one was wrong (「一緒について来た」 for 「一緒に来た」) and was fixed before the page was committed. |
S012 | S038 |
| (j) | Verify the journal's quotations too. | S019 | S024 |
| (bcu) | A rule stated over a corpus has to be checked against the corpus it describes — backwards as well as forwards. S047 found two register rows in «Jeli» that were wrong about frozen prose written before them, and found both by mechanical enumeration in under a minute after five sittings of careful reading had missed them. V7 (Verga's guillemets are kept as single inverted commas, written span 2) was never applied to the two span-1 sites, which carry the speech mark instead — and span 1's own log says in writing that it used inverted commas. V10 (reduplication is kept for adverbs of manner and degree; an adjective predicated of a person is rendered singly, written span 4) is contradicted by bianco bianco → "white, white as a dead man's" in span 2: 24 reduplications in the work, 6 predicative adjectives, 5 different treatments. The instrument was built to catch forward drift — a row set early and broken later, which is how N15 broke and was caught — and it is structurally blind backwards, because nothing rereads the finished spans. The remedy is cheap and is the note: when a row is written or revised, grep the whole artifact for the feature it governs before writing it down. Generalises past registers: any coding rule, criterion or definition stated mid-corpus. |
S047 | S047 |
| (bcv) | The option set of a forced-choice measure is a researcher degree of freedom, and a set missing an option manufactures the prediction. S047's design offered raters quoting / distancing / highlighting / cannot tell — the two-way contrast the project's own translator's log is written in, plus fillers. The independent pre-run critic said a proverb is borrowed language without being anyone's actual words and required a fifth option, fixed expression. That option took 12 of 30 answers and falsified the design's headline prediction; without it the gate would have passed and the session would have reported the opposite finding. Before freezing a forced-choice scheme, ask an independent reader what category the scheme cannot express — it is the cheapest of all the critic's jobs and it changed this result outright. Rider, and it is unfixed: the revised option set was still written by the lead. | S047 | S047 |
| (bcw) | A self-scored coding scheme is not a measurement until a second reader has applied it, and "does this rule bear on this feature" is exactly where two readers diverge. S047 froze two rule sets, translated one passage under each, coded 89 decision sites D/decided · P/permitted · S/silent, and then paid one non-Anthropic model $0.052 to code the same 89 sites blind to the lead's codes. Agreement 0.483 against a pre-registered 0.60 floor; κ = 0.173. The disagreement is not noise: 14 sites the lead coded D the classifier coded S, and 10 of the 14 cite only the two "follow the source" rules — the dispute is the single question of whether a rule that applies to everything names anything. The frozen scheme had no wording that settled it, and it also had no code for rule conflict, which surfaced independently at one site. Write the second reader into the design, and write the hard case into the code definitions before any site is coded. | S048 | S048 |
| (bcy) | "Do what the source does" is not a translation rule of the same kind as "be idiomatic," and a coverage count that treats them as peers will report the wrong thing. S047 wrote a foreignizing and a domesticating rule set from the same book and found the foreignizing one decided six times as often on the lead's coding — because two of its ten rules point at an object that exists (the source) while the domesticating set only excludes options without naming a target. The registered prediction went the other way and failed on both codings. When comparing rule sets, first ask which of them can be satisfied by pointing. | S048 | S048 |
| (bdq) | An agreement figure computed where nothing fires is a measurement of the silence, not of the instrument. RS-20260728i reported reader-versus-reader agreement at κ 0.809 / 0.714 on a three-way label and read it as a sound instrument. There were zero DECIDES in that run, so the figure was agreement about the other two labels. S056 added one candidate that can decide, changed nothing else, and sent the same 63 entries to the same two models: κ fell to 0.066. A reliability statistic must be reported with the distribution it was computed on, and a near-degenerate distribution voids it. Any future use of a label set must check that every label is reachable before quoting agreement — which is what a positive control is for, and is why this note is a sibling of (bdr). AMENDED 2026-07-31 (S070), RS-20260731b-figure-audit: the check this note prescribes is TOO WEAK, and it passes on a run where it should not. The prescribed test is was every label used at least once by at least one rater? On 43 sites built to elicit composite, all five labels are realised in the pooled run — the check passes — while P3 used two of the five labels and P2 used three. Reachability is a property of the individual rater; see note (bfq). The note's own case (κ 0.809 → 0.066) is untouched. |
S056 | S070 |
| (bdr) | A null needs a positive control or it is unreadable, and the control belongs in the same call as the null. E-20260729d was built to explain a zero and had no evidence the instrument could return anything else; the pre-run critic caught it and the fix — a rule that decides by counting words, added to the same candidate list — returned DECIDES on 58 and 12 of 63 entries and established the label as functional. Without it, both a zero and a non-zero would have been uninterpretable, and the session would have reported a finding about frameworks that was a finding about reader disposition. RS-20260728i's own zero inherited this defect and could not have detected it; it is readable now only because a later run supplied the control retrospectively. Cost of the control: ~90 words in a prompt. |
S056 | S056 |
| (bez) | Free recall and forced choice both fail to measure a model's memory of a published translation, and they fail in opposite shapes — so a "does the model hold the comparator" control cannot currently be built from either. S045 asked three models to write out Garnett's English as they recalled it: UNKNOWN at 24 of 24 cells, a refusal floor at zero. S065 asked the same three, forced choice, no decline option, matched rival, positive control at 18/18 on the King James Version: on Garnett they scored 0.467 (p = 0.856) with a 33% order-flip rate, and on a non-canonical comparator they scored 0.042 — 46 of 48 judgments picked the lead's translation as the published one, while naming the author at 8 of 8 and the individual work at 8 of 8 on the same passages. The second instrument does not decline; it answers a different question confidently — apparently which of these is the better translation rather than which was published. Two consequences bind. (i) Recognition of a work is not evidence of memory of its translation's wording, and this project has three measurements of recognition at or near ceiling. (ii) Before a "does the model hold X" control is trusted, it needs a negative arm on material genuinely not held, and this session's attempt at one inverted rather than floored — so no measured value exists for what these instruments return under true absence. | S065 | S065 |
| (bev) | A byte-identical, same-day prompt moved 24% of a classification instrument's answers, and the 7.2% figure this project quotes as reassurance was measured on a different KIND of task. E-20260730e's C1 pass was re-sent to P1 verbatim — prompt tokens 2,698 and 2,698, identical to the digit, same slug, same provider, same UTC day — and 9 of 37 codes flipped, in every direction, self-agreement 0.757 against a pre-registered floor of 0.80. S063 measured 7.2% on a byte-identical same-day repeat and called it the first repeat control in this project to pass; that was a PRESENCE task (is this feature in this passage), and this is a CLASSIFICATION task (which of three codes applies), and the two do not transfer. The S062 failures were cross-day; this one is not. Rule: a repeat control certifies the task shape it was run on and nothing else, and a classification instrument needs its own before any κ or α from it is quoted. |
S064 | S064 |
| (bex) | A mid-design amendment can make a label unreachable, and the resulting zero cell looks exactly like evidence. The pre-run critic's Finding 2 correctly showed that E-20260730e's KEEP label collapsed "requires" and "permits" and was indistinguishable from UNDECIDED; the fix redefined it as "F10's wording positively supports the figure". F10 is a prohibition, and nothing in a prohibition can positively support anything, so KEEP came back 0 of 12 — a striking-looking number that means nothing. Notes (bdq) and (beb) are about labels that never fire in the data; this is a label that could not fire by construction, created by the repair of a different defect. Rule: after amending a label set, ask of each label what answer would earn it, and if the answer is "none", the amendment has replaced one defect with another. |
S064 | S064 |
| (bff) | A pre-run critic shown CONCATENATED FILES cannot check freeze order, and is right to refuse to take it on trust. E-20260730h's critic prompt joined the frozen design, the frozen translation's tally and the rater prompt under three headings. Its first BLOCKING finding was that the design's registered predictions P5 and P6 were postdictions, because their outcome was printed above them. The premise was false — design.md contains no tally, and the freeze order is e602222 (design and predictions) → 0c812d0 → 8a11440 (the tally) — and the finding was nonetheless unanswerable from what the critic could see. A critic cannot read a repository. Put the commit hashes in the critic prompt, or accept that every freeze claim is an assertion to it and will correctly be treated as one. Cost of the fix: three lines. |
S067 | S067 |
| (bfq) | Label reachability is a property of the RATER, not of the label set, and a pooled realised-distribution check cannot see the difference. E-20260731b's positive control put 43 sites — 12 built to be composite — to the same three seats E-20260728f used, with that run's five-option block verbatim. All five labels are realised in the pooled run, so notes (bdq) and (beb)'s prescribed check passes. Per rater it is P1 five labels, P2 three, P3 two: P3 answered 43 items with two of the five options and P2 never used both or composite at all. Every composite rating in the run was P1's. A five-way instrument that one seat operates as a two-way one is not five-way, and the pooled check declares it sound. Rule: report the realised vocabulary PER RATER before quoting an agreement statistic over a fixed option set, and treat a seat that never uses a label as evidence about the seat. |
S070 | S070 |
| (bfr) | An alignment anchor must be priced for DISCRIMINATING POWER before it is built, and a block-type histogram is not evidence about an alignment. RS-20260730g §4 prescribed anchoring an RU↔JA paragraph alignment on dialogue structure after a length-based aligner drifted, and wrote that "the drift is invisible in the block-type histogram." E-20260731c built the prescribed repair. It is wrong at 3 of 7 non-1:1 blocks (0.429 against a registered 0.20) and at 8 of 68 blocks overall — and its histogram is character-for-character IDENTICAL to the correct alignment's: 61×1:1, 5×1:2, 1×2:1, 1×1:3 in both. So the histogram can be exactly right in every cell while eight blocks are misassigned. Why the anchor failed was free to measure and nobody measured it: 51 of 69 RU and 54 of 75 JA paragraphs are dialogue, so the anchor takes the same value on three-quarters of the text, and every one of the eight errors falls inside a run of same-type paragraphs. Rule: before building an anchor-based aligner, count the anchor's marginal distribution; an anchor constant on ~75% of the units carries almost no information exactly where alignment is hard. Sibling of (bfa), which fixed the void criterion; this one is about the instrument the criterion was voiding. |
S071 | S071 |
| (bfs) | A per-item adjudication does not enforce a partition, and a unit can be orphaned by every item that could have claimed it. E-20260731c's independent alignment check asked three seats, item by item, which Japanese paragraphs render a given Russian one. At the two adjacent items that between them exhaust JA ¶23's possible owners, P1 and P2 assigned it to neither — it renders nothing on their answers, and no item ever asked them to notice. The format cannot see this and the majority statistic scores it as an ordinary disagreement. Rule: when an adjudication task partitions a corpus, add a coverage check over the union of answers, or ask for the partition directly rather than item by item. |
S071 | S071 |
| (bft) | An exact-match rate over counts is carried by the cells where the count is ZERO, and the trivial matches must be reported separately. E-20260731c measured Futabatei's comma-count preservation at 0.290 over 62 aligned blocks and it reads as weak compliance. Fourteen of its eighteen matching blocks have zero Russian commas and zero Japanese ones — nothing was preserved, because there was nothing to preserve. On the 45 blocks with at least one source comma the figure is 0.089, and its relation to the null flips from near (p = 0.44) to below (p = 0.98): a translator ignoring the source commas entirely would have matched roughly twice as often. Rule: for any exact-match statistic over counts, report the subset where the source count can vary, and report how many matches are the zero-zero cell. Sibling of (bdq) — a statistic computed where nothing can vary measures the silence. This check was NOT registered in the design and it inverted the headline. |
S071 | S071 |
| (bfu) | A proper-name floor on a CJK dependence check is simultaneously too low to survive one inflected character and too high to see a real convergence. E-20260731c measured two lead Japanese limbs against Futabatei's 1888 Japanese with a same-work, non-corresponding-span reference cell — the sharpest floor this project has built. The unforced limb ran 14 against a floor of 14 and was declared none; the forced limb ran 16 and was declared high. Both runs are the character's transliterated name, and the entire margin that separates them is 「した」 attached to it after normalisation strips the comma. Meanwhile the longest NON-name shared run — 「けだそうとしたが足が」, 10 characters, written 138 years apart from the same Russian by a translator who had not read the comparator — sits below the floor and is declared invisible. Rule: on a CJK pair, strip or separately report proper-noun runs before reading a verdict, and report the longest run that contains no name. Companion to (bdw), which says know what the measurement means before reading it. |
S071 | S071 |
| (bfv) | The key-usage cross-check is void whenever another session is spending, and it fails silently by looking like drift. Note (bco)'s reconciliation compares a session's per-request cost sum against the key-usage delta. At S071 the key read 34.354965717 at the session's first call and 35.505091267 immediately before its second, a rise of $1.149 across an interval in which the session made no call at all; the session had also opened $1.198 above S070's closing figure. Two unattributable movements of over a dollar each in one session. The per-request usage.cost figures are unaffected and remain primary, as CLAUDE.md says. Rule: take an opening AND a closing snapshot around EACH call rather than around the session, and declare the cross-check void — do not report a residual as drift — when any inter-call interval moves. This is the concurrency hazard NEXT.md has carried since S048 appearing for the first time as a corrupted instrument rather than as a worry. |
S071 | S071 |
| (bgw) | A guard parameterised by a literal from the experiment that bought it fails closed on the next one. Note (bgt) was bought at S078 by a body reporting stop while truncated, and its remedy was move the item-count check into the runner. The runner carried it forward with S078's item-id pattern hard-coded — re.match(r"^\s*I\d\d\s*\|", line), I because that run's items were I01–I66. S079's items are Q01–Q22, so the guard counted zero answer lines in every body and rejected every seat, including any that would have been complete: four dispatches, $0.060159746, no accepted cell. It failed safe rather than silent, which is the only good thing about it. Rule: a check moved into a runner must take its discriminating literal as an argument, and the runner must be re-pointed at the new experiment's ids in the same commit that copies it. Sibling of (bgt), which bought the check, and of (bgu), which is the same disease in a verifier. |
S079 | S079 · S080 — fired again and correctly: line_re is now a REQUIRED argument of call.py::dispatch with no default, so an item-id pattern cannot be silently inherited. 3 of 3 seats returned 138 of 138 lines · S081 — line_re passed explicitly as the delimiter pattern for a translation task rather than an item-id pattern; the guard rejected three truncated bodies and admitted eight complete ones · S108 — SECOND FIRING, and it cost more than the first. E-20260804i's answer-line guard required each line to BEGIN with the site id; google/gemini-3.6-flash returned 24 perfectly valid lines prefixed SITE. The runner scored the body at 0 answer lines of 24, called it a truncated body, retried the same slug, and then burned two reserve calls — $0.126113 spent re-trying a $0.05795 body that was already correct. The body was recovered from disk afterwards (note (bdt) writes raw bytes before parsing) and is in the result. The rule (bgw) already states is not enough: a guard pattern must be tested against one real body of the shape it will police, from each seat, before the run depends on it — a dry run that only builds the prompt tests nothing. |
| (bgx) | An implementation whose threshold is looser than the proposition it implements will pass what the proposition fails — and S079 hit it twice, in two unrelated places, in one session. (i) E-20260730i's design.md defines CTRL-NEG as "a site where no candidate bears"; its analyse.py passes the item iff no candidate excludes. Two raters answered bears, excludes nothing, the code passed them, the design would have failed them. Adjudicated 2 of 3 for the design by independent seats — that run's F2 fires and k = 0.14 is not an estimate. (ii) E-20260731f's c_C08 implements "no ornamental attributing verbs are used" as say-rate ≥ 0.80; the anchor's second stored text scores 0.812 with six ornamental verbs, so it passes the check and does not satisfy the claim. Rule: when a check is written for a proposition, the check's own docstring must state where its operationalisation is WEAKER than the words, and any pass within one threshold-width of the boundary is reported with its margin, never as a bare PASS. A prose criterion and its code are two artifacts and only one of them was registered. |
S079 | S079 · S080 — a THIRD instance, and it was in this session's own design page. E-20260801d §3 recorded that a materials repair "moved no overlap figure at all (34/0/0/11 both times)". It moved none of the four columns dependence_check PRINTS and moved the name-excluded count 14 → 17, which that summary does not print. Corollary added: when a tool prints a subset of what it computes, "no figure moved" means "no PRINTED figure moved" and must be written that way · S081 — a FOURTH instance in four sessions, and this time the critic found it in the design before the run rather than the session finding it after: guard T, whose stated proposition is is X′ a typical lead rendering at all, is implemented as is X′ more like X than an outsider's rendering is, which catches only extreme atypicality. Accepted and stated rather than repaired, because the confidence interval the remedy needs did not exist — and the unregistered typicality check that was added instead is exactly what then complicated the result |
| (bgy) | The rule an audit validates can be the rule the audit's own repair left broken, and only a ground truth can tell you which. ARM-figure-audit spent three sessions establishing that published figures stand under the repaired name_tokens, and every one of those verdicts is a statement about stability. E-20260801d put 138 ever-capitalised types to three seats as a question of fact and got recall 1.000, precision 0.830 type-level and 0.679 token-weighted: one excluded token in three is not a name, the error is entirely one-sided over-exclusion, and 11 of the 19 false positives come from R1 — the rule S031's defect A repaired — against 1 from R2, which is the opposite of what RS-20260726c reported for its own seven cells. The mechanism is a symmetric blind spot: ends_sentence() strips closing delimiters from the previous token and never looks at an opening delimiter on the current one, so every quotation opened after a comma yields a false name. Rule: an audit that recomputes a figure under a rule has tested the figure, not the rule; when the rule is a heuristic, one ground-truth cell is worth three sessions of recomputation, and it is cheap because the candidate set is finite and enumerable. Corollary, from amendment A3: report a type-level precision and a token-weighted one — the gap was 15 points here, and the type figure alone understates the effect on n-grams by half. | S080 | S080 |
| (bgz) | The preserved-bytes guard is defeated by a label collision, and only the ledger can see it. Note (bdt) writes a response's raw bytes to disk before any parse, which is what makes a rejected body inspectable. call.py labels the primary slug's first attempt with the bare tag, so a re-dispatch of the same seat — amendment A6's cap raise on P2 — wrote over the failed body's .raw. Nothing in the runner, the analysis or the verifier noticed: the accepted body is present and correct and every check passes. It surfaced only in the spend reconciliation, as a key-usage residual of $0.064228500 against an overwritten body whose usage.cost was $0.0642285 — exact to the cent, and sayable at all only because the .meta.json had been read before it was clobbered too. Rule: a preservation guard must be keyed on something monotonic, never on a tag that a retry can reuse — and, more generally, the cost ledger is a second, independent detector of missing artifacts and should be reconciled before the run record is called complete. Sibling of (bgu): both are guards that pass while guarding nothing. AMENDED 2026-08-02 (S087) — the S080 repair does not cover the case that actually recurred. S080 labelled every attempt uniquely and fixed the collision between a primary and its reserve inside one invocation. call.py numbers attempts from 1 per process, so a stage re-invoked after an amendment reuses ..._try1. S087 raised a max_tokens cap twice and re-ran two stages, overwriting three billed bodies worth $0.211820146. Nothing detected it: the verifier's cost check sums the bodies that survive, so it reconciles perfectly against a smaller truth. The rule is now: a label must contain something a re-invocation cannot reproduce — a stage-invocation counter persisted to disk, or the declared cap — not merely the attempt number; and the ledger figure must be the sum over bodies plus anything a re-invocation destroyed, stated as two numbers. | S080 | S080 · S087 |
| (bha) | Running another session's analysis script in its own directory is a WRITE. The ARM-figure-audit sweep tried to recompute three pre-repair experiments by invoking their analyse.py in place. Two failed on missing arguments and materials; the third succeeded and rewrote E-20260726-period-control/runs/cell.json, a stored output of a session five sessions earlier. It was caught by git status and restored, and the diff turned out to be the documented S031 name-set change, so nothing was lost — which is luck, not method. Rule: an audit reads; it does not execute the thing it audits in the thing's own directory. Reimplement the specific figure from the stored materials in the auditing experiment's own tree, or run the script against a copy. And: git status is part of a verification pass, not part of committing. ⚠ Id collision: (bha) also names the S081 front-matter/senses: note; resolve a citation by its subject (S084). | S080 | S080 |
| (bhk) | A forced-choice null control between two identical texts manufactures the slot bias it is then used to police, and its floor measures identity detection rather than resolution. E-20260802c presented three null items — one Japanese, one Japanese held out, one Russian — as the same text in both slots, and required a non-tied preference. All eighteen cells recognised the identity and said so ("both translations are identical, so I arbitrarily prefer A"), scored the two slots the same, and produced a floor of exactly 0.0000 on all six senses. Two consequences, and both bit. (i) A floor of zero is not a resolution measurement: a juror that has recognised two texts as the same has told you nothing about the smallest difference it could register between texts that genuinely differ, and every "above the floor" statement built on it is vacuous — the design's own registered prediction that accuracy would sit below the floor became passable only on an exact zero. (ii) Told to choose anyway, jurors defaulted to slot A (P2 5 of 6, P5 6 of 6), which pushed their registered slot-A rates from 0.643 and 0.714 up to 0.750 and 0.800, outside the pre-registered band, so the null control triggered the exclusion of two jurors' preference data on the strength of cells where preference is arbitrary by construction. Rules: a null control must be a pair the panel cannot separate while believing they differ — a meaning-preserving micro-paraphrase with a frozen edit list, not byte-identical text; and a forced-preference instrument must exempt cells whose two arms are identical, or exclude them from every slot-preference statistic computed from it. RS-20260802c-regime-scoring §§3, 5. | S089 | S089 | · DISCHARGED S094 — E-20260803-a4-set built the meaning-preserving micro-paraphrase null this note demanded; it moved the jury by 0.139 against a 0.233 retest floor, and it is not co-presented, so recognition cannot occur.
| (big) | The statement a rendering is scored against is an arm of the experiment, and if the author of the rendering also wrote it the comparison is closed. E-20260802e established R1 by asking three seats whether an English rendering conveys a stated relation — and the lead wrote both the rendering and the statement. Two independent pre-run critic passes at S101 raised this as their first BLOCKING finding, in the same words: the grading question then measures whether the grader can detect the clause the author embedded to match the relation the author wrote. Repaired in place: a seat shown the source, a gloss and the construction and no English of any kind authored all eight statements, and the lead's were discarded. The result reversed. Against an independently authored yardstick, the filed close translation scored 0.810, an uninvolved seat's plain translation 0.810, the forced re-rendering 0.738 — and a negative control stating a different relation sat at 0.071, so the readers were discriminating and the flatness is not instrument failure. Rule: whenever a design scores prose against a written statement of what the source conveys, that statement is authored by a party that has seen no candidate rendering — and any arm frozen against the old statement (a positive control, a gloss) must be re-authored from the new one, which S101 failed to do and paid for through its own F1. Sibling of (bhz): the leak is not always in the text shown, sometimes it is in the question asked. | S101 | S101 |
| (bis) | A positive control can convict a repair the pre-run critic asked for, and this one did. E-20260805's critic (NEEDS-REDESIGN, three BLOCKING) found, correctly, that coding "does the specific sense survive anywhere in the poem" scores kept the word but moved it off the rhyme as a full save — when that displacement is the cost the hypothesis is about. The accepted repair added a fourth label, ELSEWHERE. ELSEWHERE became an absorbing state: 13 of 14 canary positions, 59 of 82 Taylor positions, 38 of 80 Egan positions — and it swallowed two of the four seeded generalisations, firing FC4 at 2 of 4 and withholding every registered prediction in the run. Two rules. (i) Adding a label to a forced-choice code is a change to the instrument and needs its own control run, not the old one — the canary was designed against the three-way code and was the only thing that noticed. (ii) A critic's finding can be right about the measure and its fix wrong about the instrument; accepting a BLOCKING finding does not license accepting its prescribed change unexamined, and a design that accepts one should say which of the two it is accepting. Sibling of (bhr), from the other direction: there a control computed before dispatch saved the run; here a control computed after dispatch convicted the repair. RS-20260805-rhyme-or-flavour §4. | S109 | S109 |
| (biu) | A positive control must be registered as a DIFFERENCE against the null arm. An absolute threshold the null arm also meets is not a control, and it will pass while demonstrating nothing. E-20260805b froze FC2 as "if fewer than 2 of 3 seats in arm 2 separate the pairs, H1 is withheld" — an arm-2-only threshold. Arm 2 separated at 3 of 3 and FC2 passed; arm 1 also separated at 3 of 3, so the loud inserted marker ("Mr Holpainen") moved no letter and no ranking, and the run ended with no demonstrated ceiling and therefore no licence to say what the target language could carry. The criterion could not have failed for the reason it existed, because it never looked at the arm it was controlling for. Rule: write every control criterion as a comparison between arms — arm2 > arm1 by a stated margin — and never as a threshold one arm can satisfy alone. Companion to (bis): there a critic's prescribed repair broke the instrument and a control caught it; here the control was written so that nothing could catch it. RS-20260805b-address-or-action §7.1. | S110 | S110 |
| (biv) | A dispatch guard rejects a FORMAT, never a CONTENT. Read every billed body it rejects, before proceeding. E-20260805b's pre-run critic wrote **VERDICT: …** with markdown emphasis; the runner's line_re required a bare line start, counted zero answer lines, declared a seat failure and re-dispatched. Both bodies were valid, full adversarial reviews and both were billed — $0.0418712 and $0.0341885. The first was read straight off disk and its six findings became the run's amendments; the second was never opened until after the nine seats had run, and it contained two findings the first did not, one of which (the scene-content confound) is the reason the run's headline is a null. So the cost of the unread body was not its $0.034 but a design change that arrived after dispatch instead of before it. Rule: when a guard rejects a body, open the body. A guard's verdict is about whether the runner can parse the answer, and this project's guards have now twice (with (bgw)) rejected work that was correct. RS-20260805b-address-or-action §7.2. | S110 | S110 |
| (biw) | A leak screen that checks TERMINOLOGY does not check QUOTATION, and quotation is the same leak. E-20260805c inherited E-20260804g's banned-lexicon screen, which fired on 0 of 10 yardstick statements. Five of the ten nevertheless quoted something out of the material: Polish tokens (sprzedano, takich państwa), an English word choice (strapping), and at one site the string "your husband" — the unmarked English rendering of that very site, verbatim. A grader given that statement is given the answer, and the frozen screen could not see it. Rule: any screen protecting a blind judgement must check for quotation as well as for vocabulary — no source token, no quoted string, no naming of a target-language word the arms might use. And when the fix is found after reading the output, make the trigger mechanical and the response uniform: E-20260805c re-requested all ten statements on any single hit, so no statement could be selected for re-request after being read, and it declared that the lead had seen the first set. Companion to (biu): there a control could not fire; here a control fired correctly on the wrong property. RS-20260805c-no-loss-to-repair §9. | S111 | S111 |
| (bix) | If every arm of a recovery design preserves the passage's content, the design measures CONSISTENCY WITH the relation, not whether any marking carries it. E-20260805c graded seven arms against an independent statement of what a Polish passage conveys. WRONG — the one arm that asserts different content — was recovered in 4 of 48 judgements, so the instrument plainly discriminates. But PARA, a free paraphrase written by a seat that was told nothing about any relation, any study, or anything to preserve, was recovered at 8 of 8 sites, 0.854 per judgement — level with the filed translation and above both R1 re-renderings. The six content-preserving arms sat between 0.667 and 1.000 and the content-changing arm sat at 0.083. Rule: before reading a recovery rate as evidence that a marking was carried, the design must contain an arm that varies the marking with the content held fixed. Without it, a high rate is a fact about the sentence's subject matter. This is the standing explanation now available for E-20260804's FR→EN null and for E-20260805c's empty gate alike, and it bounds what RS-20260802e's 5-of-6 can be read to have shown. RS-20260805c-no-loss-to-repair §4. | S111 | S111 |
| (biy) | An unbriefed arm is not thereby a NULL arm: ordinary machine rewriting has directions of its own, and one of them is towards the knowing ironist. Note (bix) prescribed an arm that varies the thing under test with the content held fixed, and E-20260805d built one — a persona shift, hand-written, against a paraphrase from a seat told only "say this a different way". Signed to the five declared directions the shift was aimed at, the deliberate change of narratorial person moved source-blind readers 0.800 and the aimless rewrite moved them 0.800: identical to three decimals, ratio 1.000 against a withholding bar of 0.75. The paraphrase raised the narrator's rated self-awareness from 2, 2, 7 to 6, 6, 7 without being asked to touch it. So (bix)'s remedy does not by itself rescue a design: an unbriefed control tests whether the measure moves for free, and when it does the primary is withheld, but the control is not a floor and cannot be read as one. Before registering an unbriefed arm, say which direction you expect its default to lie in and whether that direction is the treatment's. RS-20260805d-two-persons §3. Companion to (bix), which it extends rather than replaces. | S112 | S112 |
| (bja) | A seat probe bounds a seat's token appetite on the PAYLOAD SHAPE IT PROBED and on no other. Note (bit) bought the probe and E-20260805d ran it: qwen/qwen3.7-max passed at max_tokens 5,000 on the rating payload — one 600-word text, eight integers out — and then returned finish_reason: length with ~30,000 reasoning characters twice on the parity payload at a cap of 8,000, $0.0757 for zero content, 25% of the run's spend. The two payloads differ in kind: rating is judge one text on a fixed scale, parity is enumerate the differences between two texts, and enumeration has no natural stopping point. Probe the shape whose failure would cost the most, not the shape you happen to dispatch first, and where a run has two task shapes, probe both or price the unprobed one at its cap. Note (bhf)'s eighth firing and its sixth distinct slug. RS-20260805d-two-persons §10. | S112 | S113 |
| (bjb) | When a design delegates a judgement to a seat, the seat's own count governs the criterion — including when the lead would have counted differently. E-20260805d's F3 withholds the primary if the substantive divergences between two renderings exceed three, and it delegated substantive to the screen seats. One returned NONE; one returned COUNT: 4. Had the lead adjudicated one of the four as a wording difference the count would have been 3 and the criterion would not have fired. The rule decided and the margin was printed, and the alternative — the author of both texts ruling on how many ways they differ — is the thing a delegated criterion exists to prevent. Rule: a delegated judgement is not reopened by the delegator after the answer arrives; if the delegation was wrong, that is a finding about the design and is recorded as one. | S112 | S112 |
| (bjm) | A signed scale is not a two-sided scale until you have counted how often its negative half is used — and the count belongs in the design, as a gate, not in the discussion. E-20260806b asked three seats to score five renderings against their Italian on three signed −3…+3 scales, positive = the deformation Berman predicts. The headline came out clean and monotone — Strettell +1.833 on vernacular effacement, Dole +1.500, the 2026 arms lower, the resistancy control −0.167. Then the registered gate F1′ counted the negatives: 4 of 186 codes on explicitness (2.1%) and 9 of 186 on popular speech (4.8%), against 29 of 186 (15.6%) on register. A mean of +1.8 on a scale nobody scores below zero is compatible with a one-sided instrument, so two of the three primaries were withheld and the two best-looking numbers in the run went with them. The one-sidedness is not noise and cannot be argued away as sloppiness: the same seats returned 0.000 test–retest on a duplicated arm, 18 of 18, and 77 of 78 extreme codes quote something actually present in the span they refer to. They are precise, reproducible and textually accurate, and on two scales of three they never looked below zero. Rule: any design using a signed scale registers, before dispatch, the minimum rate of negative-half usage below which directional claims on that scale are withheld — and it does not substitute a control arm's sign for that count, because a control that fails to go negative is ambiguous between a broken scale and a true universal. That ambiguity is what the pre-run critic caught in this run's original F1, which would have withheld the primary exactly when the hypothesis was supported. RS-20260806b-berman-occurrence §4, §7. | S118 | S118 |
| (bjo) | A materials diagnostic must be computed the way the analysis will consume the material, or it will be true of a corpus the run never analyses. E-20260805f chose narration-only as its primary analysis on the ground that "dialogue share is higher in the translation in all four cells", and published the sentence in the design and in RS-20260805f §2. Its analysis blocks one work at a time, choosing that work's dominant quotation convention per work; the diagnostic was computed over an arm's works concatenated, with one convention chosen for the pair. Machen's original arm is two books with different printers' conventions, so the concatenated figure came out 13.7 where the per-work figure is 37.5 — and the cell's asymmetry runs the opposite way to the published claim (0.290 translated against 0.375 original). The statistics were unaffected, because the choice the diagnostic justified is defensible either way; what was wrong was a sentence, published twice, about the corpus. Rule: a diagnostic that justifies a design choice is computed by the same code path that will later read the material — same unit, same normalisation, same per-file decisions — and where it cannot be, the divergence is stated in the design rather than discovered by the next run. RS-20260806c-source-grammar §9; E-20260806c design §2.4 and amendment A7. | S119 | S119 |
| (bkf) | An "unlicensed strangeness" arm must be checked edit by edit against the source's own morphology and syntax AT THAT SITE, and a preposition swap almost never survives the check. E-20260807f built a factor out of eighteen markedness edits declared to answer to nothing in the French, and the pre-run critic found all six of one kind — odd prepositional government — traceable to the French preposition at the site: saturated of ← saturé de, upon them ← sur eux, unto ← à twice. A third of the factor was source-licensed, which is the one thing it may not be, and the design's flat sentence "none is a calque" was false. The kinds that did survive are the ones the source language cannot form — Germanic noun-compounding and English predicate fronting, neither available in French — and the replacement kind, broken / unresolved construction, survives because the passage completes every construction it starts. Rule: for each edit in an unlicensed-markedness arm, name the source feature at that site and say why it cannot motivate the edit; prefer kinds the source language structurally lacks; treat function-word substitutions as licensed until shown otherwise. Caught before dispatch and cost $0.03. RS-20260807f §1; E-20260807f amendment A3. | S130 | S130 |
| (bkg) | Write down an exact test's minimum attainable P before the run, because a permutation over the wrong units cannot reach significance no matter what the data say. E-20260807f §9 registered an exact permutation "relabelling the items' factor assignment"; the first implementation relabelled the four pooled cell means instead, which has C(4,2) = 15 arrangements and a floor of P = 0.0667 — so every effect in the run, including several that were unanimous across 3 of 3 seats and 6 of 6 segments, came back at P = 0.3333 or 0.6667 and would have been reported as null. Recomputed on the six segments as paired units (2⁶ = 64 sign patterns, floor P = 0.03125) the same effects sit at the floor. Rule: state the exact test's permutation units and its floor in the frozen design, and check the floor is below the alpha the design intends to use. The general form: a test whose smallest possible P exceeds your threshold is not a conservative test, it is a broken one. Caught before any number was reported. RS-20260807f §1; analyse.py docstring. | S130 | S130 |
| (bkh) | The empty-body-with-all-tokens-on-hidden-reasoning failure can be PROMPT-shaped, not seat-shaped, and when it is, neither standing remedy works. Note (bhf) has fired eighteen times and its two remedies are raise the cap (note (bhq)) and change the seat. E-20260807g dispatched three sibling generation prompts in one batch to one model; IND came back with zero content and finish_reason: length while INDC and DIRF completed cleanly at the same cap — then did exactly the same under a raised cap of 30,000, then did exactly the same under a different model at $0.212442 for nothing, whose siblings again completed. Three failures, two models, one prompt. The distinguishing feature of the failing prompt is that it asked for a mechanical transformation under a conjunction of prohibitions (keep every word grammar allows, change the speech act not at all, add nothing), which is a specification a model can deliberate about without bound. Rule: when a batch of sibling prompts has one that fails (bhf) and the others succeed, stop applying (bhq) and change the PROMPT — loosen or split the conjunction of constraints — or accept a cheaper hand. Cost of learning this the other way: $0.319654560, 39% of the session. RS-20260807g §8; config/budget.md S131 waste row. | S131 | S131 |
| (bki) | A between-hand divergence or agreement rate measured on ONE rendering per hand is not a stable quantity, and the instability is larger than the effects such designs go looking for. E-20260808a took two draws from each of two models at temperature 1.0 — the only reason it knows — and the same ten loci, judged by the same three seats, came back 0.900 different in one draw and 0.200 in the other. No class difference anywhere in that run was a third of that. The mechanism is visible in the prose and needs no statistics: an ordinary locus is a one-word synonym choice (earth against soil, hardship against misery) that resamples. Rule: a design whose statistic is how often two hands part must take at least two draws per hand, or state on the result page that its number is a single draw and that its spread is unmeasured. This bites backwards as well as forwards — RS-20260807c's divergence rates and RS-20260806d's 234 sites are single-draw numbers, and ES-20260806 §3.2 now says so. | S132 | S132 |
| (bks) | A slug with no failure record in this project is probed on the SMALLEST call of a stage before the batch, and the fallback is named in the frozen design. E-20260808f swapped its gate checker to qwen/qwen3.8-max on a critic's BLOCKING finding — a lab new to the project, so note (bkh-corr)'s check the project's own failure record had nothing to check. The design therefore registered the probe and the fallback in advance: dispatch the 27-item RU call first, and on a dead body take a named model with a record and restore the limitation the swap had deleted. The probe fired on the first call: finish_reason: length, 14,000 reasoning tokens and zero visible content, $0.087184 — for one call rather than five, and the fallback was taken without deliberation because it had been written down. The rule generalises past new slugs: in the same run, a model with a good record in this project (nvidia/nemotron-3-ultra-550b-a55b) died the same way on the Japanese call at a 20,000 cap while returning 59 Italian items on the same cap in 5,322 characters — so the source language, not the batch size, drove the reasoning chain; splitting the call in two produced both halves for $0.0334. Rule: probe-smallest-first, name the fallback in the frozen design, and treat a per-language cap as unproven until that language has returned once. RS-20260808f-placeless §7. | S137 | S137 |
| (bkh-corr) | [FIRED S136 NOT HONOURED, $0.305211000; HONOURED S137 — E-20260808f's frozen design named moonshotai/kimi-k3 as not-to-be-dispatched before any call went out, and it was not. See note (bks) for what replaced the missing-record case.: moonshotai/kimi-k3 was chosen as E-20260808e's second generation hand for MODEL-FAMILY diversity, and returned finish_reason: length with zero content twice — the exact failure this note records for this exact slug, three sessions earlier. The note's remedy was never reached because its precondition, check the project's own failure record for a slug before dispatching it, was not written down. It is now: a slug this project has recorded returning empty length bodies on generation prompts is not selected for a generation stage, whatever else recommends it.] Note (bkh)'s the failure is prompt-shaped, not seat-shaped does not hold, and the cheap remedy is the first one. At E-20260808b stage 2, moonshotai/kimi-k3 returned zero content at finish_reason: length on three of four generation bodies — $0.365763 for nothing — while the fourth body, same model, same prompt shape, same batch, came back clean; so the same prompt shape both succeeded and failed on one seat. Changing the seat to mistralai/mistral-medium-3-5 produced all four arms cleanly for $0.042, an eighth of what the failures cost. Rule: when a generation prompt (not a rating block) burns its allowance on hidden reasoning, change to a cheap non-reasoning seat before raising the cap a second time, and regenerate every arm on the new seat so the within-hand contrast is not split across two authors — the one clean body from the failing seat is discarded deliberately, and the discard is recorded. | S133 | S133 |
| (bkj) | Zero content characters with finish_reason: stop is NOT note (bhf), and neither of (bhf)'s remedies is the first move. (bhf) is a length body: the seat spent its whole allowance on hidden reasoning, and the remedies are (bhq)'s raised cap or a changed seat. E-20260808a produced a third shape — x-ai/grok-4.5 returned zero characters with stop on the primary block, $0.023064 for nothing, while the same seat completed three other blocks of the same shape in the same dispatch. A straight re-dispatch of the identical prompt to the identical seat returned 8,334 characters. Rule: on a zero-content stop, re-dispatch once unchanged before touching the cap or the seat, and preserve the dead body as always. Changing the seat here would have spent a second model's money to fix nothing and would have put one block on a different hand from its siblings. | S132 | S132** |
| (bkk) | A judgement that will define an admission condition must be shown to reproduce on a second text in a second language BEFORE the condition is written into a release. RS-20260807d measured the concordant/discordant call at Fleiss κ = 0.786 on one Ōgai courtroom scene, and ARM-marking-work step 2 was scoped to write it into a v0.2 as prediction 1′'s admission term. RS-20260808b put the identical three-question instrument to three readers of two whole Chekhov stories chosen because they are made of the phenomenon, and it came back at κ = 0.0584 — six utterances named between three readers, none named by all three, one reader finding it nowhere — while the markedness half of the same call from the same seats reproduced at 0.7214. The cause is not noise: discordance lives in a text's trajectory and the instrument asks about an utterance, and S128's κ was earned on a scene that restates the standing in every line. Rule: a κ from one text licenses this text; a condition a release will quantify over needs a second text, in a second language, before the wording is drafted. Corollary: check whether the high-agreement measurement came from a text where the property is locally restated — that is the shape that generalises worst. | S133 | S133 |
| (bkl) | A rating scale that asks how a speaker presents his own standing will be answered, on a polite isolated sentence, as though politeness meant lowness — and the error lands on the SOURCE side, where there is no other cue. E-20260808b's two largest transfer discrepancies were both the general in «Смерть чиновника»: Russian seats scored «Вам что угодно?» and «Ах, полноте…» at −0.33 (naming «вы»/«вам» as the marking form) and English seats scored the same utterances at +2.00 and +1.67 — and the English seats were right. вы marks distance and respect toward the addressee; it does not claim lowness for the speaker. Neither the design nor the pre-run critic caught it, and it inflated every all-marked mean in the run (0.368 → ≈0.20 with the two sites removed). Rule: any standing scale must separate deference shown to the addressee from rank claimed by the speaker, as two questions, or it cannot be used on a symmetric-polite pronoun system. | S133 | S133 |
| (blm) | Showing a seat every arm of a passage side by side turns a rating task into spot-the-odd-one-out, and one arm per call costs about the same. Every jury run in this project from S020 to S146 presented arms together, and E-20260809i's pre-run critic named the consequence: where one arm is built to be visibly anomalous, a seat can infer the experimenter's contrast from the comparison set and mark the odd member without ever consulting the rubric, so the run measures discrimination in a simultaneous comparison rather than how a translation scores when read as a translation. Reverse-ordering does not touch it — it presents the same contrast backwards. The run was rebuilt with one arm per call, no seat told that another rendering existed: 128 small calls for $0.740, against 54 large ones, because the per-call instructions and source repeat but the arms do not. Rule: where an arm is constructed to be conspicuous, score it alone; where arms are presented together, the run says so and reports the finding as a discrimination result. Companion to the halo rule (wiki/goodness-senses.md usage rule 6), which governs several senses in one call; this one governs several arms. RS-20260809i-terminology-drift §1, §6. | S147 | S147 |
| (bln) | Choosing an obscure text does not buy an unrecognising panel, and the successor design that needs one must measure recognition as an admission gate rather than assume it. RS-20260809d §5b recorded 11 of 12 bodies naming "The Cask of Amontillado" through an anonymisation, and set its successor the requirement of a text the panel does not recognise. E-20260810-source-asymmetry went as far the other way as this project can reach — Bret Harte, "Brown of Calaveras", an 1870 magazine story by an author out of print in every language he was translated into, every personal and place name substituted, 794 words of dialogue and no narration — and 5 of 5 seats named Harte; 2 of 5 named the story, one at confidence high. Rule: a design whose validity turns on the panel not knowing the text dispatches a recognition probe FIRST, as a gate that can stop the run, and does not treat obscurity as evidence of anything. This is the third measured failure of the obscurity premise, after S065 and S079 measured it for translator independence — same premise, different consequence: obscurity is not protective against recall for contamination, and it is not protective against recognition for judging either. Companion to the closed measurement programme on contamination (CLAUDE.md), which this does not reopen: this note is about designs, not about the rule on lead independence. RS-20260810-source-asymmetry §4. | S148 | S148 |
| (blo) | A design may register what a null WILL MEAN, but it may not register a destructive consequence for one — non-rejection is a failure to found, never a refutation, and a p-value above α may not delete work. E-20260810-legend-lexis's frozen decision table, and the translator's frozen log D47 that had already promised to honour it, both said: if the study limb finds no Károli phrasing, V12 is struck and erratum 1 flattens two spans of translation. The independent pre-run critic (BLOCKING 9) refused it before the run: non-rejection at α = 0.05 shows the effect was not found, not that it is absent; the design's "power floors" were sample-size thresholds with no registered minimum effect and no equivalence margin behind them; and a population-level association cannot license a change to five particular paragraphs in any case. The primary then did come back null (+0.0087, p = 0.4392) and the erratum did not fire. Rule: a pre-registered consequence is itself a claim and gets critiqued like one; where the consequence of a null is destructive, register instead (i) what estimate and interval would license the act, (ii) an equivalence margin if absence is to be asserted, and (iii) a unit-level analysis if the act is unit-level. The honest consequence of a second failed founding is that the artifact says so on its face — the register now labels the marking an explicitly unsupported translator's judgment — which is a stronger record than a deletion would have been. RS-20260810b-legend-lexis §4, §9; E-20260810-legend-lexis §11 A9. | S149 | S149 |
| (blr) | No reader-side question yet found turns perceived-source-carriage into evidence that anything was carried — including asking the reader to point. Three runs had shown the seven-point rating cannot separate carried form from free oddity (RS-20260807f, RS-20260808d, and RS-20260809h's word-scramble above five of six foreignizing hands). The live repair was that a rating is permissive and a quotation is a commitment: ask the reader to quote the German words a marked English span carries, or answer NONE. E-20260810d did that at ten sites of one Arnim paragraph, four versions per site byte-identical outside one marked span. The competence gate PASSED at 0.90 — 27 of 30 cells quoted the registered locus where something really was carried — and the same seats attributed source-carriage to a word-scramble in 0.60 of cells and to ordinary faithful English in 0.4333, with licensed-minus-free at +0.200, exact P = 0.156 and the licensed arm 1.37 points stranger on the run's own scale. Rule: a design that needs to know whether a translation carries a source feature measures it on the two texts; a reader-side judgment is evidence about the reader, and no reformulation of the question has repaired that. Corollary for uptake and execution measures: RS-20260809h's bar — never without a matched oddity control in the same run — stands, and the pointing task does not satisfy it. What is NOT licensed: this is a failure to demonstrate discrimination, not a demonstration that licence is invisible; no equivalence claim is registered and none may be made. RS-20260810d-carriage-naming §3, §5. | S151 | S151 |
| (bls) | A body whose finish_reason is length did not finish, and a parse rule that can read its first line will hide that. This project's dead-body practice (notes (bhf), (bhq), (bkw)) is written for the zero-content length case, and E-20260810d's registered F5 inherited it: MALFORMED or missing bodies are retried once. Eighteen bodies came back length; ten were unreadable and caught, and seven were readable — truncated part-way, with an intact first line that the design's own parse rule accepts as a verdict. They passed every check silently. Re-dispatching them at a raised cap changed six of eight verdicts, all six in the same direction, so the truncation was not cosmetic. Rule: finish_reason is checked before the parser runs, and a length body is dead whatever it contains; a run's cell-return criterion names the finish reason, not the parseability. Corollary, because the repair was decided after the direction was visible: where a protocol deviation is discovered post hoc, compute both datasets and report both — here every gate verdict and every primary reading was identical across them, which is the only thing that made the deviation harmless. RS-20260810d-carriage-naming §6.5. | S151 | S151 |
| (bmh) | A binary same/different verdict is not evidence until its reason has been read, because paraphrase drift elsewhere in the passage counts as signal. E-20260811g measured whether a one-word source edit reaches an independent English hand, with the arbiter returning {verdict, what_differs}. Of 24 bound-morpheme cells, exactly 2 came back DIFFERENT — and reading what_differs showed neither was the manipulation: one seat had seen ghost against image, another the horses against a horse, both ordinary drift at other points in the same window. Counted, the bound arm arrived at 2 of 24; read, it arrived at 0 of 24, and a 0.083 arrival rate would have been published as a small real effect. The rule: any design whose statistic is a binary difference verdict collects a free-text reason in the same call and codes it, and the result page reports arrival-by-reason beside arrival-by-count. A verdict without its reason may be dispatched but may not be a primary. Sibling of (bkt): a control that cannot be read is not a control. RS-20260811g-carrier §4. | S161 | S161 |
Sources, texts and provenance
| id | note | first seen | last fired |
|---|---|---|---|
| (bno) | Overlap with published translations is a property of the REGISTER a policy targets, not of having a policy — so a contamination measurement made for one regime does not transfer to a sibling regime of the same work. RS-20260813c §3.1 established the first half: one translator, one day, one copy-text, 19-token run and 8 shared twelve-grams with a published hand under no register rule, and 0 twelve-grams with any hand under a high one. E-20260814d supplies the limiting case from the other side. A light register rule — one step above the source, no expansion — walks straight back into the published neighbourhood: R26~BRAK 20 shared twelve-grams against R06's 5, R26~PAULL 5 against 1. And against the lead's own prior renderings of the same tale the asymmetry is extreme: 113 shared twelve-grams and a 27-token run against the unruled pass, 0 against the ennobled pass, from a translator who had read neither. Fires at: every design that reuses a contamination verdict across arms of one work, and every declaration of contamination: on a new regime's artifact. Remedy: measure per regime, and read the table as a ranking of register distance rather than of independence; a clean pair may be clean because the two texts are pitched differently, not because their translators were. Note (bhb)'s densest firing at this text length. RS-20260814d §3, T-flipperne-R26-v1 §Contamination. |
S183 | S183 |
| (bnd) | A blindness claim must be checked against the project's own pages, because result pages quote the comparator — and the pages a design must read to write its predictions are exactly the pages that quote it. E-20260813d declared span A's re-rendering blind to Mukherjea 1914 and the declaration was true of the book: it was never fetched before the freeze. It was false of the repository. RS-20260811f §6 is a table of that translator's English at six units and its closing paragraph quotes two more — eight of the span's 86 units — and those were the very sections the design had to read to register two of its predictions. A ninth unit carried a panel hand's English. The leak was found before a word was translated, recorded as a dated amendment with the unit list, and the quoted units removed from the measured cells; three of the leaked strings then showed in the new rendering where the old one had them differently — lentils, and then, by the roadside — none of them reachable by any n-gram count, all three found by eye against the declared list. Fires at: every unit that claims a rendering is blind to a comparator the project has already written about. Remedy: before translating, grep the project for the comparator's name and enumerate every quoted string on the target material; the enumeration is the artifact's contamination: basis, and none is not available once the list is non-empty. The declared boundary is what makes absorbed wording findable at all — an undeclared exposure of this size would have left the same three strings and no way to see them. RS-20260813d-first-span-again-dakghar §3, E-20260813d §10. |
S175 | S175 |
| (bii) | A clean verdict from tools/dependence_check.py does not rule out derivation — it rules out verbatim derivation, and the two came apart on the first pair this project has tested against an external ground truth. Ormsby 1885 says Motteux's Don Quixote is "a concoction from Shelton"; Fitzmaurice-Kelly 1905, a scholar with no rival version to sell, independently says it "is based on Shelton's rendering, and checked by constant comparison with the French translation of Filleau de Saint-Martin". On I.3 the tool returns clean for that pair — 7 name-excluded 7-grams, 0 twelve-grams, longest run 11 — while returning DEPENDENT? for three pairs with no such documentation. A derivation that proceeds by rewriting leaves no run for an n-gram instrument to find. Rule: where a design's validity turns on two texts being independent, a clean verdict is necessary and not sufficient, and any external record of the pair's history outranks it. Corollary, on the same run: the shared-7-gram count is as poor a proxy as the run length when names are not excluded — raw counts made lead~ormsby the top pair at 91, and name-excluding inverted the ordering to put lead~motteux first at 14. Report both columns or report neither. |
S102 | S102 |
| (bic) | A translator's own log of what it changed is not a record of what it changed, and the gap is invisible from inside. R14's operator was declared propositionally conservative and logged at 37 sites; a pre-run critic, reading only the two frozen texts, found it deleted content at five of them — "like an owl" gone with no replacement and absent from the log altogether, howled and moaned → was blowing loudly, drummed → was falling on, splashed → were moving, past counting, and hot → many times and warmly. Each site felt like the operator it was filed under. Note (bew) says re-derive a scheme's clearest cases from its definition; this is the complementary failure — the log is a record of intentions, and a mechanical diff of the artifacts is a record of acts. Rule: when a design's validity rests on what an operator did NOT change, compute the difference from the two texts (a coverage diff, a word-level absence list) before dispatch; do not read it off the log. Corollary on definitions: F4's wording, "replaced by a generic verb of the same denotation", was falsified by its own instances, and the amendment — state the manner explicitly instead of deleting it — is what the operator had always meant. |
S099 | S099 |
| (bia) | A published corpus frequency counted by substring is not a frequency, and it inflates — silently, plausibly, and in the direction that flatters the claim. A-english-tale-register is built on a mechanical attestation test (a feature is in the repertoire iff it occurs in the 51,154-word Jacobs corpus), and its first draft published said he at 31 / 21 and says the at 25 / 0. Both were str.count() figures: said he matches inside said her and said heavily. The word-boundary counts are 20 / 2 and 23 / 0 — the Kipling figure was wrong by an order of magnitude, and 21 → 2 is exactly the number a reader would quote as evidence that the feature is not one collector's habit. Nothing downstream flagged it: the figures were plausible, the class assignments were unaffected (all still ≥ 3), and the item set had already been frozen. It was caught by analysis/verify.py re-deriving every published count from the stored corpora — a check written because the anchor's whole method is a count. Rule: a frequency published from a corpus is computed with an explicit word-boundary matcher, and the verifier recomputes every figure the page prints, not a sample. Corollary, and it is the general one: when a page's method IS a measurement, the verifier must re-measure the page, not just the experiment. |
S098 | S098 |
| (bhj) | A pole definition written after the comparators were read will quote the comparators, and the independent coding it feeds stops being independent. E-20260802b gave two non-Anthropic seats a 21-site book defining an F pole and a D pole per site, precisely so the seats would code without seeing the design. The lead's first draft illustrated those poles with well-corsleted, bright-arm'd, the fat of the land, no inglorious men, death's inexorable doom — every one a phrase from one of the six renderings the seats were about to code, at eleven of twenty-one sites. The seats would have been matching strings, and the agreement rate would have measured the leak. It was caught by re-reading the assembled prompt before dispatch and by nothing else: no guard, no verifier and no critic finding covers it, because the prompt was well-formed and the leak is semantic. Rule: when a coding scheme's categories are illustrated by example, every example must come from an artifact frozen before the coded material was opened — here the two lead translations, committed at c3b41ce — and the scheme's own page must say which frozen artifact each example is from. Corollary: assemble the prompt, then read it as the seat will, as a separate step with its own place in the procedure; a payload built by string interpolation is never inspected unless inspecting it is a step. |
S088 | S088 |
| (bhe) | "The gate could not be run" names two different situations and they license different declarations. CLAUDE.md requires contamination measured before selection and says that where no comparator is reachable a session must say so; it does not distinguish (a) a comparator that exists and was not obtained — in copyright, behind a wall, not fetched — from (b) a comparator that does not exist, no rendering of the work into the target language having been made. S084 had (a) in the strong form: «Gudsfreden» was declared suspected UNMEASURED with two measured siblings at 17 and 16 tokens against an 1899 English comparator, so the risk was demonstrated for the neighbours. S085 has (b): three recorded search routes returned no English rendering of «Köyhää kansaa», of its siblings, or of any Canth prose, and contamination: none rests on that. Rule: a none declaration on an unrun gate is admissible only under (b), only with the search routes recorded as tool results on the artifact, and only with the residual channels named — here a Swedish translation that cannot put English strings in the lead's head but could shape construals. Under (a) the declaration is suspected and the reachability failure is the reason. This does not weaken the standing rule that obscurity is not protective: every measurement behind that rule (S065, S079) was against a work that had a translation the lead could have seen. workshop/translations/koyhaa-kansaa/contamination.md. |
S085 | S085 |
| (bfk) | An intralingual dependence_check.py figure may never be quoted beside an interlingual one, and a DEPENDENT? verdict on an intralingual cell carries no dependence claim. The tool returned 58 shared 12-grams and a 25-token run against the source itself on a same-language pair and printed DEPENDENT?; the tool is behaving correctly and the comparison is meaningless, because two modernisations of one text share the text. Discharges the S058 backlog row; no tool change, because T6 work is a gate and nothing is blocked. Original note (bdw). |
S068 | S068 |
| (bep) | A publisher's plain-text edition has an encoding, and guessing it corrupts words silently rather than loudly. Small Beer Press's official plain-text editions of the two S063 anchor sources are Mac Roman, not cp1252; file reported them as "Non-ISO extended-ASCII" and "ISO-8859", both wrong. A cp1252 decode turned Étienne into "ftienne", l'Hôpital into "l'H™pital", café into "cafŽ" and Sacré into "SacrŽ" — four words across four files, every one still a plausible-looking token, which is why nothing downstream would have flagged it. Caught by enumerating the non-ASCII codepoints of the extracted excerpt before freezing, which costs one line and is now the check: a stored excerpt's non-ASCII inventory is inspected, not assumed. |
S063 | S063 |
| (beu) | A predicate about punctuation is a predicate about the transcription. The same Small Beer files render em dashes as a single hyphen, indistinguishable from a compound hyphen; Gutenberg's Mansfield carries real em dashes and Doctorow's official text uses the ASCII -- convention. In E-20260730d a frozen census class "em dashes to set off parenthetical clauses" came out PRESENT in exactly the three passages whose files show a dash, ABSENT or SPLIT in the four that do not — the raters were right about the files and the files are wrong about the prose. The headline was asserted insensitive to it (analysis/sensitivity.py); a reported figure was not. Before freezing any instrument that can ask about punctuation, count the dash, quote and ellipsis forms in every passage, then drop the class or declare the artefact. Companion to (bep): that one is about characters lost in decoding, this one about distinctions lost in transcription. |
S063 | S063 |
| (a) | Transcribe an anchor claim with its hedges. | S015 | S025 |
| (d) | Bibliographic priors are wrong often enough to matter. | S021 | S031 (THE MONK: the prior was «Чернец», the poem is «Монах») |
| (s) | "Published" is not "checked". | S021 | S027 |
| (aa) | When the only free copy is uncorrected OCR, two independent scans must agree before auditing. | S023 | S031 |
| (bb) | Search evidence-first when the evidence genre is rarer than the materials. | S023 | S024 |
| (cc) | Search the scans, not the index over them — and a one-name sweep is a half sweep. | S024 | S024 |
| (dd) | A matcher over attributes cannot see a wrong referent. | S024 | S024 |
| (ee) | The bias of an unbounded pattern runs opposite to the bias of a bounded one. | S024 | S024 |
| (ff) | Do not write the negative section of a search record before the search finishes. | S024 | S024 |
| (gg) | An excerpt from a single OCR scan must mark, per quote, continuous or reconstructed. | S024 | S025 |
| (hh) | Read the geometry, not the linearised OCR. | S025 | S025 |
| (mm) | Metadata that asserts human authorship is not evidence of human authorship — read enough of a comparator to see whether a person wrote it, on a different passage so the freeze survives. | S027 | S029 |
| (pp) | Decode by the charset the file declares. | S028 | S028 |
| (cde) | A title is not an identifier. «Собака (Тургенев)» returns the 1870 story, not the 1878 prose poem. Any cross-source fetch keyed on a title needs a content gate, and the gate is what does the work. | S031 | S031 |
| (bch) | A page-break marker inside a paragraph splits it, and the word total is exactly right either way — so no length check can see the damage. source-it-full.txt was cleaned from paginated Wikisource HTML; one [p. 43] marker fell inside a <p>, and the stored file claimed 196 paragraphs where the source has 195, with 11,474 words both times. R05 §1 defines translation spans by paragraph index into that file, so the index was not reproducible and the defect was invisible to every check the arm had. Found by re-extracting each <p>, stripping page markers within paragraphs rather than at them, and aligning 195 against 196. Sibling of (bcg) — same family, opposite mechanism: (bcg) is markup silently removed, this is markup silently acted on. Where a stored text's structure is load-bearing, verify the structure against the source, not the token count; and one candidate the same check flags may be genuine (a lead-in paragraph ending in a comma, ¶30, is Verga's own setting). |
S039 | S039 |
| (beg) | A cleaner that counts what it deletes is not a cleaner that says what it deleted, and only the second is usable. tools/fetch_aozora.py prints gaiji images: N … left as ※[#…] and stripped: N — so nothing is silent — and it never says which character or where. On 芥川 「煙管」 it deleted the ê of rôle, a French word printed in Latin script in a 1916 Japanese text and part of the passage's texture, leaving rol. Before reading an extracted Aozora text, read the tool's own stripped-gaiji count, then grep the raw HTML for gaiji or ※[# and recover each alt attribute by hand. Sibling of (bcg) and (bch) and the narrowest of the three: the defect is in the report, not in the strip. This note replaces the S051 backlog row "does the gaiji-class defect exist in the project's other fetchers?", retired at S061 because its general form is unbounded T6 audit work and its operative remedy is a per-ingestion check. |
S061 | S061 |
| (bcg) | A stripper that removes markup silently produces a text that is complete and wrong, and the only symptom is a tidy negative finding. tools/fetch_aozora.py discards Aozora annotation markup; Japanese 傍点 (emphasis) lives in <em class="sesame_dot">, so the extracted text held every character and no emphasis. Reading it supported the conclusion that the translator dropped Poe's italics and small capitals entirely — the exact opposite of the truth: he marks six of seven and adds one. Caught by counting the spans in the raw HTML before writing the section. Before reading an extracted text for what a translator did with form, count in the raw source the markup your extractor throws away. |
S038 | S038 |
| (bfd) | Gating unit 1 for contamination does not bound the work, and a clean unit 1 is the least informative unit to gate. CLAUDE.md's standing rule is to measure before selecting; the practice grown around it is translate unit 1, measure, draft the rest. T-may-night-golova-R07-v1 did that: Unit A returned clean at 5 tokens and the chapter's other two units returned 7 and 19 — 19 being the project's second-highest per-unit run on a lead translation. The overlap was entirely in the expository register and Unit A is lyric, so the gate ran on the third of the text least able to produce a run. Gate each unit before drafting the next, which costs nothing but ordering; or state that the figure bounds only the unit it was measured on. |
S067 | S067 · S077 — DEMONSTRATED rather than argued: a unit-1 gate returned clean at 10 tokens with ZERO shared 12-grams and the whole work returned 17 and thirteen |
| (bgv) | A structural mark that is represented as whitespace is destroyed by any join that trims whitespace. workshop/canon/pelsen/fetch.py converted Swedish Wikisource's {{Tomrad}} — the print edition's blank line, i.e. a paragraph break — into \n\n, and the page-boundary join body.rstrip() + "\n" + page ate it wherever the marker fell at the end of a scanned page, which it did at two of three occurrences. Two paragraph breaks were silently deleted and two paragraphs of 272 and 299 words appeared where the print edition has four. Nothing in the tool reported anything: it counts what it strips (note (beg)) and this was not a strip. Caught by reading the paragraph-length distribution. Rule: carry structural marks across a join as a sentinel character, not as the whitespace they will become, and convert only after every join is done — and check a freshly ingested text's paragraph-length distribution before using it, because a tool that reports no error is not a tool that found none. |
S079 | S079 |
| (bif) | Three freely reachable copies of a public-domain text were one transcription of one printing, and the count of sources that carry a reading is not evidence that the reading is attested. Project Gutenberg 13976, Projekti Lönnrot 99 and Finnish Wikisource all reproduce «Köyhää kansaa» — and all three reproduce the same corruptions character for character, because W2's own front matter names the same transcriber and the same Otava 1917 volume that W1 carries. A page scan of the 1886 first edition was one search away, is public domain, is free, and reads Ei where all three e-texts read EL. Rule: before a source page or a translation treats an e-text as its text, establish what PRINTING it descends from and whether a different printing is reachable — and prefer an edition published in the author's lifetime to a posthumous reprint. The reprint here had silently standardised the eastern-dialect layer the translation was separately recording as a loss (collation.md §3a), which is the failure mode that makes this more than hygiene: a copy-text that pre-normalises the thing you are recording as a loss makes your record false in your own favour. | S100 | S100 |
| (bjs) | A contamination gate's verdict is a property of the PROBE as much as of the pair, and a short probe in the wrong register understates it — sometimes by enough to reverse the verdict. S121's mandatory gate ran on a 163-word narration span and returned lead~King 8 twelve-grams / run 19 (DEPENDENT?), lead~Oxenford clean, King~Oxenford clean. S122 re-ran it on the 1,421-word dialogue scene the run was actually about: lead~King 33 twelve-grams / run 23, lead~Oxenford 13 twelve-grams / run 19, and King~Oxenford 6 twelve-grams / run 15. Two of the three clean verdicts did not survive, including the one on which RS-20260806e §2 rested the claim that the two published translations are independent of each other — which is what its admission condition needed. The design's own pre-run critic had registered the worry in words (amendment A6(iii), "the dependence verdict covers narration; CLOSE and FORCE are dialogue and may be King-tinged") and nobody measured it. So: the gate's probe must match the material at issue in LENGTH and in REGISTER — dialogue for dialogue — and a clean from a short out-of-register probe is reported with its probe's size and kind attached, never as a bare verdict. This is a rule about how to run the standing gate, not a reopening of the measurement programme CLAUDE.md closed at S082. RS-20260806e §6 limit 4; T-kohlhaas-R20-v2 §Dependence measurement. | S122 | S122 |
| (bsk) | _djvu.txt throws away font information, and where typography IS the signal that is the whole finding. The Internet Archive derivative this project has read since S222 — and the *.xml DjVuXML alongside it — carry word coordinates and no font attributes; _abbyy.gz carries <formatting italic="true"> and exists for every scan the shelf uses. E-20260831's central result — that Eastwick marks Sa'di's Arabic by italic at 11 of 12 unframed loci and Ross at 8 of 10 — is invisible in _djvu.txt, and E-20260822c had already coded Arnold's verse FORM from line-break patterns rather than italics for exactly this reason without naming the cause. Rule: before coding anything typographic off an Internet Archive text, fetch _abbyy.gz; and establish the book's BASELINE italic use before reading any single italic as a marking — Arnold 1899 sets all his verse in italic (24.6% of the chapter) and Gladwin 1806 uses italic only for running heads (1.8%), so the identical observation means opposite things in the two volumes. | S235 | S235 |
| (bsl) | A "does not cover this span" exclusion is a claim about a text and must be verified in the text, not from the title. E-20260827-declared-play declared "Arnold 1899 (frompersianguli00arnogoog) is a verse selection and does not cover this span. Declared, having been checked, not assumed" — and that volume, the first four bábs, prints باب دوم whole from its tale I. The right identifier was named and the wrong book's scope was attached to it: gulistanbeingro00arnogoog, a different Arnold, is a verse selection. The study lost a hand it had, and A-gulistan-hands §5 carried the false statement for eight sessions. Rule: an exclusion on grounds of scope is recorded with the location where the span's own first sentence was looked for and not found — a page or line number, the same evidence an inclusion carries. Sibling of (cde): a title is not an identifier, and an identifier is not a table of contents. | S235 | S235 |
| (bsm) | A cap probed on a small batch is not a probe of the batch you will run. E-20260831b probed openai/gpt-5.6-terra on ten matching items at max_tokens 1500 — clean, 254 completion tokens — and then dispatched the same template at twenty items per call at a cap of 3000. Five of the first ten calls returned finish_reason: length with nothing parseable, at $0.060 each: $0.301 spent for no data. The visible answer scales linearly with the batch; the hidden reasoning does not, and on this seat it grew faster than the answer. Note (bsf) says a cap is a per-seat, per-task-shape measurement; this says the batch size is part of the shape. Fires at: any run that probes a stage on a few items and then batches it. Remedy: probe at the batch size you will actually dispatch, and treat a probe at a smaller batch as evidence about nothing but that batch. Kin to (bsf), (bsh), (brx). RS-20260831b-radif-hands §9. Second firing, S237, and it is the note working as intended: opening 167.955767498, closing 171.717925748, delta $3.762158250 against a per-request sum of $3.465563850. The session reported the pair as an observation and ledgered the per-request sum, without claiming either agreement or disagreement — which is all this note ever asks. RS-20260901-inversion-habit §8. | S236 | S237 |
| (bsn) | Before a retrieval or matching stage, ask what the seat can see of the outcome — and mask it. E-20260831b matched Payne's 199 English odes to a Persian census by showing each ode's opening couplet. Two independent critics put the same finding first: the maṭlaʿ's second hemistich carries the first instance of the very repeated tail the study measures, so the match could be made partly on the outcome. Hiding the precomputed code was not enough — the raw text carried it. The last three words of each hemistich were removed and accuracy did not move: 10 of 10 held-out anchors masked, 10 of 10 unmasked. The mask cost nothing and removed the objection entirely. Fires at: any stage where an instrument is pointed at material that also contains the dependent variable — matching, retrieval, alignment, deduplication. Remedy: enumerate what the prompt's raw text reveals about the outcome, mask it, and probe the masked version against known items before adopting it, because a mask that destroys the signal is worse than the leak. Sibling of the shared-rater rule the S234 critics raised. RS-20260831b-radif-hands §2. | S236 | S236 |
| (bso) | The OpenRouter key-usage endpoint lags, so a same-session delta is not a cross-check. RS-20260829-radif-hands reconciled its key delta against its per-request sum exactly, and the practice of reporting that as a check has stood since. At S236 the delta read $5.674416850 against a per-request sum of $1.947485275. A control call costing $0.000086 moved the endpoint by $0.000000000, so the endpoint is not counting something extra — it is behind; and the $3.727 residual is within $0.12 of the two preceding sessions' ledgered totals. Fires at: any session that reports the key delta as confirmation. Remedy: the per-request sum with "usage": {"include": true} is the ledgered figure and the only one, exactly as CLAUDE.md says; report the delta as an observation with its snapshots and never as agreement or disagreement, because an endpoint that can be days behind can produce either by accident. A second, smaller rule from the same session: write raw bodies to a tag unique per call — a fill pass that re-used its first pass's tag overwrote two bodies and $0.056 of the record. RS-20260831b-radif-hands §9. | S236 | S236 |
Translation practice
| id | note | first seen | last fired |
|---|---|---|---|
| (blx) | A collation's ⚑ is a prediction about the translation, and it should be scored as one — on the only span where that has been checked it was wrong three times in four. collation.md marks ⚑ on readings judged to change an English word. Span F re-rendered the two chapters those marks sit in and coded the result: of the four ⚑ candidates, one reached the English. ¶27's mondák/mondták — an archaic narrative past against an ordinary one — reached the same English because a register rule (V6) sends every speech tag to a plain English past; ¶32's fevde/fedve reached the same English because the copy-text is corrupt there and both renderings follow the other witness; ¶11's Máthé/Máté reached the same English because the copy-text contradicts itself six times to one. And two classes the file had never reported were both found by translating rather than by the instrument — a paragraph division (47 paragraphs against 46) and an italic — the first of which is guaranteed to reach the English, because the paragraph is the unit a serial translation is indexed in. Rule: a collation reports paragraph division and typographic emphasis as their own classes, and no ⚑ count is used as an estimate of what a copy-text change costs a translation — the register absorbs some of them and the estimate is available only by rendering. This does not weaken the gate: the same four runs of it found four real corruptions and three false published rows. ES-20260810-craft-report-legend §2.3, RS-20260810x §7. |
S154 | S154 |
| (bhc) | A translation artifact's body must be terminated by a --- rule, or a shared extractor pools the translator's log into the translation. tools/ngram_overlap.extract() takes everything after ## The translation and, having found that heading, skips later ## headings instead of stopping — so a log, a contamination section or an erratum under a ## heading becomes part of the "translation" for every overlap figure computed through it. Audited across all 86 stored artifacts at S083: 60 raise SystemExit (loud), 16 clean, 10 would pool — and no published figure is false, because every one of the seven whose pooled section is a log was read by a local slicer instead. The fix is a contract, not code: ## I inside a translation and ## Translator's log after one are indistinguishable to the function, so the rule is on the artifact. Every new type: translation page ends its body with a --- rule before any further section, and any design passing an artifact to extract() asserts that the extraction contains no front matter and no log. RS-20260801f-tierD-stages12 §7. |
S083 | S083 |
| (bhb) | The lead is not an independent sample of itself, and no design in this project has priced that. Two lead renderings of one source under one frozen brief, written in different sessions with neither seen by the other and the earlier never opened, returned longest identical runs of 37, 28 and 27 tokens and 104 / 128 / 50 shared 12-grams. The 37 is the opening sentence of Tarchetti's story, word for word, thirty-three sessions apart from the Italian alone. For scale: the largest run this project has ever measured between a lead rendering and a PUBLISHED HUMAN translation is 24 tokens, and against three blind outside models on the same material the lead's own maximum is 2.8× smaller than against itself. Rule: a re-rendering by the lead is not a fresh draw and may not be used as an independent replication, a null, or a second opinion — and where a design needs an independent rendering, the independent agent has to be a different agent. The mechanism is not identified here (weights, decoding, source forcing and a briefed regime are all live) and the finding does not need it: the quantity is the constraint. RS-20260801e-lead-carryover §4. |
S081 | S154 — fired as the ALTERNATIVE EXPLANATION a design could not exclude. RS-20260810x re-rendered two chapters the lead had rendered 22 sessions earlier and measured pooled agreement at 0.734, against this note's own references of 0.814 for a blind self-rendering and 0.317–0.650 for two different hands. Every convergence the page reports — including twelve of the translator's own logged decisions coming out identical — has self-similarity as a rival explanation the design cannot separate from rule-following, and the page says so in its one-sentence summary rather than in its limits. S128 — fired as a MISSED GATE rather than as a caught one, which is new. E-20260807d ran its contamination check against published comparators only and never asked whether the project had already rendered the passage; it had (S027, §4 entire, eleven days earlier). Measured at hand-off: 218 shared 7-grams, 74 twelve-grams, 35 fifteen-grams, longest run 24 against 160 / 27 / 12 / 24 for two independent published translators — the lead is closer to itself than two published hands are to each other, on every column but the tie. Rule added: the contamination gate's comparator set includes the project's own prior renderings of the same work, and a git ls-files on the work's translation directory is the whole cost of checking. RS-20260807d §6 limit 8. S184 — fired as a CAUGHT gate with the OPPOSITE result, which is new and bounds the note. Two lead renderings of the same 668 words, written back to back in one session under opposed frozen rule sets (R08 resistancy, then R07 fluency), measure 7 shared 7-grams, ZERO twelve-grams, longest run 10 over the whole span and 3 / 0 / 8 over the judged segments — against this note's headline 37, and against the 41 RS-20260813h §8 measured between two arms of one session under unopposed briefs. So the self-similarity this note names is not a floor set by the hand: an opposed frozen rule set collapses it, and the two arms of a paired design are not automatically the danger case. The note is unchanged as a rule — measure, never assume — and gains its first measured low reading. RS-20260814f-carriage-elevation-2 §8. S129 — fired as a CAUGHT gate, the first time. E-20260807e ran the repository check before the source was chosen (孔乙己 / Kong Yiji / 咸亨, zero hits) and recorded it on the provenance page as a precondition of selection rather than as a hand-off discovery. The check cost one grep. S138 — fired as a CAUGHT gate and at the highest DENSITY yet recorded. E-20260808g measured the lead's new chapter III against the project's own prior renderings of the same paragraphs (S077, R10, 61 sessions earlier, not read): 68 shared 7-grams, 29 twelve-grams, 18 fifteen-grams, longest run 27 contiguous tokens — in 689 words. The note's headline 37 is a longer run in a much longer text; this is the densest self-match on record, and the run is ordinary narrative prose, not a formula. The same artifact is clean against the published human comparator on the same passage (1 twelve-gram, run 12, in 3,479 words), so the asymmetry the note names is reproduced in one artifact against two comparators at once. New observation, not yet a rule: the convergence is with the period-idiomatic prior target only; the close prior target of the same session is clean at run 9. The lead wrote none of the experiment's arms in consequence.** |
| (bgm) | A model asked to re-translate a passage that is already in its context may return its previous rendering nearly verbatim under a contradictory brief. deepseek-v4-pro, given one 285-word Hungarian passage under an unmarked-contemporary register catalogue and then the SAME passage under a period-idiomatic one in the same conversation, returned 360 words against 360 words, a longest common run of 296 tokens and 294 shared 12-grams — the same text, re-paragraphed. Two of three seats did not do this; the effect at those seats is +4 tokens. Every matched pair this project has built was written by one agent in one context, and this is the failure mode that construction is exposed to. The check is free: measure the two arms against each other before measuring either against anything else, and say what the number is. RS-20260731h-carryover §4. |
S076 | S076 |
| (bfj) | A translation made AFTER reading a claim is not a test of that claim. Declare the priming on the artifact and exclude the affected sites from any test of it. The instance this generalises: T-malory-worship-R04-v1 declined a free copy at ten sites and held a thread on one lexeme, which reads as a counter-instance to A-yosano-yomogiu §1's mechanism — by a translator who had read §1. The counter-instance is not thereby wrong; it is simply not evidence, and a page that quotes it as evidence is quoting the priming. |
S068 | S068 |
| (f) | When a study limb reads other renderings of a passage the session will also translate, translate first, and commit the freeze before fetching. Fifteen sessions running; S036 kept the freeze but broke the spirit — see (abm). | S017 | S039 (span 2: nothing of any comparator was fetched, opened, printed or grepped) |
| (q) | The paired unit's translation limb is load-bearing, not decorative — a study-only session would have shipped a false pass. | S021 | S021 |
| (k) | When translating theory, the translation is the instrument. | S019 | S020 |
| (i) | A term's standard English gloss is not evidence about the term. | S019 | S020 |
| (jj) | The lead's self-estimate is uninformative — about rank, and about localisation. Now 1 for 6. S060: 2 for 7, and the second hit is the weakest kind. E-20260729h's P8' predicted that at least 4 of 6 clean-stratum units would still rank ≥ 30/42, and exactly 4 did — the prediction landed on its own line, under both variants. A self-estimate that holds at the boundary it set is evidence about the boundary, not about the estimator. |
S025 | S060 |
| (ll) | A systematic effect present in every cell has a geometric explanation before a psychological one. | S026 | S027 |
| (ss) | The independence of a published baseline is measurable — count shared 12- and 15-grams. Correction from S031: the count is not weighted by informativeness, and two of the eight longest runs in the Turgenev cell are almost all function words. S036 adds a second, measured caveat: the frozen metric does not break runs at paragraph boundaries, and on one cell that inflated the longest run by a token and the 12-gram count from 2 to 6. | S030 | S038 (a lead-vs-published Japanese cell: longest run 16 across a paragraph break, 14 within one) |
| (bce) | A frozen translator's log is a decision map, and an overlap measurement can be read against it — but compute the test's null before reporting the count. Memory and forcing make different predictions about where a shared run lands: remembered text should surface at the memorable places, which are the hard ones the log records; forced text lands where the pair leaves no room, which is where a log records nothing. S038 amendment, and it downgrades the result that created this note. The test's power scales with the shared fraction, and neither cell the project has run has enough of one. S037's Bécquer cell: 11.7% coverage, 19 located sites, 1 inside — **P(≤1 inside | uniform placement) = 0.168, suggestive rather than the separation the original wording claimed. S038's Poe EN→JA cell: 2.12% coverage, 28 sites, 0 inside — P(0) = 0.377, i.e. no power at all, and an all-outside result there means nothing. So: run it, and report the null probability with the count; where the null is large, report the direction and the content of the runs, not the tally. The qualitative reading — is the run at a place the pair forces? — is what carries, and it is a judgment. n=2 cells; the log is self-report, so "outside the log" is not "nothing to decide" (S038's 16-character run contains an explicitation both translators made and neither logged). runs_vs_decisions.py (whitespace) / runs_vs_decisions_cjk.py (character). |
S037 |
| (bcf) | A hand-maintained column beside a computed one will drift, and a checker that reports both without comparing them will pass while wrong. check_balance.py computed arm staleness from last_worked and read track sessions since from a typed column; at S037 hand-off three of six tracks were one short and the run still exited 0. Repaired by checking the column against the arithmetic. Where a tool prints two numbers that must agree, make it assert they agree. |
S037 | S037 |
| (bcd) | Measure the translator's contamination on the material before designing the experiment, not as a diagnostic inside it. One longest-common-run call per candidate unit, before any locus is selected and before anything is translated. A selection gate, not a flag. S036 sharpens the wording: for a work chosen for serial translation the measurement cannot precede all translating — it is run on span 1, and the session must be prepared to discard span 1. S042 adds the other half, and it is the half that bit: measure-once is not enough on a serial work — re-measure at the midpoint. «Jeli» was declared suspected at S036 on span 1, which had been primed by accident. Spans 2 and 3, both unprimed, measure 14 and 11 tokens against span 1's 15, with shared 12-grams going 10 → 9 → 0; span 3 alone would be declared none. Span 2 had never been measured at all. A declaration set on one span of five describes that span, and the drop within the unprimed spans was as large as the primed-to-unprimed step. RS-20260727d §5. S043 records the case the note does not cover: the gate cannot be run at all when no comparator is reachable. 「最後の一句」 has exactly one English translation, in copyright and unobtainable, so T-saigo-no-ikku-R04-v1 declares none UNMEASURED, with the bibliography as its basis and a free pre-registered self-probe recorded as the weak substitute it is. An unmeasurable work is not a disqualified work — the gate's binding clause is about canonical translations — but the declaration must say which of the two it is. S044 is the first time the gate ran on a canonical comparator and returned a verdict against the artifact: Lu Xun 1930 against Yang & Yang 1959, longest run 12 tokens, and the forced-vs-recalled control (bcp) refused the forcing defence, so T-yingyi-jiejixing-R04-v1 is declared high on a measurement rather than suspected on a prior. Two sessions of none-UNMEASURED were followed by the first measured high, which is the note working. |
S031 | S050 (second consecutive session in the prescribed order — Unit A translated, gate run, Unit B drafted only after; 8 tokens against a null-control floor of 5. And the first priming declaration the note has produced: the lead had seen the comparator's English of the story's first third while checking one existed, so the passage was chosen beyond it and the single term that crossed the line is named on the artifact) · S063 — same ordering as S061/S062, stated as the limitation it is; the gate returned an 11-token run with no lexical choice in it and the prior of suspected, recorded separately before the measurement, was held rather than overturned · S071 — the ordering could not be obeyed (the comparator IS the object of study) and is stated as the limitation it is; the prospective declaration suspected was overturned in BOTH directions by the measurement, to none on one limb and high on the other |
| (bcn) | A count is only a count if the class it counts over is partitioned — and "the class the finding was recorded in" is usually not that class. S042 found a criterion that had never defined a set (bcl); S043 found the mirror case, where the criterion was fine and the class was not. RS-20260727-log-typology reported "twenty-three decisions sit on that seam", over the class C4a. The frozen corpus holds 26 such rows, and measurement shows 13 of them were never contested at all — the non-social half (tense, middle voice, grammatical gender, impersonal on, pre-nominal participles) agreed at 0.923 under the very wording the seam was alleged in. Before reporting a count over a class, ask whether every member of the class is actually subject to the thing being counted; annotate the split in advance if not. Costs nothing when done at build time: S043's social flag was frozen before dispatch and an independent critic then named eleven of the same thirteen. |
S043 | S043 |
| (bco) | Write API key-usage snapshots to disk, never only to stdout. S043's critic call took an opening and a closing snapshot and printed them; the terminal output was later truncated and the cross-check for that call was lost permanently, leaving it on per-request cost alone. The twelve-call run, whose script wrote cost.json, cross-checks exact to 1e-7. Same session, same endpoint, two practices, one survivor. Also confirmed again: the endpoint lags — it returned the unchanged opening figure immediately after the twelfth call and had settled on re-read, which is the S037 behaviour, so a snapshot taken at the instant of completion is not evidence of anything. Note (abf) says take the snapshot; this says persist it and re-read it. Applied S044: the session-start snapshot and the probe run's snapshots were written to disk (snap/, runs/cost.json); the ratification calls went through tools/ratify_vote.py, which still does not persist one, and are on per-request cost alone. The tool, not the operator, is where this note has to land. |
S043 | S062 (two settling lags, 0.0042 and 0.0075, both asserted as bounds) · S063 — the excess direction for the first time, which is note (bet) · S067 · S071 — VOID, see (bfv): the key moved $1.149 between two of this session's own calls while it made none · S075 · S077 — short by $0.060538300 against a per-request sum of $0.580589442 · S078 — EXACT to 1e-9, delta 1.323350575 against a per-request sum of 1.323350575 |
| (bcp) | A shared run between your translation and a published one is not explained by "the source forces it" until you have made the source force it for somebody else — and the test costs three sentences and three calls. S030 refused the forcing defence on Latin by argument; S044 refused it by measurement, on a lead translation, for the first time. T-yingyi-jiejixing-R06-v1 shares a 12-token run with Yang & Yang 1959 (but there is still a great deal of paper in the world). Three independent panel models, given only the Chinese sentence, all wrote plenty of and none reached 7 contiguous tokens of the run — while at a different site, 於社會上有些用處, all three produced the lead's seven tokens verbatim. The probe separates loci rather than blessing or condemning them together, which is the only reason its verdict is worth anything. Pre-register the failing branch: S044's design said in advance that a NOT FORCED result on the headline run means contamination: high, and it fired. Two riders. (i) One-sided: it can refuse a forcing defence, never establish independence — recall, a shared idiom prior, and two translators freely choosing alike all predict the same result. (ii) Running the measurement between draft and revision is itself a priming event — the lead then knows which of its own strings match — so the matched sites must be left exactly as drafted, and that rule must be written into the log before the revision starts. Changing them is optimising against the instrument. |
S044 | S044 |
| (abl) | (fired S038: the Japanese extent figure was produced in Python.) wc -w is not a safe word count for non-ASCII text. GNU coreutils 9.4 under LC_CTYPE=POSIX undercounts across multibyte characters — 542 whitespace-delimited tokens reported as 533 on one French page, reproducible on b\xc5\x93ufs tranquilles, \xc3\xa0 la robe (4 against 5). Count with python3 … .split() or awk NF, which agree. Whether earlier extent figures in this project's Russian, Japanese, German and Latin artifacts were produced this way is unchecked — see wiki/backlog.md. |
S035 | S035 |
| (abm) | Verifying a fetched comparator is itself a reading event. Note (f) says download-then-translate is safe if the comparator is not read; S036 downloaded a published English translation, printed 900 characters of it to check that the extraction offsets were right, and thereby primed 9% of the span it then translated — the measurement's longest run, 15 tokens, landed inside that 9% verbatim. Verify a fetched text by counts, offsets and heading positions only; never print its prose. The habit that caused it is the good habit (check your extraction) applied to the one class of file where looking is the thing forbidden. Applied cleanly S044 and it worked: a 560,882-character OCR scan was reduced to the right 8,289-character slice using only all-caps heading lines and standalone roman-numeral markers, with no prose displayed, and the resulting measurement was identical on the matched span and the whole essay — so the offsets were right. The note is now positive evidence, not only a scar. | S036 | S062 (seven comparator windows located by heading offset, no prose displayed) |
| (bdf) | A markup-stripping tool must turn what it cannot represent into something visible, and must count what it turned. tools/fetch_aozora.py deleted every JIS X 0213 gaiji in Aozora's HTML edition with its generic <[^>]+> strip, and counted the ※[# form that appears only in the plain-text edition — so it reported 0 deletions while making 22, and three canon manifests then recorded "no gaiji notes" as verified fact. Twenty sat inside stored members; one was the noun in the sentence a paragraph turns on, leaving a copula with no complement. The failure was silent at 20 of 22 sites because the surrounding text stays grammatical-looking. Two rules follow. (i) A cleaner's unrepresentable case degrades to a loud marker, never to nothing. (ii) A count of what was stripped must match the form the parsed edition actually uses — a counter aimed at the wrong edition's convention reports zero forever and reads as a clean bill of health. RS-20260728h-gaiji-deletion. |
S051 | S051 |
| (bdi) | An anchor set written in your own words is biased toward finding agreement, because at the sites where the comparator diverged it is your words that are least likely to be there. S052's pre-run critic replaced a discretionary "no counterpart" rule — which let the lead rescue its prediction by declaring an inconvenient site non-corresponding — with lexical anchors frozen before the comparator was opened. Seven of nine hit. The two that missed were the only two sites in the study where the comparator disagreed with the lead's frozen prose, i.e. exactly the two cells that falsify the design's central prediction: one anchor required sulphate where the comparator substituted quinine (a realia substitution, strategy three on cultural-mediation's own list), the other required one of five verbs where the comparator wrote "do for himself". Applied mechanically the control would have confirmed the prediction it was written to protect against. The mechanism is general and is not bad luck: an anchor is a list of the words you chose, divergence in wording is largest exactly where divergence in method is, so the instrument's misses are not random with respect to the hypothesis. Anchor on the source, not on your own target text — or accept that a miss is evidence and go and read the passage. |
S052 | S052 |
| (bdj) | A verifier that has never disagreed with the analysis has not been shown to be capable of disagreeing. S052's verify.py re-implemented a frozen grading rule independently and, on its first run, contradicted the stored table at three cells and produced two formatting mismatches. All five were the verifier's defects — a token lookup grading "would know" as a finite present, and a subjunctive regex that caught only the inverted form — and fixing them changed no datum. The disagreement is the evidence that the check is live. Record the verifier's own first-run failures on the result page: a bare "N checks, 0 failures" from a verifier written after the numbers were known is compatible with a verifier that cannot fail. S053 discharges the note by demonstration rather than by luck, and the method is one function call: its verifier passed 102 of 102 on the first run, so two mutations were injected into the stored results — one Fisher p shifted by 0.01, one cell count changed from 92 to 91 — and the verifier caught both, named both, and exited non-zero; restoring the file returned it to 102/102. A clean first run is not evidence; a mutation test is, it costs one command, and it should be run whenever the verifier does not disagree on its own. |
S052 | S058 (mutation-tested a second time and by default: 67/67 clean, 65/67 under two injected errors — an agreement ratio and a Jaccard. The note is now standing practice rather than a reminder) |
| (bdk) | A page's prose and a page's tally are two readings, and nothing makes them check each other — so grep your own prose for the cells your table scored negative. A-beowulf-ingeld §4.3 scored Morris refuse at the rǣdan site and published "at total drift, nobody is tempted, 0 of 12"; §4.4 of the same page, written by the same reader in the same session, quotes *"thou shouldest arede" (rǣdan)* as an example of Morris's archaism. Two independent raters found the cell at S053 and one string match confirmed it, seven sessions after the page was written. Two of the three cells they found are this shape. The mechanism is that a tally is built by scanning for a form and prose is built by reading for an argument, and the two passes have different failure modes — the negative cells of a table are exactly the ones the prose is most likely to have already contradicted, because a negative cell is "I did not find it" and prose is where finding gets written down. Cost of the check: one grep per negative cell, against the page's own text. Sibling of (bcu), which is the same failure between a rule and the corpus it governs; this is between a count and the argument built on it. |
S053 | S058 (a second anchor, five sessions later: A-yosano-yomogiu §2's table records in a parenthesis that Yosano writes 赫耶姫 for かくや姫, and §4 of the same page then states that "all three titles copied" and uses that uniformity to narrow C4. Two independent readers both score the title a SUBSTITUTE. The note is no longer a fact about one page) |
| (bdn) | A comparator's tail, printed during material triage, primes the site the triage was for. S055 considered Andreyev's «Баргамот и Гараська» as a translation limb because the story's hinge is an address form — and established that the stored Lowe comparator ran to the end of the story by printing its last 600 characters, which are that hinge. The story was rejected as a limb on the spot. The general shape: you check a comparator's extent by looking at its edges, and the edge you look at is chosen by the thing that made the work interesting. Cost of the check done safely: one wc -c and one paragraph-count, neither of which prints prose. Sibling of (bcd), which says measure contamination before selecting; this says the measuring can itself contaminate. |
S055 | S055 · S063 — the comparator's extent was checked with wc and a word count only, printing no prose, before the translation was written · S067 |
| (bdo) | A re-derivation defeats every priming guard the derived artifact was built to hold. SR-20260725b-nation-1904-excerpt.txt §8 deliberately did not transcribe the twenty paired Hapgood/Garnett extracts the 1904 review prints, on the stated ground that "storing them here would prime any future blind reading of this novel", and recorded only where they are. S055 discharged condition (iii) by regenerating the whole page from the item id — which reproduces those extracts in full, and this session has now read them. The guard was correct, the re-derivation was owed, and the two are incompatible. Neither is the error. What is owed when they collide is a declaration, on the artifact whose guard was defeated, naming who read what: a future session designing a blind Turgenev reading must know that S055 is contaminated on A Nobleman's Nest. |
S055 | S055 |
| (bdp) | A frozen regex checklist over generated prose measures vocabulary, not carriage — and it fires hardest exactly where the point is subtlest. E-20260729c froze a 17-point coverage gate before its summaries existed, which is the right discipline, and its most important item — does the summary draw the consequence that every printed exhibit comes from one work? — registered HIT on both non-lead summaries. One of them asserts the opposite ("drawn from those works", plural, which is false); the other states the volume numbers and draws nothing. The gate cannot distinguish naming a thing from carrying it, and its headline number is therefore an upper bound. The registered prediction that depended on carriage was tested separately, on the voices rather than the summaries, and failed — so the two instruments disagreed and the weaker one was the one with the number. A completeness gate over paraphrase needs entailments, not phrases, or a second reader; a regex list is a presence check and should be reported as one. |
S055 | S055 |
| (bdu) | A positive control's MAGNITUDE is a design parameter, and a control weaker than the material it calibrates cannot calibrate it. E-20260729e froze sixteen controls before its reader pass — the right discipline, and the pre-run critic had strengthened the gate they feed. The eight surface controls (a contraction, a misspelling, a comma for a dash, an adverb moved) scored E = 12.87, against the corpus's own median real edit at E = 27.33: they were smaller than the things they were meant to bound, so the E-side criteria failed and the session's principal reader-based finding was withheld under its own registered rule. The meaning controls, built at full strength, separated at 89.4 against 0.0 on the same axis — so the instrument was fine and the control was not. Set a control's size from the material's measured distribution, not from the designer's sense of what the category means. Companion to (o): a threshold must be reachable by the thing it is applied to, and a control must be comparable to the thing it calibrates. |
S057 | S062 — NOT the explanation of the case it was written about (RS-20260730c §2); true in general, narrowed in application |
| (bdv) | Between a translator's log and a mechanical diff, containment runs in BOTH directions, and a rule that assumes one direction silently reports a near-zero. A log quotes whole phrases ("moisten the spring in secret" → "moisten the springtime in secret"); a diff reports the minimal difference inside them (spring → springtime). E-20260729e's frozen §6 tested only quote ⊆ span and matched 2 of 12 decisions in a log that itemises ten changes — a result that would have read as a dramatic finding about self-report rather than as a broken rule. Two further traps in the same rule, both found by running it: a one-token span like a is contained in almost every quoted phrase and matched eleven of fifteen decisions until spans were grown by their own unchanged context; and a keyword test for declared non-changes run before the match test misfires on kept/keeps/holds occurring inside descriptions of real changes. When a near-zero comes out of a string-matching rule, suspect the rule before believing the zero. |
S057 | S057 |
| (bdw) | On an INTRALINGUAL pair the contamination instrument measures the study's own variable, and its verdict vocabulary inverts. tools/dependence_check.py run on T-malory-worship-R06-v1 against its Middle English source returned 58 shared 12-grams and a longest common run of 25 tokens, and printed DEPENDENT?. Every published overlap figure in this project is between a lead translation and another translation, where a long shared run is evidence of dependence; between a translation and its own source in the same language the same run is the free-copy option being taken, which is the phenomenon A-yosano-yomogiu §2 exists to study. The same number means opposite things and the tool cannot tell which case it is in. For scale: this project's measured lead-versus-published runs span 0 tokens (Ovid) to 21 (Turgenev); this is 25, against the source. Do not carry an overlap figure across the intralingual/interlingual boundary, and do not read the tool's verdict column on an intralingual pair at all. Companion to (bcd), which says measure before selecting; this says know what the measurement means before reading it. |
S058 | S058 |
| (bdy) | A census Jaccard cannot separate a vague category from incomplete enumeration, and without a completeness control the low number gets read as the first when it may be the second. E-20260729f asked two models for a determinate list — every number written in words in an 821-word English passage — and got nine items and six, Jaccard 0.667, on a task with a right answer. On the same call the vague task (culture-bound items) landed at 0.316 on that passage and 0.778 on another. The vague-task figure is only interpretable where the determinate one is high. RS-20260729-drift-window-verify §7 read a census Jaccard of 0.4375 as evidence that "the criterion cannot be applied by a reader who is not its author" and ran no completeness control at all; that inference is narrowed as a result, and what survives is the two lists differ. Any design that measures set agreement must, in the same call, measure agreement on a set whose membership is determinate — and must say what it will conclude if that one fails. Sibling of (bdu): that note is about a control being the wrong size, this one is about the absence of a control that says whether the instrument can enumerate at all. |
S058 | S058 |
| (bea) | Repeat every condition, not only the treatment — a one-condition repeat measures the treatment's noise and then gets read as the instrument's. RS-20260728f-nonlead-items repeated its R1 condition only, measured a byte-identical swing of 0.243 in Krippendorff α, and closed ARM-sense-boundary on the reading that "the instability belongs to the item format". E-20260729g repeated all three of its conditions on the same forty lead-free items and found the page's own wording swings 0.075 — a third as much — with per-rater self-agreement 0.839–0.919 against S049's 0.817. The 0.243 belongs to the rule condition, not to the items. The narrowing is partial: self-agreement on lead-free items is still below the 0.956 measured on lead-written ones, so some stability was lost; the instrument was not destabilised. A design that repeats one arm of a comparison cannot say which arm the noise is in, and the cheapest possible fix — double the calls — is what turns a swing into an attribution. Companion to (bdu) and (bdy), which are about a control's size and a control's absence; this is about a control's coverage. |
S059 | S059 |
| (beb) | An option nobody ever takes is not an option, and a five-way instrument in which two categories go unused is a three-way instrument being reported as five. composite was offered in every condition of E-20260727c and E-20260728f and used zero times; E-20260729g offered both composite and both on 240 cells across two source languages — including 22 Korean sites chosen because grammatical deference and social role coincide there — and both were used zero times again. Three experiments, one unused option throughout and a second joining it. Consequences: a chance-corrected agreement figure computed over a five-category vocabulary is corrected against categories that do not exist in the data, and any derived label that CAN assign the unused categories is not comparable to the direct one. Before quoting an agreement statistic over a fixed option set, print the realised distribution and say how many options were actually used. CORRECTED 2026-07-31 (S070), RS-20260731b-figure-audit, from the stored bodies. Three of this note's factual claims are wrong and its consequence clause is false. (i) E-20260727c never offered composite — it offered four options (run.py:105–108). (ii) In E-20260728f, composite was used 4 times and both 3 times, all by P1 in conditions A and C; the note read the two zero columns (B, B2) and dropped the two non-zero ones, which the result page itself prints correctly. (iii) E-20260729g offered them on 372 cells, not 240, and composite was used once; only the both-at-zero claim holds. And the consequence clause — "a chance-corrected agreement figure computed over a five-category vocabulary is corrected against categories that do not exist in the data" — is false: Fleiss' κ, Cohen's κ and Krippendorff's α all take expected agreement from observed marginals, so a zero-count category contributes exactly zero. Thirteen published figures recomputed, thirteen nominal-versus-realised deltas of 0.000000000. What survives is the instruction: print the realised distribution. What does not survive is the reason the note gave for it. |
S059 | S070 |
| (bec) | Repairing an instrument does not repair the figures it already produced, and this project has never once re-run one. tools/ngram_overlap.name_tokens was repaired at S031 with two defects diagnosed and a legacy function retained specifically so old figures could be reproduced. Twenty-nine sessions later nobody had used it. S060 re-ran S027's own pipeline behind a reproduction gate — the legacy rule reproduces the stored rows 42 of 42, the repaired rule differs in 31 of 42, every difference in the same direction — and the published 17.8× became 14.23×, a number three result pages quote. The reproduction gate is what makes this cheap and safe: reproduce with the old rule first, so a difference is attributable to the repair and not to the re-run. Rule: a repair to a shared instrument carries an obligation to name the figures computed with the old one and either recompute them or record that they were not recomputed. A repair that leaves its outputs standing has converted a known defect into an unknown one. |
S060 | S060 |
| (bed) | A subset defined by removing 'bad' observations is not a neutral subset when the badness criterion correlates with the measured quantity — and it can bias in the OPPOSITE direction from the bias it removes. RS-20260726b-baseline-dependence prescribed reporting every rank twice, over all units and over the dependence-clean subset, on the reasoning that a dependent pair inflates the baseline and so biases a rank against finding the lead unusual. True of the all-units figure. But the flagged units are the high end of the distribution — published median 122.85 against 58.48 clean — so deleting them deletes the top of the reference distribution and any rank inside the remainder is mechanically easier to be high. The prescribed repair errs the other way, and the page that prescribed it asserted it could not. Before reporting a statistic over a filtered subset, measure the filter against the statistic and say which way the filtering pushes. Related to (bee): both are about a rank whose reference set was chosen rather than sampled. |
S060 | S060 |
| (bee) | One observation placed inside a measured distribution is a draw from a distribution of its own, and a rank has no error bar until the numerator's spread is measured. RS-20260726-period-control §3 called a single figure — the lead overlaps Garnett more than 41 of 42 published pairs — "the strongest overlap result the project has", and it stood for thirty-three sessions as a property of the lead. S060 translated six further units of the same book blind: the lead's own rate varies by 5.72×, the six rank 6, 16, 37, 41, 42, 42 of 42, and 41/42 is the ceiling of the lead's range, not its middle. The instrument was never wrong; the inference from one placement was. Rule: before a rank inside a distribution is reported as a property, either measure the numerator at several points or state in the same sentence that it is one draw. The project spent great care on the denominator — 42 published pairs, a 17.8× spread, a downgrade of CI to ordinal — and none at all on the fact that there was one number on top. |
S060 | S060 |
| (bef) | A harness that skips an item on falsy input makes a gap invisible; an enumeration gate turns it into a warning. dependence_check.run() skips any unit whose text is empty and records nothing. E-20260726b/build.py fed it a {title: body} dict built over a heading list containing eight duplicate titles, so A CONVERSATION resolved to an empty speaker label; the unit vanished, RS-20260726b published "13 / 41" without naming the forty-second, and four sessions of ARM-baseline-restate being named and skipped went by with the reference distribution's own minimum unchecked. The pre-run critic added gate F6 — the results must contain exactly N enumerated units, checked against the manifest, or every downstream figure is blocked — and it failed as stored. Two rules: never use a name as a key without checking it is unique, and assert the expected count of a computed collection against its source manifest. Sibling of (bdy), which is the same failure one level up: that note is about a model enumerating incompletely, this one about the harness doing it. |
S060 | S060 |
| (h) | A program's framing of an unbuilt cell can be wrong in ways that matter. | S017 | S019 |
| (bie) | A rule that says "count it as you go" produces the number late and produces it wrong, because the sites already behind the counter are not counted. register.md V18a made a rule about a grammatical feature the translation cannot carry and instructed spans 4–8 to record every further occurrence so the craft report could state the total. The register's own running count said two sites. A whole-text search, run in one command the moment the question was asked, returned four — and the one it had missed was in span 1, already translated, behind the counter. The rule would have delivered the figure at span 8 and delivered it halved. Rule: when a feature becomes the subject of a register rule, count it over the WHOLE source text immediately, not span by span. This does not collide with the no-rule-for-an-unread-site rule (register.md V17): a count of where a feature occurs is not a decision about how to render it, and V17 forbids only the second. The same command also found a third verb in a family this project had twice declared closed at two (raahtia/hennoa/raskita, D74). | S100 | S100 |
| (bim) | A mechanical diff over an OCR layer is a finding aid and never a witness — and it manufactures plausible substantive variants at about one in seven. S105 collated 5,372 words of a 1917 reprint against an 1886 first edition's OCR text layer: 62 rows survived noise filtering and nine dissolved on the page image, the printings agreeing exactly. Three of the nine would have been published as substantive divergences — laidallahan for laidalla hän, tuollaiseen for tuonaiseen, and murjottivat (sulked) for muljottivat (goggled), which would have changed an English word. The same run's four real corruptions were all confirmed on the image, so the discipline costs nothing it does not buy back. Rule: no reading enters a collation record, an erratum or a copy-text on the strength of a text layer; every divergence a diff reports is adjudicated against the page scan or is recorded as unadjudicated by name. | S105 | S105 |
| (bjn) | Two machine renderings of one source from two different labs are not two independent renderings, and the gap is an order of magnitude. On 740 tokens of Verga, tools/dependence_check.py measured the lead's R04 against an unbriefed deepseek-v4-pro single pass at 90 shared 12-grams and a longest common run of 30 tokens — against 11 and 17 for Strettell 1893 versus Dole 1896, two published human translators three years apart. No shared prompt, no shared instruction, different labs. Note (bhb) said the lead matches itself at up to 37 tokens; this extends it across labs, so a panel model's rendering is not an independent second opinion on a lead rendering of the same source, and any design that needs a second modern translator must either measure the pair or say it did not. Two companions from the same run, both of which bound how the tool may be read: nine of ten pairs came back DEPENDENT?, including S~D, the one pair independent by construction if any is — on a passage this constrained a close rendering produces long runs whoever writes it, the recall floor of (bez)/(bgf)/(bgi) met again — and the informative cells were the zeros: the resistancy arm shared 0 twelve-grams with either published rendering. Read the table as a ranking against the published pair's own figure, never as a verdict. RS-20260806b-berman-occurrence §3. | S118 | S118 |
| (bry) | A second rendering of one text under a DIFFERENT declared policy, made with the first rendering open, is pulled toward it far enough to stop being a different arm — and the size of the pull is a fact about composition ORDER, measured. «المقامة الحلوانية» has two arms by one hand under two policies written to differ: restraint (R43) and Preston's balanced period (R50). On «المقامة الصنعانية», where the balanced arm was written last, they share 19 twelve-grams and a 15-token run. On «الحلوانية», where it was written first and was open on disk, they share 216 twelve-grams and a 48-token run — "and the glittering of his show, I looked hard into his features, and let my eye run over his marks, and there he was — our shaykh of Sarūj". Eleven times the overlap from reversing the order and nothing else. This is note (bhb)'s hardest firing and the first on two texts the regimes were written to make different. Fires at: every multi-policy design in this project — the shape that produced RS-20260825b, RS-20260825c and this run. Remedy: run tools/dependence_check.py on every pair of the lead's own arms before designing an evaluation on them, exactly as the contamination rule already requires against published hands; treat a DEPENDENT? pair as a near-duplicate control rather than as two arms, and drop it from the design by the measurement — which is what E-20260827b did to its BAL arm. And where the order can be chosen, write the arms in the order that puts the instrument arm first. RS-20260827b-shown-or-told §10; workshop/translations/maqamat-hulwan/R43-v1/translation.md §Contamination. | S227 | S229 |
| (brz) | A de-chime must be checked for the chimes it CREATES, not only for the ones it removes — and a hand-checked one will ship both. R49 rule 2 says change the colon-end word of one colon per graded chime. Applied to «الحلوانية», the first pass left two chimes standing (examination ⁄ consideration, men ⁄ woven) and created two that were not in the treatment text at all (meet ⁄ brought, well ⁄ people) — a control arm that chimes where the treatment does not is worse than no control. dechime.py caught all four before dispatch because it re-grades every consecutive colon-end pair in both texts and fails on either direction; it also enforces rule 3 by failing if any colon outside the substitution table moved. Seven loci additionally needed a second substitution, because R49 rule 2 is written for two-colon loci and this maqāma has runs three, four and eight deep. Fires at: every subtractive control this project builds — the R49 family, and any same text with feature X removed arm. Remedy: make the control's construction a script that re-measures the whole text and exits non-zero on a survivor or a creation; extend the rule explicitly where the source's runs are deeper than the rule assumes, and say so on the artifact. workshop/translations/maqamat-hulwan/R48D-v1/translation.md. | S227 | S227 |
| (bsa) | A compound prediction must be checked for ALGEBRAIC ENTAILMENT against what the design itself says was already visible, before it is registered. E-20260827c registered P4 in two clauses — Chappelow's mean clause within 20% of Chenery's, and his clauses per Arabic colon at least 50% above Chenery's — and presented the second as independent confirmation of a mechanism. Within a panel the Arabic denominator is shared, and total words = mean unit length × unit count, so with the word totals the design's own §6 declared already visible (ratios 2.11 and 2.65), the first clause forces the second to at least 1.76 and 2.20. Both pre-run seats derived the identity independently and both called it BLOCKING. The clause could not fail. Fires at: any registered prediction with two or more clauses over quantities linked by an identity, and especially where the design has a what was already visible section — that section is exactly the input the entailment runs through. Remedy: write the identity out before registering; if one clause is implied by another plus a known total, delete it and report the quantity as arithmetic. RS-20260827c-declared-page; E-20260827c-declared-page/critic-response.md. | S228 | S229 |
| (bsb) | A reproduction check on a SIMULATED statistic cannot demand exactness, and one that does is a control set to fail. E-20260827c's C1a required the pipeline to reproduce RS-20260826's published figures exactly, and the design halted the whole run on a C1a failure. Its deterministic half reproduced exactly — 107 and 123 printed lines, longest 14 and 13 words, CV 0.198 and 0.177 — but the six shuffle ratios are Monte-Carlo estimates over 10,000 draws and a different call order gives a different stream, so they landed 0.0001 to 0.0017 away. The pre-run critic named this as a MINOR finding about tie tolerances and it arrived in execution as a control criterion instead. Fires at: every extraction check, replication check or figure audit that re-runs a permutation, bootstrap or shuffle null. Remedy: split the criterion — exact for deterministic figures, within a stated tolerance derived from the number of draws for simulated ones — and state both in the design, not after the run. RS-20260827c-declared-page §8. | S228 | S228 |
| (bsc) | Before declaring that a comparator volume does not contain a work, open its contents. A-hariri-hands §4a stated that Chappelow 1767 does not contain al-Ḥarīrī's second Assembly. It does: his Assembly II is headed HULWANENSIS, and his six are al-Ḥarīrī's first six in order. The claim was inferred at S223 from the shape of his page — running prose, no colon marking, so he had been excluded from that session's measurement — and then written down as a fact about his book. It cost the panel its second declarer for a session, and the second declarer turned out to carry the arm's headline. Fires at: every anchor and every comparator table that records what a published volume does and does not contain. Remedy: a claim of absence about a printed book is a claim about its contents page; check it there, and record where it was checked. Note that the check is free — the volume OCR was already downloaded. RS-20260827c-declared-page §1. | S228 | S228 |
| (bsd) | A preference item whose two members differ by a few words gives a seat nothing to be about, and position fills the vacuum — build the item from two WHOLE renderings and the same seats become order-consistent. Note (brs) recorded the failure at S224: on four line-end arms of one text, differing in three words, the same three seats returned the same label under an order swap at 0.548, one of them below chance, and every registered quantity was withheld. E-20260828 put the same question to the same three seats — which of these two do you prefer to read as English verse? — on two complete renderings of the same ghazal by one hand, and order consistency ran 0.852 at baseline and above chance in all five conditions, with a pooled first-position rate of 0.5296 inside its bar, 0 void of 270, and a 12-of-12 operational floor. So the remedy note (brs) proposed — give the judgment something to be against — is not the only remedy, and on this evidence not the first one to reach for: giving it something to be between worked, and the disclosure conditions that were supposed to supply the criterion moved nothing. Note the direction of the effect on the instrument: the condition that told the seats exactly what the source does had the lowest consistency of the five (0.593) and the most undecided windows. Fires at: any A-vs-B preference task built by editing one text into two. Remedy: where the question is about a policy rather than a word, render the policy twice and whole; a minimal pair is the right instrument for a minimal claim only. RS-20260828-forced-half §4; RS-20260826b-radif §5. | S229 | S229 |
| (bse) | Ask whether a translator supplied a word and you must ask it against the HALF-LINE the English is rendering, not against the couplet — at couplet grain the answer saturates at 95% anchored and the measurement dies. E-20260828 set one English word against the Persian bayt its block renders and asked three seats whether the word answers to anything in it. The seats were not the problem: 12 of 12 on known-answer calibration items, 19 of 24 decoys called supplied, zero UNSURE in 104 study items, pairwise agreement 0.94–0.97, and the same low rate in each seat separately (0.029, 0.038, 0.087). The problem is the haystack: fifteen to twenty-five Persian words in two hemistichs offer something for nearly any English word to be pointing at, so the pooled supplied rate came to 0.048 and the run's own saturation gate withheld both primaries. The design had chosen couplet grain deliberately, to make supplied a conservative verdict; it bought the conservatism with the whole primary. Fires at: any design that codes addition, omission, or correspondence between a translation and its source. Remedy: align to the smallest unit the source itself divides — the hemistich, the colon, the clause — and pay the alignment cost up front; and where the alignment cannot be had, expect a base rate too low to test and say so in the pre-flight rather than after. Kin to (bsd), which is the same lesson at the other end of the instrument: there the item gave the seat nothing to be about, here it gives it too much. RS-20260828-purchased-figure §4. | S230 | S230 |
| (bsf) | A max_tokens cap probed on one item does not verify it for the next, on a model whose hidden reasoning is variable. E-20260829 bought 27 short, uniform, structured classifications at a cap of 1200 and, as note (brt) requires, checked the first three returns for truncation before dispatching the rest. They passed. P2 then truncated on the seventh call — same task shape, same prompt template, a different item — having spent 1,123 of 1,196 completion tokens on hidden reasoning. The visible answer needs about seventy tokens; the reasoning is what varies, and it varies by item, not by shape. So (brt)'s remedy — probe the shape — is necessary and not sufficient. Fires at: any run buying a fixed-format answer from a reasoning model at a tight cap. Remedy: size the cap for the reasoning budget rather than the answer, which on this panel means several times what the visible output needs; and structure the prompt so the load-bearing fields come first, because a truncated body that already carries every field it is scored on is re-parsed at $0 under (brx) rather than re-bought — which is exactly what saved this call. RS-20260829-radif-hands §8. Fired again S232 and the second firing is instructive: at a cap of 300, 92 of 270 rating bodies from the same panel returned finish_reason: length, almost all P2 — but 86 of them had already delivered the scored line and were re-parsed at $0 under (brx), and only six lost it and had to be re-dispatched at cap 1200, for $0.016. So the failure mode is not "truncated" but "truncated past the answer", and only a finish_reason check plus a tolerant parser distinguishes it from a dead call. The remedy below is confirmed, and its second half — put the load-bearing field first — is what turned a 92-call problem into a 6-call one. RS-20260829b-inversion-price §11. | S231 | S243 — eighth session running; S239 was the seventh; S237 was the sixth; see the tail of the note. Earlier: S234 — fourth session running, and now on four seats at once. At max_tokens 2500 on a nine-line answer, google/gemini-3.6-flash and qwen/qwen3.7-max both returned finish_reason: length on the first item probed, and z-ai/glm-5.2 returned an empty body after 6000 completion tokens; at 6000 the two panel seats came back clean and six of thirty-eight stage-S bodies truncated anyway. The cap that works is now a per-seat, per-task-shape measurement and nothing else. RS-20260830b-rhyme-family §7. Fired again S236, fifth session running, and this time it cost money: the cap was probed on this shape for P2 (3000 truncated, 8000 clean) and not re-probed for P1, which had returned a ten-item probe of the same template clean at 1500. P1 then truncated on five of ten batch calls at 3000, at $0.060 each for nothing — $0.301 that bought no data. See the new note (bsm): the batch size is part of the shape. RS-20260831b-radif-hands §9. Fired again S237 — sixth session, and the first run that honoured the note in full and was bitten anyway. Every seat was cap-probed at the batch size it would actually be dispatched at, as (bsm) demands; the probes cost $0.297 and were worth it (4000 truncates P1 and P2 at 24 items; 12000 clears P1 and still loses six of P2's 24; P2 needs batch 12). Then one P2 batch in 44 truncated at the probed cap and the probed batch size, losing 12 items. The note cannot be discharged by probing: probing bounds the loss, it does not remove it. What remains is to parse what came back, count what did not, and report it. RS-20260901-inversion-habit §8. Fired again S239, seventh session running, and in a direction the note does not cover — see the new note (bst), where a LARGER cap made the answer worse. Two of 54 stage-G bodies returned EMPTY after spending the whole 2500-token cap on hidden reasoning and one stopped mid-list; the three dead bodies cost $0.092926 and the re-buys at cap 6000 cost $0.062950, so the abandoned calls were the more expensive half — a body that spends its whole cap and returns nothing bills for the whole cap. RS-20260902-rhyme-slot §8.2. Fired again S243, eighth session running, and in its sharpest form yet: google/gemini-3.6-flash truncated 6 of 12 batches at cap 2000 on a task whose visible output is eight lines of the form LI.11/GAR KEPT, pinning at exactly 1,996 completion tokens each; the identical items at four per batch and cap 6000 finished 4 of 4. Batch size is the variable — (bsm) confirmed from the other direction. config/budget.md 2026-09-04. |
| (bsg) | Never gate a primary on a pooled main effect when the hypothesis predicts an INTERACTION — the gate fires exactly when the hypothesis is most true. E-20260829b v1 registered a manipulation check M1 (pooled inversion penalty ≥ 1.0) and made it a withholding gate on both primaries. Its pre-run critic showed that the design's own best case — a large penalty in one register and none in the other — drives the pooled penalty down toward the gate, so the strongest possible confirmation would have withheld the result. Fires at: any crossed design that gates on a marginal quantity the interaction is defined as varying. Remedy: gate on controls only — a positive control, a compression or floor rule, a completeness rule — because those are properties of the instrument and are independent of which way the hypothesis falls; report manipulation checks as descriptive and let them withhold nothing. RS-20260829b-inversion-price §6, critic-response.md P3-1. | S232 | S232 |
| (bsu) | When you count whether a source rhyme can be answered in English, count what the rhyme is MADE OF — repetition crosses every language boundary, a grammatical ending crosses none, and pooling the two makes a class that behaves two ways. ARM-gulistan span C explained a nought — 0 end-rhymes in 44 bayts — by counting 19 of 44 as rhyming on a Persian grammatical ending, and filed radif (a word or phrase repeated after the qāfiya) in that same class. Span D counted its own 38 bayts before translating, registered the prediction, and the prediction half failed: the count was exactly right about rhyme (1 taken, in a lexical-qāfiya bayt, at no cost) and blind to chime, because every line-end chime the English took came from repetition — the radif «گو مباش» becomes say, let there be none and both couplets end identically in English exactly as in Persian, while the radifs made of «را» and «کن» carry nothing because English has no word there to repeat. A lexical radif is free; a grammatical radif is unanswerable; and a grammatical ending is unanswerable. Fires at: every census of could this rhyme be carried — the ghazal work, ARM-radif, ARM-radif-hands, §7.42 and §7.49 — and at any classification whose category is a position (the line end) rather than a material. Remedy: classify by material, not by position: split radif from ending, and record separately how often the target chimes by repeating a word rather than by finding a rhyme, because that column is the one a translator can actually act on. spanD-qafiya-precount.md; T-gulistan-bab2-R05-v1 D47, D48; register V18. | S240 | S240 |
| (bsv) | When you code whether a hand marks a foreign quotation, code the quotation's type AGAINST the type around it, not "italic / not italic" — a hand whose surroundings are already italic must mark by setting the quotation ROMAN, and an italic-only measure scores that as unmarked. E-20260903-latin-in-italian coded 21 Latin loci in «Vita Nova» across four English hands. Three of the four set Dante's divisioni in italic as a body and Martin extends it to the whole of chapter XXV, so 10 of 84 cells sit inside italic. On the italic measure the retained runs are marked 62/70 = 0.886 and §7.45.5's registered prediction fails on two of three marking hands; on the contrast measure they are marked 69/70 = 0.9857, and each of those two failures was the same single cell. Fires at: every marking census on a printed book — §7.45, §7.50, A-gulistan-hands, and any future bilingual source — and generally at any measure whose category is an absolute property of the marked span rather than a relation between the span and its context. Remedy: record the setting as <span type>-IN-<context type> and derive distinguished from it; report the direction of the contrast alongside the rate, because the direction is a fact about the volume's typography and the rate is a fact about the hand. The Persian material had no italic environment, which is why §7.45 could not see this and why the note is not a criticism of it. RS-20260903-latin-in-italian §2, §3; E-20260903-latin-in-italian/coding.md. | S241 | S241 |
| (bsw) | Read the price of every seat a design dispatches from the API, in the session that dispatches it — config/models.md is a cache and nothing re-reads it. On 2026-09-03 GET /api/v1/models returned $2.00 / $12.00 per M for openai/gpt-5.6-terra against the $1.00 / $6.00 the table had carried since S106: the first upward move this project has recorded, and the first that can breach a declared ceiling rather than merely over-price one. The three earlier corrections (S061, S106, S182) were all downward and all went unnoticed for weeks precisely because a downward error is invisible — the run comes in under. Fires at: every pre-flight estimate, not only the ones that dispatch a reserve seat, which is where the S182 gate stopped. config/models.md §2026-09-03; E-20260903b-line-end-order §11. Honoured S243, and this is what honouring it looks like when nothing has moved: all four seats re-read from /api/v1/models before the ceiling was set, all four unchanged from the S242 re-read, ceiling built on the read and not on the cache. A no-op reading is the note working, not the note idle. config/budget.md 2026-09-04. | S242 | S243 |
| (bsx) | Separate enjambment from inversion before counting line-end word order — they are confounded at the line end by construction. When a marked line's clause continues into the next line, its last word is not syntactically final, so a “would-it-be-last-in-prose” coding calls it displaced whether or not the hand displaced anything. In E-20260903b-line-end-order's fresh 60-item gold, all 21 enjambed items were coded INVERTED, against 21 of 39 among complete-clause items. Two hands who enjamb at different rates would then show an inversion-rate difference that is really an enjambment-rate difference. Fires at: any count of line-end word order (§7.47, §7.51, ARM-line-end-order and any successor) — add a clause-completeness field and report the primary on complete-clause items, or stratify. RS-20260903b-line-end-order §5. | S242 | S242 |
| (bsy) | Buy a blind second coder on the one dimension a quoted passage can settle — it is cheap, and the coder who wrote the rule is the coder who breaks it. Two pre-run critics made it BLOCKING that E-20260904 had one coder who was also the predictor, the locus classifier and one of the translators. The repair — 96 cells to one masked seat, hand identity hidden, no prediction and no framework shown, five codes — cost $0.119 and agreed at 92 of 96 (0.958). It also cost a registered prediction: one of the four disagreements flipped P4 and is a lead error, at a cell where the coding instruction the lead wrote and handed the seat verbatim lists Monsieur among the French forms that count as kept, and the lead had coded one hand's Monsieur Pierre as englished while coding another hand's M. Pierre as kept. Fires at: any run where one reader codes every cell of a comparison — which is most of this project's coding studies. Remedy: send the quoted cells out blind on the single dimension the quote settles (retention, presence, direction), register the disagreement gate and the flip rule before dispatch, and report a flipped prediction as unresolved rather than correcting the cell into a pass. It is not a jury and it does not check location — the seat sees the passage the lead extracted — so it checks classification only, and that is worth saying on the result page. RS-20260904-french-in-russian §6. | S243 | S243 |
API, cost and panel
| id | note | first seen | last fired |
|---|---|---|---|
| (bst) | A bigger max_tokens cap can make a seat's answer WORSE, on the same item at the same temperature — so a cap is a parameter of the answer, not only a pre-flight measurement. E-20260902's locating stage was probed on one poem. google/gemini-3.6-flash located 8 of 14 senses at cap 2500 and 4 of 14 at cap 6000, losing three it had already found and spending 5,833 completion tokens against 2,486 to do it. This is not truncation and note (bsf) does not cover it: (bsf) is about a cap too small for a seat's hidden reasoning, and here the seat had more room, used all of it, and got less right. The practical consequence is that a run cannot raise a cap "to be safe" — raising it changes the instrument, and a run that probes at one cap and dispatches at another has not probed its instrument. Fires at: any stage whose cap is chosen after a probe, which is every stage this project buys; sharpest on seats with unbounded hidden reasoning on a task that has a right answer to lose. Remedy: probe at least two caps on the same item and compare the answers, not just finish_reason; dispatch at the cap that was probed; and where two caps disagree, treat the seat as unusable on that shape rather than picking the reading you prefer. Here P2 was dropped from the shape entirely and P3 x-ai/grok-4.5 at low effort took it — 12 of 14, 1,659 tokens, $0.011 — while P1 truncated at 2500 for $0.031 and stayed on glossing. Kin to (bsf) and (bsq): all three say that a setting is an instrument choice and not a discount. RS-20260902-rhyme-slot §8.1; E-20260902-rhyme-slot/amendment-v2-1-probe.md. |
S239 | S239 |
| (bsq) | reasoning: {"effort": "low"} on x-ai/grok-4.5 is a thirteen-fold saving that buys a seat which is not looking — probe it against a KEY, never against its own full-effort answers. config/models.md has carried low effort on P3 since S234 as "worth probing", on the evidence that it agreed with its own full-effort answer at six of eight positions. E-20260901 is the first run to probe it against a key. Cost per 24-item batch fell from $0.094 to $0.0073, and the seat then failed both calibration gates: on 40 keyed items it called an INVERTED line CANONICAL 7 times of 19, against P1's 3 and P2's 0 — a one-sided miss, not noise. Its whole 523-item coding cost $0.277 against P1's $1.197 and was worth nothing, because it did not vote. Fires at: any stage that reaches for low effort to fit a ceiling. Remedy: a reasoning-effort setting is a different instrument, not a discount on the same one; probe it on keyed items from the run's own material and read the direction of its errors, because a seat that systematically under-detects the construct will still look plausible on agreement with another seat. Self-agreement is not validity. Corollary to (bne): an effort pin is a request to the provider, and here it was granted and cost the run a seat. RS-20260901-inversion-habit §4, §8. |
S237 | S237 |
| (bsh) | A body that returns finish_reason: length with content: null carries NO bit to re-parse, so note (brx) cannot fire — and a reasoning: {"effort": "low"} passthrough does not stop it. E-20260830's adjudication asked for a two-line answer at max_tokens 900; 8 of 115 openai/gpt-5.6-terra bodies spent all 900 completion tokens on hidden reasoning and returned nothing, at $0.011 each. The run was stopped, reasoning: {"effort": "low"} was added, and the provider billed 900 reasoning tokens anyway — the parameter had no effect on this model and provider, which is note (bne)'s corollary holding for a second seat. Re-buying the ten dead bodies at cap 3000 returned all ten clean, at $0.159115200. Fires at: any short-answer stage on a seat with unbounded hidden reasoning, which is now every frontier seat this project uses. Remedy: for a stage whose answer is under ~50 tokens, set the cap from the seat's reasoning habit and not from the answer — 3000 was ample here and, because billing is on tokens produced and not on the cap, a generous cap on a short answer costs nothing when the seat behaves and saves a whole re-dispatch when it does not. Do not price a short-answer stage as if the cap were the cost; price it from a probe. Kin to (bsf) (a cap probed on one item is not verified for the next) and (brx) (re-parse, don't re-buy), whose precondition this note names. config/budget.md S233. |
S233 | S233 |
| (bsi) | When a bought judgment requires the seat to SEARCH, split it: buy the retrieval and do the maximisation yourself. E-20260830b asked one call to gloss a ghazal's rhyme senses and then find the single English rhyme covering most of them. That second half is a search, and every seat probed spent thousands of reasoning tokens on it: openai/gpt-5.6-terra $0.030 a call and an EMPTY body at cap 2500, with reasoning: {"effort": "low"} making no difference; x-ai/grok-4.5 no return in 100 s, twice; google/gemini-3.6-flash $0.0195 clean only at cap 6000. The same seat, asked instead for up to eight ordinary English words that carry this sense — pure retrieval, no search — answered in 5.2 s at $0.0052, and the maximisation over CMUdict rime classes was then exhaustive, mechanical and free. The split was six times cheaper and a better measurement: it made the predictor an actual maximum over an enumerated inventory, which is what the pre-run critic's first BLOCKING finding had demanded and the design had overruled as unreachable. Fires at: any stage that asks a seat to optimise, rank a large space, or find a best-covering anything. Remedy: name the part of the task that only a language model can do — usually a gloss, a list, a judgement about one pair — buy exactly that, and put the combinatorics in the analyser where it can be verified. The tell that a stage needs this: a probe whose completion tokens are five to ten times its visible answer. RS-20260830b-rhyme-family §1; E-20260830b-rhyme-family/amendment-v3.md. |
S234 | S234 |
| (bsj) | A mismatch control built inside a genre with a stock lexicon can INVERT — it must be graded against an out-of-field arm, and the gate must sit on the harder one. E-20260830b built two grades of negative control: X-MIS, one Hafez ghazal's Persian against another Hafez ghazal's English, and X-SHUF, an English list assembled one word at a time from four different odes. The pre-run critic predicted X-MIS would be weak because Hafez's rhyme senses and a Victorian translator's rhyme vocabulary are both stock — heart, wine, dust, dawn-wind — and the design demoted it to descriptive before the run on that reasoning alone. Measured: X-MIS scored 0.1111 against the true items' 0.0384, i.e. higher, while X-SHUF scored 0.0000. Had X-MIS carried the gate, a working instrument would have been declared dead. Fires at: every design whose negative control is the same kind of thing, wrongly paired — the commonest shape of mismatch control there is. Remedy: build two grades, put the gate on the one that destroys the genre's own vocabulary, and report the in-field arm as a measurement of how much the genre alone supplies. Kin to (bnu) — a control domain defined as the absence of the treatment is not a control domain. RS-20260830b-rhyme-family §6. |
S234 | S234 |
| (brr) | openai/gpt-5.6-terra (P1) returns an EMPTY body at max_tokens 1500 on a 250-word prompt, and google/gemini-3.6-flash (P2) truncates a long critic response at 3500 AND AGAIN at 9000 — so the two cheapest seats both need prose caps sized well above what the visible answer costs. S223's nine calls include four re-dispatches, 54% of the session's spend, for two stages that asked for nothing but prose. P1 at 1500 gave finish_reason: length with no content and answered fully at 3500; both seats truncated the critic prompt at 3500, P1 completed at 9000 and P2 truncated again at 9000, so two of its findings are all that exists of it. Fires at: any stage that wants written prose from P1 or P2 — critics, operationalisation probes, free-text reasons. Remedy: budget a short prose stage at 3500 and a long-context critic at 9000 from the start, on both seats, and treat a P2 critic pass as possibly partial however high the cap: record what was lost rather than reporting the findings as the seat's whole response. NEXT.md's pre-spend list; family of (bph) and (bhf), and the fifth form of the empty-body failure. |
S223 | S223 |
| (brf) | A dispatcher that resumes by reading its own output file computes done ONCE, at start — so two of them running at the same time will re-buy nearly everything. E-20260824 launched stage D in the background, was told the wrapper had exited, and launched it again; the first process was still alive. 390 designed calls billed as 1,271, about $1.60 for nothing, inside the declared ceiling only because the ceiling was generous. Fires at: every run.py --stage X in this project — the resume-by-jsonl pattern is the house pattern and none of them holds a lock. Remedy, two lines: before dispatching, take an exclusive lock on a .lock file beside run.jsonl and exit if it is held; and never launch a dispatcher without first confirming no prior one is alive (pgrep -f run.py), because in this container a background job's exit notification does not mean the job it wrapped has finished. And check the row count against the design's call count before writing the ledger — the billed total is what catches this, and it caught it here only because someone looked. config/budget.md S217. |
S217 | S229 |
| (bqu) | A seat that renders its deliberation into the CONTENT truncates before its answer, and the failure looks like a dead cell rather than a cap that is too small — size the answer cap to the seat's habits, and make the parser find the answer line anywhere in the body, not only at the start of one. E-20260821c stage A opened at max_tokens: 100 with reasoning: {max_tokens: 80}, which is ample for a one-line gloss. qwen/qwen3.7-max (QR) writes “Thinking Process: 1. Analyze the Request…” into the message content, so it spent the cap on visible deliberation: 7 of the first 52 cells died, 13.5%, every one of them QR, against a 5% bar. Raising the cap to 350 was not enough on its own — QR writes its answer as * ENGLISH: lowest inside a bulleted list, which an anchored ^ENGLISH: regex misses, so bodies that CONTAINED the answer were being scored dead and re-bought. Unanchoring the regex then created ten false positives, because QR also restates the format template (ENGLISH: <translation>) inside its deliberation. The working combination is all three: a cap sized to the seat, an unanchored last-match parser, and a filter that rejects a value which is the template. Fires at: any dispatch whose reply format is a keyed line and whose panel includes a seat that narrates. Remedy: probe the seat's output shape before pricing the run; parse the last keyed match anywhere in the body; reject template echoes explicitly; and re-parse the stored raw bodies rather than re-buying them — this run recovered five cells that way at $0. Kin to (bpv) (the two caps are a ratio) and (bph) (truncation seen from the content end): those describe a cap too small for the answer, this one a cap consumed by something that is not the answer. RS-20260821c-echo-availability §7. |
S210 | S210 |
| (bqs) | A per-call worst case must add reasoning.max_tokens to max_tokens and price the SUM at the completion rate — a pre-flight built from the answer cap alone underprices a reasoning seat by about a factor of two. E-20260821b §10 priced 160 short-prompt bodies at ≈$0.62 worst case, correctly per note (abc) in every respect except this one: the requests carried max_tokens: 900 and reasoning: {max_tokens: 400}, and OpenRouter bills the hidden reasoning at the completion rate, so the true per-call cap was 1,300 completion tokens and not 900. Realised: $0.008685 per main body, against an assumed ≈$0.0039. The runner's stop-loss fired three cells short of the design's ninety, which is the guard working — the declared $1.20 ceiling was never approached and the daily budget was never at risk. Fires at: every pre-flight on a seat with a reasoning budget, which is every seat this project now uses. Remedy: worst case per call = (max_tokens + reasoning.max_tokens) × output price + prompt tokens × input price, per seat, summed — and note that (bgk) says even that sum is not always an upper bound. Kin to (abc) (build the worst case from the cap the request permits) and (bqk) (the stop-loss must sit above the worst case): (bqk) held only because the stop-loss had been set below the ceiling, so the arithmetic error cost three cells rather than a breach. RS-20260821b-matched-heard §7. |
S209 | S209 |
| (bqk) | A stop-loss must sit above a worst case that assumes EVERY call needs the doubled-cap re-dispatch, not a fraction of them — otherwise the guard fires part-way through a frozen design and leaves an arm of it short. E-20260820's pre-flight priced twelve calls plus "four worst-case re-dispatches" at an average seat rate, giving $0.17, and set the stop-loss at $0.24. The dearest seat truncated at the 1,500 cap on all four hands and re-dispatched at 3,000 every time; the run halted at $0.243000 with one of four hands holding a single seat, which under its own aggregation rule makes every cell of that hand INDETERMINATE. The stop-loss did its job; the estimate did not. Fires at: any serial dispatch whose seats differ in price by more than ~2× and whose reply format is long enough that the dear seat may truncate. Remedy: build the stop-loss from n_calls x (cap cost + 2x cap cost) at the dearest seat's rate, not at an average and not for an assumed subset; and where the guard does fire mid-design, revise it only with the arithmetic written into the runner and only for calls whose content is already frozen and whose outcome cannot be steered. RS-20260820-rhyme-bearer §7. |
S205 | S205 |
| (bps) | A reasoning-budget parameter is NOT portable across providers, and the fix for one seat can break another — probe it per seat, on the census's LONGEST item, before dispatch. E-20260816d set reasoning: {max_tokens: 150} on every call. qwen/qwen3.7-max writes its chain of thought into the content channel and hit the cap on 4 of its first 8 bodies; moonshotai/kimi-k3 ignores the budget entirely, spends whatever max_tokens it is given on reasoning and returns an empty string — 8 of its first 18, and still failing at a 557-token cap. Switching globally to reasoning: {effort: low} fixed kimi (500 → 138 tokens) and BROKE qwen, which had been passing 18 of 18. The seat was eventually dropped, at $0.174286 of discarded bodies, because a cap large enough for it does not fit any stop-loss the day's headroom allows — and dropping it forced an overlap between raters and hands that the pre-run critic had specifically made the design remove. Fires at: every multi-provider run, and hardest at the expensive seat, where the wasted cap costs the most. Remedy, three calls and under $0.05: before dispatch, send the design's own prompt at the design's own cap to each seat on the longest census item, under each budget variant you are considering, and read finish_reason and completion_tokens. Set the budget parameter per seat, never globally. Kin to (bph), which is the same defect seen from the truncation end; this note is what makes (bph) checkable in advance. RS-20260816d-lexical-channel §8. |
S198 | S199 — FIRED AGAIN, AND HALF-HONOURED IS NOT HONOURED. E-20260816e probed both rater seats before dispatch, as this note requires, but on the longest classification item rather than on the census's longest item — a third of the length. z-ai/glm-5.2 then returned finish_reason: length on the census at caps of 3,000 and 6,000, twice, and was replaced mid-run at $0.017566 of discarded bodies. The note says the census's longest item and it means it: probe the longest prompt the run will actually send, not the longest of the short ones. |
| (bne) | deepseek/deepseek-v4-pro is not usable as a judging or coding seat on ANY task shape, and effort: low does not bound its hidden reasoning. NEXT.md has carried a standing finding against this seat since S174 — five dead bodies of forty-two — scoped to structured-output judging. E-20260813e gave it a plain-text coding task, one line per site, with reasoning: {effort: "low"} set, and it returned content: null with 21,497 characters of reasoning, finish_reason: length, the whole 6,000-token cap consumed before any answer began, provider DigitalOcean. $0.018416160 for nothing, on the first call, which is why the design's second coder had to be replaced mid-run. Fires at: any design that budgets this slug as a cheap volume seat. Remedy: do not seat it for judging, coding or any task whose value is the answer rather than the prose; the practical panel for those is P1, P2, P3. Where a second independent coder is wanted and every remaining seat also produced a subject in the same run, cross-assign — each seat codes every hand but its own — which is what this run did and which costs nothing extra. Corollary against (b)/(bmb): an effort pin is a request to the provider, not a guarantee, and a design must not treat it as a cost bound. RS-20260813e-slot-typology-ja §8. |
S176 | S176 |
| (bmz) | finish_reason is not a truncation detector — a provider will return "stop" on a body it cut off at the cap, so a re-dispatch guard keyed on "length" silently loses the call. yp__S03 came back with finish_reason: "stop" and 796 completion tokens against an 800-token cap, 640 of them hidden reasoning, ending mid-word inside its JSON. It did not parse, the guard never fired, the segment lost its yardstick, and that one segment is the margin by which E-20260813b's manipulation check P3 missed its second clause — 10 of 15 against a bar of 11 — so F3 fired and the run's primary was VOID. One mislabelled body voided the primary of a 329-body run. Fires at: every dispatcher that decides whether a body is complete from provider metadata. Remedy, both halves: (i) treat a body as truncated when it fails to parse into the shape the prompt demanded, whatever finish_reason says, and re-dispatch on that; finish_reason == "length" is a hint, not the test. (ii) The same code path leaks money: dispatch() writes the successful attempt's cost and overwrites the discarded one, so 48 re-dispatches left $0.293451, 19.8% of billed spend, unattributed — accumulate cost across attempts instead of assigning it. Kin to (bgk): a worst case built from max_tokens is not a guarantee, and neither is a completion signal built from the provider's own word. RS-20260813b-affect-yardstick §§6.2, 8. |
S173 | S174 — BOTH HALVES APPLIED, AND THE MONEY HALF IS DISCHARGED. E-20260813c's runner re-dispatches on parse failure irrespective of finish_reason (12 of 54 dispatches were re-dispatches, and 7 of those recovered a body that finish_reason alone would have kept or lost wrongly), and it accumulates cost across attempts. Result: sum of parts $0.398367110 against a key-usage delta of $0.398367107 — an unattributed residual of 3 × 10⁻⁹, against 19.8% on the run that raised the note. The truncation half remains live: it prevents silent loss, it does not make an unparseable seat parse. RS-20260813c-ennoblement-direction §9. |
| (bmb) | [FIRED S162 in a NEW MODE, and every existing guard missed it: TRUNCATED content, not NULL content. 14 of 210 bodies on google/gemini-3.6-flash returned finish_reason: length with a JSON object cut off mid-string — {"answer": "A", "cue": "spir — and the licensed re-dispatch never fired, because the caller tested if content: and truncated content is not empty. $0.058131. The signature this project has learned to catch is an empty body; the cheap general fix is to treat finish_reason == "length" as the trigger rather than emptiness. A post hoc recovery found the answer field intact in all 14 and moved no figure by more than 0.033, and none could enter the primary in any case because each lost the cue the cue-attributed reading needs. RS-20260811h §7.1.] A model that expands hidden reasoning to fill its cap cannot be repaired by raising the cap, and the standard re-dispatch reflex makes it worse. E-20260811-floor gave google/gemini-3.6-flash a 1,400-token cap; it spent 1,345 on hidden reasoning and truncated the answer. The cap was raised to 2,600 — note (abc)'s own remedy — and it spent 2,494, truncating again and costing 1.7× more per body. Four blocks × two rounds = eight dead bodies, $0.150 of a $0.329 run, 46% waste. The slug also rejects reasoning: {"enabled": false} with HTTP 400, so the project's standing NO_REASONING switch does not reach it; reasoning: {"effort": "low"} does, and fixed it at the third attempt with reasoning at 1,057. Rule: when a body truncates and completion_tokens_details.reasoning_tokens is within ~5% of the cap, the cap is not the constraint — cap the REASONING, and if the slug refuses enabled: false, try effort: low before spending a second round. Corollary for pre-flight (note (abc)): a reasoning slug's worst case is max_tokens, and raising max_tokens raises the worst case and the expected case together, which is not true of a non-reasoning slug. First occurrence on this slug; the family is (bkw)'s. RS-20260811-floor §9.8. [FIRED S158, on TWO FURTHER SLUGS, and the remedy is now two things and not one. E-20260811c lost 51 bodies, $0.525140707, 28.4% of a $1.85 run: deepseek/deepseek-v4-pro spent 2,999 of a 3,000 cap on hidden reasoning with null content, and moonshotai/kimi-k3 killed 11 of 18 parity bodies the same way at 2,000. Raising the cap alone did not fix either — the note's own point — and reasoning: {"effort": "low"} alone was not tried in isolation. What worked on both new slugs was the PAIR: pin the effort AND raise the cap, deepseek clean at 4,000 with effort pinned, kimi clean at 4,000 with effort pinned. So the rule generalises past the one slug: on any reasoning-capable seat, pin the effort in the FIRST dispatch and size the cap from note (bhq), not after a dead round. Three slugs now, three labs.] [FIRED S168, and the failure was reading past this note rather than a new mode. E-20260812f's design cited (abc) and (bmp) by name and set the judging cap at 300 tokens with no effort pinned on moonshotai/kimi-k3 and deepseek/deepseek-v4-pro — the two slugs this note had already named. 26 of 30 and 20 of 30 bodies came back empty at finish_reason: length, the whole cap spent on hidden reasoning. The repair: kimi dropped from the judging seats, because the 4,000-token cap this note prescribes prices thirty of its calls at $1.80 and would have broken the run's declared ceiling; x-ai/grok-4.5 replaced it; deepseek pinned at effort low with a 4,000 cap then returned 30 of 30. Cost of not applying an existing note: 61 dead bodies and 47.9% of a $1.42 run. Corollary that is new: a seat this note cannot afford to repair is a seat to replace, and the replacement is chosen from slugs that have already returned clean bodies in the same run. RS-20260812f-affect-unprompted §7.] |
S156 | S174 — the paired remedy was applied in the FIRST dispatch and was still not enough, and this time it cost a run. E-20260813c pinned deepseek/deepseek-v4-pro at effort low from the first call and sized the cap at 1,500 with a doubling re-dispatch. The seat still returned finish_reason: length with unparseable bodies on 5 of its 14 sites, at a 6,000-token cap. Those five deaths are 11.90% of 42 bodies and fired F3, voiding every primary of a run whose other gates all passed. The rule this adds: on a task requiring STRUCTURED output, a seat with this failure history is not made safe by pinning and sizing — budget for it to fail and register the run as two-seat, or replace it before dispatch. RS-20260813c-ennoblement-direction §5. |
| (blj) | A skip-if-exists guard is not a lock, and a stage the harness cannot see cannot be stopped. S146's retrieval probe was first launched with nohup … & inside a tool call, which returned immediately and reported success while the python process kept running untracked. The stage was relaunched in a tracked call; the first process was still alive, and the two walked the same job list with only a file already exists check between them. Both dispatched the same seat's probe; the body that lost the write race was billed and never stored, surfacing as a $0.010518579 gap, 1.03%, between the key-usage delta and the per-response sum. Note (blf) says dispatch in the background rather than a killable foreground shell — this is its converse and both are the same defect of not knowing what is running. Rule: launch every stage so the harness owns the process; never nohup … & inside a tool call; and make a resume guard a lock file taken before the request, not an output-file check made after it. RS-20260809h-rule-execution §8. |
S146 | S146 |
| (bhd) | A re-dispatch must mint a fresh tag: a retry that reuses its label overwrites the preserved raw body of the attempt it replaces. S080's runner filed a higher-limit retry under the same name as the failed attempt; the save-everything-before-reading safeguard was silently defeated, and only the spend reconciliation (short by exactly the discarded call's $0.0642285) noticed. Every runner that can retry derives the filename from attempt count or timestamp, never from the request id alone — and the verifier's call-count check is the backstop that catches it when forgotten. | S082 (from S080's defect) | S156 — nine bodies rotated into runs/discarded/ across two re-dispatch rounds, none overwritten |
| (bgk) | max_tokens is not always an upper bound on billed completion tokens, so a worst case built from it can be exceeded through no error of arithmetic. x-ai/grok-4.5 returned 4,231 completion tokens against a declared cap of 2,200 on one cell of E-20260731h; OpenRouter did not enforce the cap for that slug. Note (abc) says build the worst case from the cap the request actually permits, and that is still right — but the cap is a request, not a guarantee, and a reasoning-model slug may bill past it. Where a run's headroom is tight, price the reasoning-inflated figure (S063's 3× ratio) rather than the cap-literal one, and check completion tokens against the cap after the fact. Companion to (abc) and (b). Fired again at S086 on a different model and a different provider — deepseek/deepseek-v4-pro via SiliconFlow, 12,237 completion tokens (12,069 reasoning) against a 10,000 cap, finish_reason: stop, billing $0.0414 against a declared per-call worst case of $0.038775. One call in 96, and the only one of ten providers routed for that seat to do it. Two firings on two vendors makes this a property of the market, not of a slug. The after-the-fact check is now enforced rather than advised: E-20260801f's verifier asserts the breach is exactly one call, on that provider, at that token count (checks 20), because a stage-level reservation tripwire cannot see a single call overrun inside it. |
S076 | S185 — third firing, third vendor family, and this one cost nothing because the mean stayed under the cap. x-ai/grok-4.5 returned 2,167 completion tokens against a declared 800-token cap on one cell of E-20260814g, finish_reason: stop, body clean and parsable. The seat's mean was 702, so the run came in at $0.420799 against a cap-literal worst case of $0.642 and nothing broke — but the per-body worst case built from the cap was exceeded 2.7×, exactly as this note says it can be. The after-the-fact check is what saw it; no tripwire fired, because none can see a single call overrun inside a stage reservation. |
| (bid) | A killed client does not cancel a billed request, and the per-request re-sum cannot see the orphan — only the key-usage delta can. A 120-second shell timeout killed a dispatch to openai/gpt-5.6-terra while it was in flight. The request was still billed; its body was never received; it appears in no usage.cost field, in no .err file and in no stored raw body. The session's re-sum read $0.077852200 and the key-usage delta read $0.097558200 — a residual of +$0.019706, which is a whole call at that seat's rates. Rule: when a dispatch is interrupted by anything other than the API returning (a timeout, a kill, a lost connection), record it as a dispatch. The key-usage delta is the ledger's actual for that session and the re-sum is the number that is wrong — the reverse of this project's usual ordering, and the case the cross-check exists to catch. Corollary: run long dispatches with a client timeout shorter than the harness's, or in the background, so that the client — not the shell — decides when a call is over. |
S100 | S100 · S102, and the corollary was the whole cause: a 110-second shell timeout was set around a runner whose own socket timeout is 300 s. One stage-0 dispatch was killed in flight, billed, and left no usage.cost, no .err and no raw body. Session residual $0.020428749 against a per-request re-sum, of which this call and one IncompleteRead are the whole of it. The runner carried the fix; the operator wrapped it in a shorter timeout anyway. S138 — third firing, same cause, same seat family. A 2-minute shell timeout killed the translate stage mid-gpt-5.6-terra; per-request sum $0.156794026, key delta $0.168686275, residual $0.011892249 — the size of a call of that shape at that seat. Ledgered at the key delta and named on the result page. The corollary is still the fix and was still not applied: run the dispatch in the background, or give the client a timeout shorter than the harness's. |
| (bgl) | Note (b)'s remedy — raise the cap — WORKED, once, and the same seat then failed at the raised cap. deepseek/deepseek-v4-pro returned finish_reason: length with 2,200 tokens and zero characters of content on three cells; the cap was raised to 8,000 and eight of nine cells returned clean at 5,072–7,427 completion tokens. The ninth burned all 8,000 and returned nothing. So the cap is a real constraint and raising it is a real fix, and it is not a reliable one — which is the first evidence this project has that separates the cap was too low from this seat sometimes returns nothing whatever the cap is. Twenty-first and twenty-second firings of (b), and the first on a translation SUBJECT rather than a critic or rater. |
S076 | S076 |
| (bgp) | A line count is not a format check, and note (bfb)'s reasoning fallback must be put through the same test as the content it replaces. E-20260801's runner accepted a body if message.reasoning had at least n lines. google/gemini-3.6-flash, truncated at length, returned a prose thinking-summary — "I'm now confirming that Rule F6 indeed bears on Q04…" — with plenty of lines and not one answer in the answer grammar, and the runner wrote it to disk as the seat's answers. Acceptance must count lines matching the response format, and the (bfb) fallback must pass the identical test; otherwise the fallback that exists to rescue a good answer from the wrong field launders a bad one. |
S077 | S077 |
| (bgq) | A runner that snapshots key usage inside its first stage overwrites the opening snapshot every time that stage is re-run. E-20260801's run.py critic takes session-open; the stage crashed on note (bgc)'s whitespace body and was re-run, and the second write replaced the first. Here both reads happened to be identical to the digit so nothing was lost, which is exactly why it would go unnoticed the time it matters. Snapshot with a write-once guard, or name the file by attempt. |
S077 | S077 |
| (bfh) | A DECLARED reserve is not a WORKING reserve, and note (bfc) protects against the wrong thing. (bfc) requires every dispatch stage to name a reserve before it runs; E-20260730i did, and the reserve then failed twice in the same way the primary had — google/gemini-3.6-flash truncated at max_tokens 4,000 and again at 10,000, spending 8,213 characters on reasoning and returning 27 and 67 of 69 lines, after deepseek-v4-pro had returned an empty body at 4,000. $0.15167092 wasted, 31% of the session's spend, on one seat that was never filled. The operative rule: a reserve must differ from the primary in the failure mode it is exposed to, not merely in the slug — and a stage that has burned three dispatches is abandoned and declared, not bought a fourth time. |
S068 | S068 |
| (bfo) | A hung call that you abandon still bills at its max_tokens cap — abandoning is a client-side act and the server finishes without you. moonshotai/kimi-k3 held the socket open past twenty minutes on E-20260731's critic seat and was abandoned with no response body, no usage row and nothing written to disk. The session's key-usage delta came back at $0.309273057 against a per-request sum of $0.1154717578 — an excess of $0.1938, against a pre-flight worst case for that seat of $0.191. The abandoned call billed at essentially its full cap and left no per-request record to ledger it against. Two consequences. (i) Note (b)'s wasted-call accounting must include abandoned hangs, which until now were assumed free because nothing came back. (ii) A session whose per-request sum is the only cost record will understate its spend by the size of any hang; the key delta is the figure to ledger, and this is the (bet) excess direction with a cause. Companion to (beh): a fall-through chain does nothing against a hang, and now the hang is known to cost the cap. AMENDED 2026-07-31 (S070): not always. S070 abandoned a pre-run critic call at 120 s of a generation that took 144 s at the same seat the day before, and the session's key-usage delta equalled the per-request sum of its four accepted calls to 1e-9 — the abandoned call added zero. What separates the two cases is not established. The note's own case, a twenty-minute hang billing ~$0.1938, stands; the generalisation to every abandoned call does not. |
S069 | S070 |
| (bet) | A per-stage upper bound on the key-usage delta is not a valid check; only the session-level bound is. E-20260730d's verifier failed one check of 208 on its first run, and the failure was the check rather than the run: the repeat stage's delta (0.10137) exceeded its own per-request sum (0.06696), because the census stage's outstanding billing settled inside the repeat window — the census delta was short of its sum by 0.10138, the repeat window's entire delta to five decimal places. Note (bco) had only ever been seen in the shortfall direction. At session level the arithmetic is exact: opening 29.516047048, closing 30.068816756, delta 0.552769708 against a sum of 0.5527697100, residual −2e-9. Assert monotonicity per stage and the bound at session level. |
S063 | S063 |
| (beq) | A response can carry no payload at all — only SSE keep-alive padding — and that is a transport failure, not note (b)'s failure mode. qwen/qwen3.7-max at Alibaba returned 1,892 bytes of blank and :-comment lines and no JSON mid-run, and the harness died on a JSONDecodeError rather than recording it. Note (b) says never retry a slug that failed and fall through to a declared reserve; that is right for a model returning length or an empty completion, and wrong here — nothing was served, and the seat was the materials enumerator, which cannot be replaced without putting enumerator identity inside the study's own statistic. The request was re-issued to the same slug and accepted first time; the failed call shows no billing in the delta. call.py now raises a PayloadFree exception so the case is logged instead of fatal. Distinguish "the model answered badly" from "nothing came back", and let only the first trigger a fall-through. |
S063 | S063 |
| (beo) | A cross-session comparison must resolve every cell to the slug that actually answered it, because a declared reserve makes "the same seat" a different model. Note (bdt) said a verifier reading <tag>.raw reads the rejected body wherever a reserve fired; at S062 that fired on real data and changed a number. S057's P5-M reader cell had fallen through to the declared reserve, so its answer of record came from google/gemini-3.6-flash and not from deepseek/deepseek-v4-pro; comparing it against today's P5 would have been a between-model difference reported as cross-day drift, and it was the largest of the six cells. Found by this session's verifier, not by its analysis. Corollary caught at the same time: a cell can keep its slug and change provider (GMICloud → SiliconFlow here), which the S022 routing caution says is not the same served model either. |
S062 | S062 |
| (beh) | A fall-through chain protects against a call that returns a failure and does nothing against a call that hangs. In E-20260730-grain-clause's dispatch, deepseek/deepseek-v4-pro held its first socket open past seven minutes twice without returning; call.py's reserve chain never fired, because there was nothing to fall through from, and the whole 21-call run sat behind one optional third reader. Two remedies, both cheap: order a multi-seat dispatch so every call the primary statistic depends on completes before any optional seat is attempted, and treat an abandoned socket as a possibly-billed call — declare it and check the key-usage delta rather than assuming an unreturned call is free. Companion to (b) and (bdl), which are both about returns. |
S061 | S061 |
| (m) | State cost estimates as a range whose worst case is built from per-call maxima — and see (abc) for what "maxima" has to mean. | S020 | S031 |
| (abc) | [fired S047 — and the band moved for a structural reason: 47%, against 15–31% in the four previous sessions, because this run's cost is dominated by input (prefixes of 896 to 11,661 words), so a max_tokens worst case is a much tighter bound. The note works; the band was never the target.] A worst case built from an assumed output length is not a worst case; build it from the cap the request actually allows. S031's critic pass billed $0.091925 against a stated worst case of $0.090 — the first run in this ledger to land outside its estimate — because the estimate assumed output ≤4,500 tokens while max_tokens permitted 7,000. S037 sharpens what "the cap allows" means: the cap bounds reasoning tokens as well as visible output, so a request can consume its entire budget and return nothing. See (b), which already said so and was not applied. |
S031 | S062 — fired in a new place: an amendment that ADDS a call invalidates the pre-flight fraction, and both figures (35.1% with the added call, 31.2% without) are stated · S063 — declared the reasoning-inflated worst case (3× the cap, the S062-observed ratio) rather than the cap-literal one for the enumeration stage, and stated both; the run then landed at 43.8% of the amended figure, ABOVE the 15–34% band, because the census calls were prompt-dominated rather than output-dominated · S067 · S071 — the declared worst case was RAISED in session from $0.25 to $0.30 with the reason written (max_tokens 3,000 -> 4,000 against note (b)), and the actual came in at 48% · S075 · S077 — worst case $1.10 built from max_tokens, actual $0.580589442 = 53%, not raised · S079 — APPLIED AND STILL NOT ENOUGH. The gate's worst case was built from max_tokens per seat and overran by 3%, because the caller can dispatch one logical seat 3 attempts x 2 slugs. A worst case is max_tokens x attempts x slugs. Attempts capped at 2 for the same session's main design, and the retry structure written into its estimate, which then came in at 49% of declared · S081 — declared worst case $0.57 from max_tokens, actual $0.187826153, 33%. The estimate held because the attempt cap was in the arithmetic: S079's overrun came from pricing one dispatch per seat when the caller can dispatch several · S156 — the CRITIC prompt's own cap truncated at 6,000; and see (bmb) on why raising a reasoning slug's cap is not the fix |
| (b) | Reasoning models spend 1.7k–11k tokens before output; size max_tokens for reasoning plus output. Fired against S037, which did not apply it: a 60-item coding call capped at max_tokens: 2000 was consumed entirely by unreturned reasoning on both panel roles, returning finish_reason: length and no usable content, at a wasted $0.057517. The note was seventeen sessions old and correct. Where a request only needs a short structured answer, send reasoning: {"effort": "low"} as well as a generous cap — that is what made the re-run cost less than the failure. Fired again S044, on the fifth session to be bitten: a moonshotai/kimi-k3 ratification vote capped at max_tokens: 6000 consumed the entire budget on unreturned reasoning, returned finish_reason: length and no content, at a wasted $0.130665, the single largest waste in this ledger. The note was correct and was not applied. tools/ratify_vote.py now takes --reasoning-effort, so the remedy is in the tool rather than in the operator's memory — which is the only form of this note that has ever worked. Fired again S045, sixth session, on deepseek/deepseek-v4-pro at max_tokens 6000: finish_reason: length, content: null, $0.019293 for nothing. Two sessions of the same failure on two different labs is enough to state the rule the note has been circling: a reasoning-heavy model given a long prompt and a max_tokens in the low thousands returns nothing at all, reasoning.effort is honoured by some providers and silently ignored by others, and the remedy that has actually worked twice running is to fall through to the declared first reserve rather than to retry. FIRED AGAIN S093, and this is the note's costliest firing: the ratification review was dispatched to moonshotai/kimi-k3 at max_tokens 14,000 with NO reasoning cap, and returned finish_reason: length, 14,000 completion tokens, ZERO content characters, $0.221532 — 82% of that gate's spend, buying nothing. The retry with --reasoning-effort low returned a complete review in 50.9 s for $0.047. This is the SAME MODEL the note was written about at S044. Amendment, and it is the only new rule of conduct: on a seat with a recorded history of this failure, the reasoning cap is not advisory. --reasoning-effort low is set on the first dispatch, not the retry. |
S010 | S049 (seventh session, and the first where it cost a materials build rather than a call: deepseek/deepseek-v4-pro at max_tokens 12,000 with reasoning: {"effort":"low"} returned finish_reason: length and no content, $0.039651 for nothing, and the declared fall-through worked again — third time running) · S054 (eighth session; deepseek/deepseek-v4-pro at max_tokens 12,000 with effort: low, $0.0498 for nothing — and this time the fall-through was NOT the remedy: see (bdl)) · S066 (FIFTEENTH and SIXTEENTH firings, both deepseek/deepseek-v4-pro at max_tokens 8,000, length with an empty body, at TWO different providers in one session — GMICloud $0.0145573 and StreamLake $0.0113267. Both fell through and both fall-throughs were accepted first call. One of the two stages had no declared reserve: note (bfc)) · S065 (thirteenth session, and the most expensive firing in this ledger: moonshotai/kimi-k3 at max_tokens 16,000, provider Fireworks, finish_reason: length with content: null, $0.4149495 for nothing — 41% of a session's whole spend. The declared reserve deepseek/deepseek-v4-pro was accepted first call at $0.0233. The rejected body was preserved and not overwritten, which is the S045 runner defect not repeated. The operative lesson is the one the note already carries and nobody applies: a bigger cap is not a safer cap on a reasoning model, it is a bigger bill for nothing.) · S067 · S068, eighteenth firing: deepseek-v4-pro at max_tokens 4,000, empty body, 15,897 characters of unreturned reasoning, provider AtlasCloud · S073, nineteenth · S074, TWENTIETH — and the first time a cap RAISED because of this note was still not enough. moonshotai/kimi-k3 as pre-run critic at max_tokens 12,000, chosen at 12,000 precisely because S073 lost the same seat at 8,000: finish_reason: length, 11,997 reasoning tokens, ZERO characters of content, $0.195747 for nothing. The declared reserve returned a six-finding verdict for $0.030372. On this seat the cap is not the variable to tune. · S077 — TWENTY-THIRD firing, and the first time a whole SEAT was lost to it: google/gemini-3.6-flash returned length with a truncated prose summary on ALL THREE batches at cap 8,000, $0.186061500 for nothing; raised to 20,000 with effort: low it answered all three first call at a third of the wasted cost · S078 — three length truncations on one seat, and a FOURTH body that reported stop while being truncated; see (bgt) · S081 — three firings in one nine-seat run, all deepseek-v4-pro, all finish_reason: length at caps of 3,000 and 1,600 on a plain translation task. Two of them exhausted a logical seat's two attempts and the design's own §9.3 (fewer than 2 seats -> NO RESULT for that passage) absorbed it without an amendment: pair K's Floor rests on two seats and says so. $0.013038552 wasted, 7% of the session · S102 — TWO seats in one stage, on two labs and three providers, all four dispatches lost: deepseek/deepseek-v4-pro (SiliconFlow, GMICloud) and qwen/qwen3.7-max (Alibaba) each burned an ENTIRE 3,000-token allowance on reasoning and returned finish_reason: length with zero content, WITH effort: low set on the first dispatch as the note's own amendment requires. $0.0289 wasted. Re-run at 8,000 both answered first attempt — so (bgl)'s reading holds and the amendment's effort: low did not bound the trace. The operative addition: size the cap for the reasoning trace, and note that a stage whose answer is 7 short lines is exactly where a too-small cap looks safe.** |
| (bdl) | One slug is not one instrument: OpenRouter routed deepseek/deepseek-v4-pro to three providers inside one session, with three different behaviours. S054 sent the same model four times. Cloudflare burned 10,755 reasoning tokens against max_tokens: 12000, returned finish_reason: length with an empty body and billed $0.0498 for nothing; GMICloud, on the sibling call with the same cap and the same reasoning: {"effort":"low"}, returned 10,224 tokens cleanly; Together, on the re-run, returned 17,303 tokens cleanly at max_tokens: 32000. What worked was raising the cap, not falling through to the reserve — which is the opposite of the remedy note (b) had settled on after S044/S045/S049, and the reason is visible in the numbers: the failure was not the model refusing to answer, it was one provider spending the budget before answering. So (x)'s "list price is not billed price" has a companion: the provider decides whether the call returns at all, and the provider field must be read off every failure as well as every cost. |
S054 | S054 · S067 |
| (bdt) | A verifier that opens <tag>.raw reads the REJECTED body wherever the declared reserve chain fired. call.py writes the first attempt to <tag>.raw and each fall-through to <tag>-reserveN.raw. Note (b) fired on P5's M-axis call in S057; verify.py read read-P5-M.raw — 12,000 tokens of nothing — and scored the run against it. It was caught only because an empty body crashed the script on a missing key; a merely truncated rejected body would have verified silently against the wrong text, and would have done so while reporting a clean check count. The bug is latent in every verifier this project has frozen whose reserves happened not to fire. A verifier must resolve the accepted body by walking the reserve chain and asserting finish_reason == "stop", and it should also assert that the rejected attempt is still on disk and still billed. |
S057 | S062 — fired on real data for the first time and changed a number; see (beo) · S067 · S071 · S075 — swept across the whole archive and found NOT LIVE: 5 of 5 mutations of a rejected body are silent, no frozen verifier reads one. The static signature it prescribes over-calls: two verifiers with no chain mention are correct anyway. |
| (v) | Instruct brevity explicitly on reasoning models. | S021 | S229 |
| (w) | The blind second-reader shape works and costs ~2¢ a source. | S022 | S022 |
| (x) | List price is not billed price — though P1 has now billed exactly list five times running, provider OpenAI. |
S015 | S050 (P3 x-ai/grok-4.5, provider xAI, billed $0.0197964 against $0.020014 at list — list honoured, and the first time this note has been checked on P3 at all) · S063 — a failing instance on one call of fourteen: census-F-P1 routed to Azure rather than OpenAI and billed $0.0216525 for 980 output tokens, the highest per-output-token cost in the batch; the other thirteen honoured list |
| (c) | moonshotai/kimi-k3 is the priciest juror. |
S010 | S020 |
| (zz) | A snapshot taken immediately after a failed call can under-report it; take the cross-check at session end. | S030 | S030 |
| (abf) | A key-usage delta can appear in a session that made no call at all. Snapshot at session start as well as end, ledger the difference against the current day rather than writing it off, and do not treat a single occurrence as a pattern. | S033 | S049 (opening snapshot $0.060274 above S048's close, of which S048's own $0.051713 settling late accounts for 86%) · S156 — key delta lagged by exactly one body, $0.012654000 |
| (bcx) | When two sessions run concurrently on one key, a usage snapshot cannot be attributed to either — and a delta of exactly zero is a dead instrument, not a passing cross-check. S048's delta was 0.000000 against a per-request cost of $0.051713, on three snapshots including one taken later in the session; the endpoint had not moved. S048 also read its opening snapshot as $0.664526 of unexplained non-project drift and said so in its own budget entry — and that was wrong: the money was concurrent S047, whose settled closing figure is S048's opening figure to the digit. The attribution was only discoverable because the two sessions collided at merge. Read a zero delta as "no signal"; and never attribute a between-session gap without checking whether another session of this project was running. | S048 | S048 · S081 — key delta $0.157485256 against a per-request sum of $0.187826153, the ordinary settling-lag direction; per-request sums ledgered |
| (bfa) | A void criterion must not conflate "the instrument failed" with "the phenomenon is real". E-20260730g's F1 voided condition B when the 1:1 paragraph alignment covered 51.4% of the source against a registered 90%. Two entirely different states produce that number: the aligner drifted, and the translator did not preserve paragraph boundaries — and only the first is a reason to void. Both were in fact present: the arithmetic was consistent (69 + 7 − 1 = 75, i.e. seven genuine splits) and hand-verification found the aligner wrong at three of eight non-1:1 blocks. A coverage threshold cannot tell them apart, so a design that voids on coverage must also register a check that separates them — here, hand-verifying every non-1:1 block, which the design did require and which is what made the verdict readable instead of merely negative. | S066 | S066 · S071 — its own repair registered as F1a/F1b and DISCHARGED: the coverage threshold was unreachable by construction (a perfect alignment reaches 0.568) while the instrument failed independently, and splitting the criterion is what made both readable |
| (bfb) | (fired again S074 — the same cell that had just produced note (bgc)'s whitespace body returned this shape on its retry; a plain third dispatch to the identical slug returned complete output.) A third dispatch failure shape: finish_reason: stop, content: null, and the whole answer — one line of thirty-seven — inside message.reasoning. x-ai/grok-4.5 billed 466 completion tokens, did not truncate, and did not return nothing. Note (b) is about length; note (beq) is about a payload-free socket; neither covers this. The remedy is neither fall-through nor blind retry: dropping reasoning: {"effort": "low"} and re-issuing to the same slug returned a complete, correctly formatted answer at 429 tokens. That is note (bdl)'s shape — change the parameter, do not change the seat — applied to a failure note (bdl) did not describe. Read message.reasoning on every null-content body before deciding what failed. | S066 | S066 · S077 — its fallback LAUNDERED a bad body; see (bgp) |
| (bfc) | Declare a reserve for EVERY dispatch stage, not for the critic alone, or a stage that fails makes you choose a seat after seeing the failure. E-20260730g declared a reserve for its critic and for one reader seat. Note (b) then fired on a third stage that had none, and the fall-through seat was picked after the failure — which is the one thing a declared reserve exists to prevent, because the choice can be made to suit the result. The seat chosen (P1, whose other task could not leak into this one) was defensible and the point stands: defensible-after-the-fact is not the same as declared-in-advance, and the difference is not visible in the output. Cost of the fix: one line per stage in the design. | S066 | S158 — and in the place the note does not yet name: the stage that overran was a GATE. E-20260811c declared $1.80, spent $1.846625863, and the $0.047 overrun is exactly the parity re-dispatch. A failed judge stage can be scaled down; a failed PARITY control cannot, because a run with no parity data is a run about damage. A gate's reserve is not discretionary and must be priced at the full cost of dispatching it twice. |
| (bfe) | A FOURTH dispatch failure shape — HTTP 504 "The operation was aborted", no body, no usage object — and it is NOT free. moonshotai/kimi-k3 at max_tokens 12,000 returned it twice; max_tokens 8,000 with reasoning: {"effort":"low"} returned a complete verdict from the same slug at the same provider. That is note (bdl)'s remedy — change the parameter, not the seat — on a failure neither (b) (truncation), (beq) (payload-free socket) nor (bfb) (answer inside message.reasoning) describes. The expensive half is the billing. Key usage was identical to the digit immediately after the second 504, and by session close the delta EXCEEDED the per-request sum by $0.0437483996 — the first excess this ledger has recorded, where every previous cross-check ran short (note (bco)). S061 built exactly this check — assert the delta as positive and not exceeding the sum, which would also catch a billed call with no stored body — and recorded that no excess appeared. It appears now. So: an error response is a lower bound of zero, not a cost of zero, and the per-request sum is a lower bound whenever a dispatch returned no usage object. | S067 | S067** |
| (bgb) | A recognition YES is not evidence that recognition occurred, and on non-canonical material it is usually the tradition rather than the work. E-20260731f asked both raters to identify a lead translation of Kielland's «Karen» (1882, Norwegian, no canonical English). Three YES answers, all three wrong — twice "The Poet, Isak Dinesen", once "Himmerland Stories by Johannes V. Jensen" — both Danish, both roughly the right literary neighbourhood, neither the author. A fourth, "Karen by Johannes V. Jensen", looks like a near-miss and is not: the name Karen appears eleven times in the passage, so the title is readable off the page. RS-20260730d §5 established that stripping authorship from a lead translation of a canonical work leaves the raters knowing the work; this is the other failure mode, and a design that scores recognition as a binary would have counted all four as hits. Score recognition against the truth, never as a YES/NO rate. | S074 | S074 |
| (bgc) | A FIFTH dispatch failure shape: HTTP 200 whose body is nothing but keep-alive whitespace — no JSON object at all. x-ai/grok-4.5 returned 2,167 bytes of spaces and newlines, which raised JSONDecodeError in the caller and aborted a chained multi-stage run, losing the two stages behind it. Not note (b)'s truncation, not (beq)'s payload-free socket, not (bfb)'s answer-inside-message.reasoning, not (bfe)'s 504. Two remedies, both taken: write the raw bytes before parsing (already standing, note (bdt)) and retry a non-JSON body up to three times, preserving each bad body; and never chain stages with && when any link can raise, because the cost of the crash is the stages that never ran, not the call that failed. The same cell then failed a SECOND time in a DIFFERENT shape — (bfb)'s null content with the tokens inside reasoning — and a third dispatch to the identical slug returned complete output. One cell, one seat, two shapes, two retries. | S074 | S074 · S077 — fired TWICE in one session, the first time that has happened: 1,584 bytes to openai/gpt-5.6-terra and 1,331 bytes of pure whitespace to x-ai/grok-4.5. It crashed the runner on the session's FIRST dispatch, before any guard existed, and that body was lost; the guard built in response cost the second occurrence one cell rather than one stage · S102 — qwen/qwen3.7-max at Alibaba, 924 bytes of whitespace on a source-only stage; the retry guard held and the run continued, which is the first time this shape cost nothing but one dispatch |
| (bgd) | When a claim's truth depends on a tokenisation choice, print both counts rather than picking one silently. A-mchugh-presence asserts 1,511 words. Whitespace tokens return 1,511, exactly; a letter-token regex returns 1,506, the difference being numerals and P&G. Run 1 of the check used the letter rule and would have reported the anchor wrong. The rule that reproduces a published figure is evidence about which rule produced it, and the check now declares the rule it uses and prints the alternative beside it. | S074 | S074 |
| (bge) | An instrument repaired after seeing its failures must ship its first run beside its second, or a zero yield is unreadable. E-20260731f's stage-2 checks failed 3 of 18 claims and 20 of 53 quotations on first writing; every failure was the instrument (substring matching that fired ere on were and hell on Hello?; the anchor pages' own trailing commas; ellipsis-elision; a two-text anchor). After four repairs, yield 0 of 18 and quotations 53 of 53 — and "I repaired the checker until the failures went away" is indistinguishable from "the pages are clean" unless run 1 is stored. Each repair must be justified by a convention of the audited page, stated as a fact about the page, and run 1 kept as a file. | S074 | S074 |
| (bgr) | A control item's qualification does not transfer between runs, and note (bfy)'s own prescribed remedy is what failed. (bfy) says "a control item is qualified by DATA — a prior run in which it behaved as declared." E-20260801b did exactly that: six items reproduced verbatim from E-20260730i's payload, read out of the frozen prompt file rather than retyped, where P1 scored 6/6 and P3 scored 6/6. The same two seats scored 3/6 and 4/6, and the failures are item-specific rather than diffuse — X2 (British-1890s register, options car / automobile / truck / carriage) went 0 of 5 payloads for both, having been correct at S068. What differed between the runs is small and enumerable — standing register American → British, surrounding items Portuguese → Polish, item numbers — and none of it touches X2, which carries its own register note. Rule: a control block must be scored on the run that uses it and reported there; a prior run's pass is provenance, not qualification. A design that leans on inherited controls has to be able to survive their failing, which means F1 must be written before the run and obeyed after it. | S078 | S078 |
| (bgs) | k is a mean of the unstable judgment over a set fixed by the stable one, and nobody had reported it with that structure named. On a byte-identical repeat of one rater in one session, which candidate bears reproduced at 30 of 30 cells (Jaccard 1.000) and how many options it excludes did not — 0.6333 exact agreement, and the second pass excluded +0.467 options per cell more than the first. This is the reverse of RS-20260730i §4's ordering, which measured the cross-rater version and found BEARS unreliable at 0.21–0.51 on positive cells and the exclusion judgment reliable. The two are compatible because they are different quantities, and the pair is the point: the denominator of k is set by the reproducible step and its numerator by the irreproducible one. Rule: any statistic that is a mean over a selected set must report the reliability of the selection and of the value separately — a single agreement figure for "the instrument" hides which half is broken. | S078 | S078 |
| (bgt) | A SIXTH dispatch failure shape: finish_reason: "stop" on a body that is unambiguously truncated. google/gemini-3.6-flash returned a body ending …I30 \| C3:- and then the two characters I3 — mid-item-id — with stop, at 5,996 completion tokens, the identical count to two bodies from the same seat in the same session that reported length. Note (b) is about length; (bga) is stop with content: null and the answer in reasoning; (bgp) is a line count is not a format check. None of them covers a stop that lies. What caught it was the registered failure criterion F4 — fewer answer lines than the payload has items — applied in the analysis; the runner's guard (finish_reason == "length") let it through and wrote it to disk as an accepted seat. Two remedies, both taken mid-run as amendment A6: move the item-count check into the runner so a short body is a seat failure whatever the provider says, and raise the seat's cap rather than change the seat (note (bdl)) — 6,000 → 20,000, which returned 46 of 46. Unlike S077, the higher cap cost MORE ($0.138051 against $0.053262), so (b)'s remedy paying for itself is not a rule. Rule: finish_reason is provider testimony, not a measurement; the only truncation detector a design can trust is its own expected item count. | S078 | S078 · S079 — THREE more, at two labs in one stage: bodies stopping mid-list at stop / native completed, 303 / 219 / 244 content tokens against a max_tokens of 8,000. No cap explains it. The remedy that worked was a required END terminator plus an explicit line count, both enforced in the runner: 6 of 6 seats returned 22 lines afterwards |
| (bgu) | A mutation that does not mutate is a passing test that tests nothing, and a verifier that reads the first non-empty body reads a REJECTED one. E-20260801b's verifier failed three of its own checks on first run and all three were the verifier's or the analysis's fault, not the run's. (i) A mutation declared as "turn every C1 dash into two exclusions" rewrote the literal string C1:-, which the target body does not contain, so it changed nothing and the harness reported MISSED — it was reporting honestly only because a paired identity (must NOT change) mutation existed to give the comparison meaning. (ii) The verifier resolved a seat's stored body by taking the first .raw with non-empty content, which for a seat whose first attempt failed with a non-empty payload is the rejected body; the analysis had the mirror defect and silently dropped that seat entirely. Rules: assert that each mutation actually changes the bytes it claims to change, and resolve a seat's body by the artifact the runner marks as accepted (here: the label that has a .txt beside it), never by scanning for the first one that parses. | S078 | S078 · S079 — the inert-mutation assertion is now code in the verifier's harness rather than a rule in its author's memory; 6 of 6 mutations changed bytes and 6 of 6 were caught · S081 — 6 of 6 mutations changed bytes and 6 of 6 were caught, with the byte-change assertion inside the harness |
| (bio) | A mutation test must target the finest-grained stored quantity, not the headline: thresholded counts absorb the very defect the test exists to catch. E-20260804g's third mutation flipped one graded answer inside a raw billed body, and the verifier did not notice — recovery is a >=2-of-3 threshold over a both-orderings rule, and one flip moved a site from 3 seats to 2 without moving any reported number. It caught only after a per-site, per-condition seat-count check was added alongside the totals. Rule: for every reported aggregate, the verifier must also recompute the un-aggregated quantity it is built from, and the mutation test must be aimed at the latter. Otherwise a verifier certifies that the arithmetic on top of the data is right while remaining blind to the data. Companion to (bgu) — that note makes a mutation actually mutate; this one makes the mutation observable. | S106 | S106 |
| (bhd) | A mutation test that re-runs a scoring script rewrites the results file the next check reads, so the mutation survives the restore. S084's verifier restored each mutated input TSV and asserted the bytes back, then reported six failures on the second run — because score_pair.py and score_controls.py write pair_results.json / control_results.json, and those were left holding the last mutation's output. The inputs were clean; the artifact was not. Restore is not restore unless it covers every file the mutated run can write, not just the file that was mutated — snapshot them before the subprocess and assert them back after, exactly as (bgu) already requires for the mutation itself. Cheap general form: run mutations against a copied tree, or treat any file the scoring script opens for writing as part of the restore set. | S084 | S084 |
| (bhl) | A cadence arm's alarm is invisible to check_balance.py, and its own self-test says so. ARM-atelier-cycle declares one visit at least every three sessions, and under-visiting is the alarm — the cadence analogue of a step budget. It was last worked at S085 and the tool reported no violation at S086, S087, S088, S089 and S090, because a budget: cadence arm is excluded from the budget check by construction: tools/check_balance.py's own test asserts violations(... budget="cadence", used="9") == []. So the arm ran two sessions past its declared alarm with a clean tool run every time, and the arm-staleness list — which did print ARM-atelier-cycle 4 — is a count, not an alarm, and names no threshold. Rule: at session start, read every budget: cadence arm's declared cadence off its own page and compare it to N since worked by hand; the tool does not do it. The tool fix was considered and declined as method work under the subject rule — each arm states its cadence in prose, nothing published is false, and no deliverable is blocked. | S090 | S090 |
| (bhm) | A declared loss that names no attempt is an assertion, and one counterexample is enough to retire it. register.md V14 said the child-to-parent honorific plural had "no compensation available that would not be a fabrication". E-20260802d put the passage cold to three independent translators in two framings and 6 of 6 cells marked nothing, which supports the claim — and Hertzberg's authorized 1886 Swedish marks it with a kin term («Blir mamma länge borta?»), in a target that also had the pronominal option and set it aside. A device used by the only human translator of the work on record is not a fabrication, so the rule was false while every machine result agreed with it. Rule: a translator's log may declare a loss, but the declaration must state what was tried and rejected; a bare 'nothing is available' is a claim about the language and needs a witness, not an intuition. Corollary, measured here: agreement among cold machine translations is not that witness — six of them missed what one 1886 translator found. | S090 | S090 |
| (bhn) | A rule written for a site outside the material in hand was wrong the one time it was tested, and its stated ground was false about the text. Span 1's D11 decided the rendering of «Jumal' antakoon» at ¶79 — a site in span 2 — on the explicit ground that "the greeting recurs and the register must bind". It does not recur: the formula occurs exactly once in the 573-paragraph novella, and the one occurrence was the site being pre-decided. The cost was not merely a wrong rendering (the comparator in fact took D11's side, RS-20260802d §3.4) but a binding register rule carried for five sessions on a false premise about the source. Rule, now register.md V17 and general to any serial work: no rule is written for material not yet read in place; a decision that reaches forward must be recorded as a question for the span that owns the site, never as a binding. | S090 | S090 |
| (bhp) | A grammatical loss declared without a forced second attempt is not a finding about the language pair, and across this project's whole record the declaration was wrong at five sites of six. E-20260802e took every site in the filed logs where a marking was said to be carried by a grammatical category the target lacks and declared wholly lost — 8 sites, of which 2 had already been refuted by published comparators (Dole 1896, Hertzberg 1886) — and re-rendered the remaining 6 under a brief requiring the marking to appear in any category. Three independent seats, blind to provenance, recovered the relation from the forced rendering at 5 of 6, from a length-matched unmarked decoy at 0 of 6. The one failure (Spanish grammatical gender) was predicted before the run and is metalinguistic: the relation was female and automatically so, and no English category encodes automaticity. The same census found 8 Class B sites where these same logs record the move being made — pronoun → benefactive verb, suffix → possessive, pronoun → address noun — so the rule was being applied and denied by the same hand, eight times each, with nothing connecting them. Rule: a log may not record a grammatical marking as lost until a rendering has actually been attempted that requires it to appear somewhere else; the absent category is a fact, the absent marking is a hypothesis. This is framework/v0.1/'s R1 and the first prescriptive recommendation this project has admitted. Corollary from T-wang-liulang-R04-v1 D13, which named two English renderings and rejected them as "costume" while asserting "no English slot": impossibility and ugliness are different claims and a log must not make the second while writing the first. | S091 | S091 |
| (bhq) | A seat's hidden reasoning is a capacity constraint, not only a cost one, and max_tokens sized from the answer produces zero-content bodies. Note (bhf) has fired three times on price — P2 gemini-3.6-flash at 4–6× P1 for the same passage. E-20260802e fired it on capacity: at max_tokens 4,000, sized generously for 24 one-line answers, P2 returned finish_reason: length on four of four attempts (4,177–5,691 reasoning characters) and P4 kimi-k3 returned length with 16,803 reasoning characters and ZERO content characters for eighteen one-line ratings. Both recovered at 16,000–20,000 with byte-identical prompts. The failed bodies billed $0.196 and are ledgered. Fired again S106: P4 at max_tokens 12,000 returned length with 12,000 reasoning tokens and zero content, billed $0.2128698 — 39% of that session's spend — and the seat was replaced rather than re-sized, which is what (bhf)(i) prescribes once a slug has failed the shape twice. Rule: size max_tokens from a seat's measured reasoning appetite plus the answer, never from the answer alone; and when a body returns length, raise capacity and re-dispatch as a declared amendment rather than accepting a truncated body (note (b) already forbids the latter). The worst-case estimate must then be built from the raised cap — note (abc). | S091 | S091 · S092 — fired as a precaution, not as a failure: caps of 16,000 and 24,000 were set from this note and six of seven bodies then spent over 95% of their completion tokens on reasoning for a 143-character answer; 7 of 7 accepted first attempt, 0 length failures. |
| (bhs) | A failure criterion built as a ratio is not a test when its denominator can be measured as zero, and this project wrote one. E-20260803-a4-set registered the paraphrase floor must not exceed 2 × R(s, j), the same juror's retest floor on identical text. P1 reproduces itself exactly on accuracy, voice and affect — retest floor 0.000 — so 2 × R is zero and any movement whatever fires the criterion. It fired on four of P1's six senses and three of P2's, on differences that are every one of them exactly 0.500: one juror, one pass, one point of a seven-point scale. The pooled comparison the criterion replaced says the opposite and is right — 0.139 against 0.233. The amendment that introduced the ratio form was an accepted BLOCKING critic finding and was a genuine improvement in one respect (it paired the two quantities per juror) while introducing this defect in another. Rule: a criterion stated as a ratio, a fold-change or a multiple must state what it does when the denominator is zero or near it, before the run; the cheap general form is an additive floor (P ≤ R + c) or an explicit clamp. Corollary, and the reason this is worth a note rather than a line in a limits section: an accepted critic amendment is not thereby a correct amendment, and the acceptance record must not be read as one. | S094 | S094 |
| (bht) | A predictor test needs a no-information benchmark registered beside its null, and this one reversed the reading. E-20260803-a4-set asked whether a translator's frozen log predicts the sense a blind jury scores lowest. It registered a criterion (r̄ ≤ 2.20) and an exact null (each item's rank multiset, 46,656 outcomes). The test fired: r̄ = 2.000, p = 0.0098. It then took nine lines of arithmetic, computed after the run and registered nowhere, to establish that a predictor which never opens the log and names naturalness every single time scores 1.500 — better. The null asks does it beat chance; the benchmark asks does it beat the cheapest thing that is not the predictor, and only the second question was load-bearing. The same shape would have caught a majority-class classifier reported as accurate. Rule: any design whose claim is that some source of information predicts an outcome must register, before the run, the best predictor that ignores that information — usually a constant, a marginal, or the outcome's own base rate — and report the two side by side. A p-value against chance is not evidence that the information was used. | S094 | S094 |
| (bhu) | contamination: is a property of the SPAN, not of the work, and a measurement on one span does not transfer to another. T-mare-au-diable-R04-v1 is filed contamination: none on a ch. II measurement against Sedgwick and Ives. At S094 the same translator, the same work, the same two comparators, on ch. XI, returned 16 tokens and 13 shared 12-grams against Ives and 13 / 2 against Sedgwick, with same-book nulls at 2–4 — a discard, twice over. The ch. II figure was then recomputed on the filed text and reproduced exactly (9 and 10 tokens, 0 twelve-grams), so nothing published is false; what is false is the reading of the front-matter field as a fact about the work. Daudet fired the same session at 14 / 3. Three of the project's last five gate runs have fired, and the two firings here are on works whose other spans are clean. Rule: the gate is a per-span measurement and must be re-run for each new span of an already-gated work; a contamination: value states the span it was measured on, or it states nothing. The standing rule in CLAUDE.md already says one call per candidate unit — this note says the unit is the span. | S094 | S094 |
| (bhv) | A registered criterion that reverts a decision must be ORDERED against the criteria that vindicate it, and the ordering must be written before the run. E-20260803b measured one string in one sentence: a compensation for a Finnish honorific English cannot carry. It worked — +1.167 on the target rating, exact permutation p = 0.0368 over all 924 splits, 3 of 3 seats moving the same way, 5 of 6 quoting the device unprompted, the specificity control unmoved at 0.167. It was reverted, because a naturalness tripwire registered at ≤ 0.50 came back at 0.833 and the design's decision table made that row unconditional and first. On the raw scale the device gained more than it cost (1.167 against 0.833) and went anyway. The pricing — a defect weighted at half a benefit — was set before any call, on the ground that a translation answers for its faults before its bonuses; the point of the note is not that pricing but that the pricing existed and was ordered in advance. Without the ordering, the same six numbers support keeping the sentence, and the argument for keeping it would have been written after they arrived. Rule: a design with more than one gate states which gate wins, in what order, before dispatch; a design with only gates that can vindicate has no gates at all. Corollary from the same run: the pre-run critic's first BLOCKING finding was that no outcome of the original design could have changed anything — that is a checkable property of a design, and it is worth checking on every design that tests the lead's own work. | S095 | S095 |
| (bhw) | A no-information arm can leave a passing result intact and still change what it means. Same run. (bht) required a benchmark that ignores the information; here it was a seat shown no text at all, only a novel set in a Finnish town, 1886; a nine-year-old wakes her sleeping mother because two visitors have come in. It rated the girl's deference at 3.667, five of six cells at 4 — above both English versions. So P − N = −2.167 and M − N = −1.000: both translations under-deliver against what a reader brings to the scene, and the compensation closes just over half the gap. The recovery result (+1.167, p = 0.0368) is unaffected and still true. What changed is the sentence it licenses: not the device works, keep it but the loss is worth about 2.2 points on this instrument and the best available compensation recovers half of it — a price on a grammatical loss where the project had only ever recorded a declaration. Rule: a no-information arm is not only a veto; register it wherever a design will otherwise report an effect without reporting the scale the effect sits on. | S095 | S095** |
| (bhx) | Two accepted critic amendments broke each other, and the collision was invisible to the critic that produced both. E-20260803c carried a POSITIVE arm — the target relation stated outright in the English — as its reading check: a seat that misses it is not reading. The same critic pass then required amendment A12, a header reading "Default to NO … answer YES only if the relation is POSITIVELY MARKED in this English", to stop the instrument reading relations into unmarked prose. A12 worked — the registered floor check passed at 0.095. It also made POSITIVE fail by being read correctly: P2 answered NO to ten of twelve POSITIVE items with reasons of the form "English does not positively mark …", and YES only at the two sites where the marking sits in the speech rather than in the gloss. The registered bar was 11 of 12 per seat, so the run is void, and the void is not the seat's fault. Rule: when a critic pass changes the grading instruction, every registered gate that depends on that instruction must be re-derived under the new wording before dispatch — an amendment is a change to the instrument, not only to the prose. Corollary, and this is what makes it a note rather than a limits line: a reading check that a correct reading can fail is not a reading check, and this one had passed three earlier designs unexamined because no earlier design told its seats to default to NO. | S096 | S096 |
| (bhy) | A blind coding payload can hand the seat the answer in the GLOSS, with no target text shown at all. E-20260803d asked three seats, from the Russian alone and with every English rendering withheld, whether a transparent English calque is available for each of fourteen items. The design's own §5 certified that no English was shown. The glosses then read "the spirit of the forest" for леший, "the spirit of the water" for водяной, and quoted the formula outright for крестная сила — i.e. each item's gloss WAS the calque whose availability was the question. The pre-run critic called it BLOCKING and was right; every gloss was rewritten to describe the referent by what it is believed to do and where it is met, with no paraphrase of the word's parts. Rule: when a design withholds one representation of an item, audit every other field for a paraphrase of it — the gloss, the context sentence, the item's own id. Corollary the same pass forced: the repair is reduction, not removal (among trees still carries half of wood-), and the record must say which it achieved. | S097 | S097 |
| (bhz) | A probe run to SELECT material is priming, and it primes on exactly the material the study is about. Before translating a word of «Бежин луг», S097 ran a string-frequency count over both published comparators — domovoy, nymph, sprite, water-, wood-, rusalk and twelve others — to decide whether the class was worth a session at all. It printed counts, no sentences, and it still told the lead which comparator transliterates and which substitutes, on the exact fourteen items the run would then measure. The project's standing rule (never prime a workshop translation with a stored published rendering of its own source) was not broken by the letter and was gutted in effect. The design absorbed it — the lead is excluded by construction from the run's primary and the priming is declared in full on the artifact — but that was a repair, not a plan. Rule: a selection probe over a comparator is part of the translation's provenance and must be run in a form that cannot leak, or run AFTER the translation is frozen. A count is not automatically safe; what makes a probe safe is that its output cannot be read as a handling. | S097 | S097 |
| (bih) | A write-once guard that holds across a script's retries does not hold across the script being run twice, and the second run silently overwrites a BILLED body. Note (bgz) made every attempt inside one invocation uniquely labelled. At S101 run_critic.py was launched, appeared not to have started, and was launched again; both dispatches were billed, and the second wrote to the same label and destroyed the first's stored .raw. Its cost survived only in the run console, and the per-request re-sum was short by $0.0586302 until it was added by hand — the key-usage cross-check is what would otherwise have carried the discrepancy. Rule: a runner refuses to overwrite an existing .raw and labels a repeat dispatch _re2, _re3; write-once is a property of the FILE, not of the loop. Second half of the note, and the more interesting one: the two dispatches were byte-identical prompts to one slug at temperature 0, routed to two providers, and returned different verdicts (NEEDS-AMENDMENT, four findings / NEEDS-REDESIGN, six findings) with only two findings in common. How much a single pre-run critic pass warrants is now an open question — the accident was worth more than the call it duplicated, and the union of the two was accepted. | S101 | S101 |
| (bik) | An answer-line guard must accept every format the instruction actually permits, or it converts good bodies into seat failures and spends the reserve on nothing. E-20260804d told seats to output ITEM=<id> ERRORS=<n> and set line_re=r"^ITEM=". P2 wrote ITEM I01 ERRORS=1 — the separator, not the format — and both of its stage-1 bodies were complete, terminated, correctly populated, billed, and rejected as seat failures, triggering two reserve dispatches at $0.0575388 that bought nothing. The bodies were recovered from disk afterwards and scored. This is note (bgw)'s failure mode from the other side: (bgw) says a guard must not carry a stale pattern forward; (bik) says a guard must not be stricter than the task it guards. Rule: write the acceptance pattern to match on the item identifier, tolerating any separator (^ITEM[= ]), and — because a rejection is expensive and silent — print the first 200 characters of any body a guard rejects, so a format mismatch is visible at the moment it costs money rather than at analysis. | S104 | S104 |
| (bil) | The key-usage delta LAGS the per-request sum by minutes, so a closing snapshot taken immediately after the last dispatch under-reports — and that lag, not a discrepancy, is what has been reported as "not closing". S105 snapshotted at the last dispatch and got a delta of $0.078312 against a per-request sum of $0.132372; the same key re-read a few minutes later returned a delta of $0.132371900, equal to the sum to 1e-9. The intermediate snapshot showed the reason exactly: it had caught P1's $0.0085345 and neither P2's nor P3's. S104's non-closing cross-check ($0.482830081 against $0.534977746) is the same shape and the same direction, and there is no evidence it was anything else. Rule: take the closing snapshot when the ledger is written, not when the last call returns, and re-read once before reporting a residual. Note (bfv)'s void-on-concurrency rule is unaffected and still governs; this note only removes the false positive that a prompt snapshot manufactures. [fired S107, second time, and the same magnitude: an immediate read gave a residual of −$0.0786, a twenty-second-later read gave 0.000000000. Twice now the lag has been large enough to look like a real billing anomaly, so the rule is restated as an instruction rather than a caution: never report a residual from a snapshot taken at the last response — re-read, and report the second figure.] | S105 | S107 |
| (bit) | A pre-run seat probe is worth what note (bhf) has been claiming since S087 — and a token ceiling learned from it must be applied to EVERY stage, not to the one the author was worried about. E-20260805 ran the probe for the first time. It fired immediately: P2 returned finish_reason: length with empty content twice at max_tokens 64 — S108's exact failure, which cost that session 60% of its spend — for $0.000188. Re-probed at the run's own ceiling P2 answered normally, at $0.001188 against P1's $0.000094 for the identical 24-character reply: 12×, all hidden reasoning. So the probe's yield is not "is the slug alive" but a per-seat minimum token budget, and that is what the probe should return. The lesson then cost money anyway: the amendment raised stage 2's ceiling and left stage 1 at 4,000, and stage 1 burned $0.147 on six length bodies from the two heavy-reasoning seats before every stage was raised. Rule: probe at the ceiling the run will use, record the per-seat minimum, and apply it wherever the probe's condition holds. Seat-failure waste fell from S108's 60% to 23%. | S109 | S109 |
| (bjc) | A difference statistic can shrink from either end, so a design that changes a rating item's wording and compares a DROP must report the two scores that make the drop, separately, or it will report an improvement that is a demotion. E-20260805e re-ran Tier D's failed drop(naturalness) under the revised naturalness string and the number fell from 1.083 to 0.750, across the ≤ 0.75 bar. The whole of that fall is the reference losing points: reference 6.125 → 5.417 (−0.708) against the damaged copy's 5.042 → 4.667 (−0.375), and the reference's ceiling ratings went from 6 of 24 at 7 to 0 of 24. Read as a drop alone it looks like accuracy damage no longer bleeding into fluency; read as two numbers it is a criterion string that marks the undamaged text down. And the same run shows the mirror image: on two new items the identical 0.333 shrinkage came entirely from the damaged copy scoring higher while the reference did not move at all. Same effect size to four decimals, opposite mechanism. Rule: report the reference mean and the comparison mean beside every drop, in every cell, and never let a ceiling criterion stand in for this — the run's FC4 watched only for reference scores at the ceiling and would have passed a cell whose reference had simply been marked down. | S113 | S113 |
| (bjd) | Resolvability must be defined WITHIN a run. Distance from a previously published point estimate is a random throttle, not a noise floor. E-20260805e v1 made a wording effect reportable only if it exceeded |this run's retest − the published 1.12|. The pre-run critic showed the failure mode in both directions: a lucky retest landing near 1.12 licenses almost any effect, an unlucky one suppresses a large one, and the comparator is itself a high-variance random variable on 12 correlated units. The replacement — an exact paired permutation test on the per-unit differences, plus a frozen minimum effect size — is what withheld the run's primary: the effect cleared δ_min at 0.333 and the permutation test returned P = 0.138. The two properties are genuinely different and this run measured both: the cell mean reproduced across three separate runs to 0.037, while the per-unit drops swung by up to 1.5 scale points. Rule: a gate on "is this effect real" reads the run's own paired structure; a published prior figure is a retest STATISTIC to report, never a threshold to gate on. | S113 | S113 |
| (bje) | Navigating a comparator file is priming, and the remedy is to move the locus, not to flag the artifact. Note (bhz) covers a selection probe over a comparator; this is the plainer case that had never been named. At S114 the lead located Quijote I.20 by searching the Gutenberg files for chapter boundaries and printed Ormsby's and Smollett's English of the passage it was about to translate — not a count, the prose itself. R04 §1 was breached before a word was rendered. The temptation is to translate anyway and declare the priming, which the regime permits; that produces an artifact that is evidence about nothing, because a primed rendering cannot speak to independence, to process, or to what the source forces. R04 §1 supplies the real remedy in its own text — "the passage selected must be one whose rendering has not already been seen" — and it was applied: the locus moved forward within the same chapter, to prose whose English had not been opened. Rule: locate the locus in the SOURCE file and freeze its boundaries before any comparator file is opened; if a comparator's prose has been read, move the locus rather than declaring the priming. The declaration is the fallback when nothing can be moved, not the first response. | S114 | S114 |
| (bjf) | Two works by one author can be the same prose, and a corpus arm that does not check for reuse weights one text twice while reporting two. E-20260805f built Machen's "original English" arm from The Great God Pan, The Three Impostors and The House of Souls — and The House of Souls (1906) reprints The Great God Pan whole, along with The Inmost Light. A 500-token probe from the Pan file appears verbatim in the Souls file. The arm was counting the same prose twice, and weighting it by whichever collection happened to be longer. Nothing in the design would have caught this; it surfaced only while checking the premise of an unrelated critic finding about control-cell text reuse. Anthologies, collected editions and "complete works" files are the common carriers, and a title that names a collection rather than a work is the signal. Rule: before blocking a multi-work corpus arm, cross-check every pair of member texts for a shared n-gram run — one contiguous run of a few hundred tokens is a reprint, not a coincidence — and record the check. The cost of missing it is silent and invisible in every downstream statistic. | S114 | S114 |
| (bjg) | A multilingual panel cannot proxy a monolingual reader, and the questions where that matters are exactly the ones about what a reader does NOT know. E-20260805g put a passage containing one untranslated Swedish line to four seats and asked, uncued, for a retelling. All four reported the content of the line — one naming the language with no warrant on its page and translating it, one translating it silently — where the artifact's own declared reader is "a reader with no Finnish … no facing text, no notes". The translator's log had declared a cost that depends entirely on the reader being unable to read that line, and no prompt can make a model not know Swedish; instructing it to pretend would measure compliance, not comprehension. The same trap waits wherever a design turns on ignorance rather than judgment: an unglossed foreign phrase, a dialect a target reader would not have, an allusion the reader is meant to miss, a name whose connotation is supposed to be opaque. Rule: before designing any comprehension measure, ask what the seat would have to NOT know for the measure to work, and check whether the panel knows it. If it does, the panel cannot answer, and the honest output is that the cost is unpriced rather than a number. Corollary: this is the first thing in this project that a bigger, better panel makes worse. | S115 | S115 |
| (bjh) | A guard that records THAT a dispatch failed but not WHY is half a guard, and the half it keeps is the useless one. call.py wrote repr(e) to a .err file on any transport failure. An HTTPError's repr is <HTTPError 400: 'Bad Request'> — the provider's actual message is in the response body, which urllib makes readable exactly once, from the exception object, and which was being discarded. Six mistral-medium-3-5 dispatches died undiagnosably at S115 until one call was reproduced by hand outside the runner, whereupon the provider said "top_p must be 1 when using greedy sampling" — a constraint that fires only when the reasoning field is present, and a one-line fix. Rule: store the error BODY beside the repr, and print the first 200 characters at the moment of failure. This is note (bik)'s rule — print the first 200 characters of any body a guard rejects — extended from bodies that arrive to bodies that do not, which is where it was missing. Repaired in call.py the same session. [FIRED S116, second session running, and the repair paid: mistral-medium-3-5 returned the same 400 twice as a screen seat and the stored body named the cause at the first failure, so the declared reserve took the seat under (bfc) with no hand-made reproduction and no lost time.] | S115 | S116 |
| (bjp) | A serial arm's cadence must be declared in units of the ROTATION, not of sessions, or it is unmeetable the day it is written and every visit reads as an overrun. ARM-atelier-cycle declared one visit at least every three sessions and was visited at S085, S090, S095, S100, S105, S110, S115 — every fifth session, seven times, without one exception. That is not drift: the selection rule gives one principal unit per session to the most neglected of five eligible tracks, so a track returns on a period of about five, and no discipline available to a session could have produced three. Two consequences followed and both were costly. check_balance.py structurally cannot see it — its own self-test asserts that a budget: cadence arm never fires a violation — so seven overruns produced zero warnings and the alarm was raised by hand, first at S090. And the unmeetable cadence was twice used as an override reason to take T1 over the tool's named track, so an alarm nobody could clear became an argument for taking the arm. Rule: a serial arm declares its cadence as one visit every N rotations, where a rotation is the number of eligible tracks; an arm that wants to be visited faster than the rotation is asking for an override, and must say so at birth rather than accrue an overrun that reads as neglect. ES-20260806-craft-report-koyhaa-kansaa §2. | S120 | S120 |
| (bjq) | A device you exclude from your own arm by argument is a device a published translator may have used, and the census will tell you — so measure published practice BEFORE ruling a device out, not after. E-20260806-published-loss §5 excluded the archaic second-person pronoun from its FORCE arm, before translating and before any published English was read, on the argument that thou is one-sided: it can mark the down-address and English has no pronoun for the up-address. Oxenford 1844 used it, at four of nine sites, and it works — because once thou is in play, plain you becomes the marked contrast term. The exclusion's provenance is verifiable (75cd449) and it stands as registered; what it cost is that the run's one constructive arm used a weaker device family than the only translator who carried the relation. Rule: when a design must choose among the devices a recommendation licenses, extract what the published comparators actually do FIRST — a device count is textual, needs no panel and no jury, and is available before any translation is written — and let the exclusion be informed rather than argued. The general form: an argument about what a target language can do is a hypothesis, and the published record is the cheapest test of it that exists. RS-20260806e-published-loss §4.3. | S121 | S121 |
| (bjt) | A fix applied to a run inside the session that needed it, and not committed to the runner, does not exist for the session that resumes the run. S121 hit repeated provider hangs, pinned deepseek/deepseek-v4-pro's routing to the providers that had returned cleanly, wrote that it had done so in RS-20260806e §5 and in the ledger — and the pin was never in run.py. S122 inherited a runner identical to the one that hung, and the published sentence "routing for that slug was pinned" was false of everything that survived the session. Sessions here are ephemeral and only what is merged exists, which the charter says about work and is exactly as true of configuration: a mid-run mitigation is written into the committed runner, or it is written into nothing. Costed nothing this time — the fix took one edit and the three calls returned in seconds — and the false sentence is the real damage. [FIRED S124 and honoured: the deepseek-v4-pro provider pin was written into the committed call.py, not into the result page alone, before the run was restarted.] | S122 | S124 |
| (bju) | [FIRED AGAIN S125, predicted in advance: the same selection rule was reused deliberately so the instrument would be unchanged, intrusion was declared inert-by-construction before the run, and it came back at 4 of 16 — chance.] A selection rule that fixes a property destroys that property as a discriminator, and the rule usually looks like good design when it is written. E-20260806f chose all four loci by one source-side rule — a span in which the narrator presents himself and almost nothing happens — because that maximises persona signal and minimises plot. It also guarantees that every narrator scores high on intrusion: the four came in at 4.50–6.00, and the single-axis matcher on intrusion reached 3 of 16, worse than chance, while irony reached 13. The axis was not weak; the sampling frame had no variance left in it. Same shape as RS-20260806e §7, where an admission condition defined on whether the English conveys the relation selected for sites whose content already conveyed it. Rule: before freezing a selection rule, write down which of the measured variables it holds constant, and either vary it or declare it unmeasurable in this run. RS-20260806f §6. | S123 | S123 |
| (bjv) | A request killed before its response reaches the process can still be billed, so killing a hung call is a cost decision and not a free one. Note (bhf)'s S116 entry records the opposite — "the key-usage delta across that killed request is exactly 0.000000000: a request killed before its response returns bills nothing" — and S123 measured $0.0011020 and $0.0165294 across two killed pre-run critic attempts on moonshotai/kimi-k3, the whole of that session's unexplained key-usage residual. The S116 sentence is true of the request S116 killed and is not a general fact, and it is the sentence a future session would rely on before deciding to kill rather than wait. Both entries stand; neither generalises. Practical form: the cross-check does not close after a kill, and the residual is the kill, not a mystery. RS-20260806f §7 limit 11; config/budget.md S123. [FIRED AGAIN S124: killing run_coding.py mid-flight to apply a provider pin left a $0.016075170 residual — the whole of that session's unexplained key-usage gap, identified as the killed call rather than reported as a mystery.] | S123 | S124 |
| (bjr) | Provider latency is a design risk, not just a cost risk, and a long run needs resumability before it needs anything else. config/models.md carries a standing caution that OpenRouter routing can move the price of a slug by ~4×. S121 found it moves wall time by more than 20×: deepseek/deepseek-v4-pro routed to eight distinct providers in nine calls, and several hung past a 600-second timeout, twice on the same block, stopping a 36-call grading stage at 33 and forcing the session to land an incomplete run. Two things saved it and one was added mid-run. Every raw body was written to runs/ before anything was computed from it (standing practice), and a resume clause — reuse a stored body instead of re-billing the call — was added while the run was stalled, which made three restarts free and leaves the remainder at three calls. Rule: any stage of more than ~10 calls is written resumable from its stored bodies from the start, and a stage that stalls is restarted rather than waited on. Corollary, learned the same session: pinning provider.order for a slug is an infrastructure fix, not a design change — the slug is what config/models.md specifies and provider is logged as provenance either way — but it must be declared on the result page. | S121 | S121 |
| (bjw) | A scale's apparent one-sidedness is a property of the CORPUS put in front of it, not of the scale — and the cheapest test is to make the missing arm. E-20260806b withheld two of three primaries when its gate F1′ found the negative half of the explicitness and popular-speech scales used in 2.1% and 4.8% of 186 cells. Two explanations were open — a scale that cannot go negative, or five translations none of which went down — and no amount of re-reading the run could separate them, because both predict the same table. E-20260806g separated them by writing one deliberately downward translation under a frozen regime and putting it in the same line-up: the rates went to 17.0% and 15.7%, register to 32.7%, and the new arm read −1.500 with all three seats agreeing. The instrument was never the problem. Corollary, and the sharper half: the other arms' negative rates moved by a factor of three on byte-identical texts once the line-up changed, so negative-half usage is a contrast effect and no figure from a comparative coding run is an absolute magnitude. Rule: when a gate fires on scale one-sidedness, the first hypothesis is an empty region of the corpus, not a broken instrument; constructing one instance is a translation, costs $0, and settles it — and any design quoting a magnitude from such a scale states what else was in the line-up. RS-20260806g §10. | S124 | S124 |
| (bjx) | A free-text justification tied to "your most extreme code" cannot support a gate defined on ONE named scale, and the mismatch is invisible until someone counts. E-20260806g's F4 — the gate separating the coders read it as low from the coders read it as bad — classified WHY lines attached to register codes, while the prompt (inherited verbatim, deliberately) asks each seat what drove its most extreme of three. Any version whose top code was explicitness or popular-speech therefore yields no register evidence at all, and the gate would have "passed" on whatever happened to be left. The pre-run critic found it; the repair kept the instrument frozen: restrict the gate to cases where the target scale IS the (co-)most-extreme, compute that coverage as a number, and declare a floor below which the gate is UNDECIDABLE rather than passed. Coverage came in at 22 of 45 against a floor of 15. Rule: a gate that reads free text names which text it needs and demonstrates that the prompt produces it — a justification field answers the question it was asked, not the question the gate wants. E-20260806g A1. | S124 | S124 |
| (bjy) | A fixed-vocabulary lexical detector reports a PRESERVED pattern as absent, and it can only err in that direction. machine.py's signifying-network check looked for adorn|deck|furnish. Arm V preserved Verga's adornare network — "does the place up for you" → "that's for the place you did up for me" — and the detector returned empty, which reads on the page as network destroyed. Nothing in the run would have caught it: a zero from a keyword detector looks exactly like a true negative. The pattern was extended and the network found, and RS-20260806b's "one network survives and one does not" is a lower bound for the same reason. Rule: every count from a keyword detector is a LOWER bound, a zero from one is never evidence of absence, and an arm that scores zero is read by hand before the zero is recorded. RS-20260806g §4. | S124 | S124 |
| (bjz) | A recognition confound is broken by rebuilding the corpus, not by arguing about it — and the gate must be probed on BOTH sides, because a prior can only anchor a match if the seat identifies both texts. RS-20260806f measured a source→target persona crossing at 16 of 16 and had to report in its own headline that both recognition seats named all four works and all four authors from the English, leaving carriage and shared-prior equally consistent with the table. E-20260807 rebuilt the corpus out of authors with no reachable English translation record, held the instrument byte-identical so fame was the only intended variable, and promoted recognition from a reported number to a gate: 0 of 8 from the English, 1 of 8 from the source, crossing 16 of 16 anyway. Three parts of the rule. (i) The English-side probe alone is not the gate the confound needs; the pre-run critic's BLOCKING 1 is why the source-side probe exists, and it is where the run's only true hit came from. (ii) The scoring rule must be the correct surname, not the recognised flag: when these seats guess they guess a famous author of the right language and period — de Maistre for Karr, Saltykov-Shchedrin for Levitov, Shiga Naoya for Miyaji — so four confident wrong names read as four recognitions under the flag and as zero under the rule. (iii) The gate is a tolerance, not an α: naming a correct nineteenth-century author by chance has no estimable rate, and a design that promises to calibrate one is promising the assumption it is calibrating. What the gate buys is bounded and must be said in those words: it excludes identification-mediated recognition, never a diffuse shared prior. RS-20260807-narrator-unknown §§1–2. | S125 | S125 |
| (bka) | Matching your TARGET lengths does not defeat a size-surrogate matcher, because the surrogate compares target size against SOURCE size — match the sources. RS-20260806f §6 found a one-dimensional nearest-neighbour on word count getting 3 of 4 works right and told the next design to match on length. E-20260807 matched the four English renderings to 448–466 words, ±2.6%, and wrote into its own design that the matcher was "expected to be at chance". It came back 2 of 4 — because the four source spans were 352 / 431 / 466 / 391 in the design's units and the surrogate reads across the pair, not within a side. One of its two hits is free by construction: the frozen rule fits the CJK character-to-word divisor to the item so the surrogate is as strong as it can be, which makes that item match itself exactly. Rule: a length-confound control is defeated by matching the sizes the surrogate actually compares — for a cross-language matcher that is the SOURCE sizes, in the surrogate's own units, including any fitted conversion. Corollary: a design's stated expectation about its own control is a prediction like any other and is reported when it fails. RS-20260807 §6. | S125 | S125 |
| (bkb) | A mutation test that cannot change the decision it tests is not a test, and a uniform perturbation cannot change a nearest-neighbour argmin. verify.py's inherited mutation set added +3.0 to every axis of one profile vector and asserted the recomputed match count changed. At S125 it was caught 2 of 3 — and the miss is a property of the mutation: a uniform shift moves a point equally away from all four consensus points, so the argmin need not move at all. The verifier was not weak; the perturbation was orthogonal to the decision rule. Five mutations chosen to be decision-relevant — four label swaps within a seat, one flattening of a seat's whole target side — were added and caught 5 of 5, and both sets are reported. Rule: pick each mutation by asking what the analysis's decision rule is a function of, and perturb THAT; a mutation set that has never been checked against its own decision rule can report a clean verification of nothing. Corollary: report the mutations that did not bite, with the reason, rather than only the count. RS-20260807 §8. | S125 | S125 |
| (bkt) | A positive control must manipulate the SAME COMPONENT OF THE GRAMMAR as the phenomenon whose absence is being claimed. E-20260808g set out to show that a Hungarian tag's tense morphology is not carried into English, and gated that null on a control that changed the tag's lexeme (mondá → kiáltá, said → shouted). The pre-run critic's BLOCKING 2: models carry a lexeme change trivially, so passing it would have licensed nothing about morphology, and the null would have been unreadable. The control was replaced before dispatch with a same-lemma, same-slot tense change (mondá → mondja, said → says). Both controls then returned exactly 1.000 — but only the second one is the reason the 0.000 means anything. Rule: write the control by asking what kind of thing am I claiming is not carried, and manipulate a thing of that kind; a control one component away is a demonstration that the instrument works on something else. Companion to (bkb): pick the perturbation from what the decision rule is a function of. RS-20260808g §3. | S138 | S138 |
| (bku) | Before claiming a translation loses a source's marking, manipulate the marking and see whether independent hands move. This project has retired one prediction after five attempts in five language pairs (framework/v0.2 S1), and every one of them began by assuming the source's marked form meant something and going to look for its loss. E-20260808g inverted the order on a new device: it took eleven paragraphs, changed one token in each — the tag's tense — gave the two versions to four independent hands who never saw both, and asked only for a translation. Cross-arm disagreement 0.0000 against a same-arm floor of 0.0000, with the component-matched control at 1.0000. The cost was $0.169 and four API calls, against the several thousand dollars and dozens of sessions the loss-hunting shape has consumed. Rule: where a design is about to assert that English cannot carry a source form, the cheap first move is a minimal-pair generation probe — the loss claim is only worth pursuing if independent hands' English actually moves with the source form. Corollary, and the substantive one: a form bound to a construction may be marking something the target marks by other means — Hungarian marks the attribution slot morphologically, English marks it with a bleached said, and neither loses anything to the other. | S138 | S138 |
| (bkv) | A gate must be measured at least as precisely as the quantity it gates. E-20260808h's primary was read on 6 ratings per site (3 rater seats × 2 calls); its source-side gate — the precondition that the manipulation was legible as a change of manner at all — was read on 2 ratings per site, one per arm, on the same five-point scale, whose observed between-seat spread is 0.696 scale points. The gate's bar was +1.00. It returned +0.667, F1 fired, and a primary computed from three times as much data was withheld by a coin. Worse, the same gate returned the wrong sign on the positive-control block (−0.75 on 8 ratings) while the English side returned +0.786 on 24 — the two sides disagreeing about a manipulation both were reading. Nothing in the design was wrong except the allocation: the manipulation was grammatical (0 residual forms, 0 grammaticality flags in 32 site-arms) and every other check passed. Rule: when a design has a gate, count the ratings behind the gate before dispatch and give it at least the precision of the thing it can withhold — a gate is not a cheap formality bolted to an expensive measurement, it is the measurement's licence, and an under-powered gate does not fail safe, it fails loud. Corollary, and the one that costs money: a gate is the wrong place to economise seats. RS-20260808h-trajectory §1, §4. | S139 | S140 — FIRED AND DEMONSTRATED, which is rare for a note of this kind: the repair was carried out and the number moved. E-20260809a stage R re-measured S139's source-side gate on the byte-identical texts, the same prompt string (asserted against the stored request before dispatch — the builder refuses to run otherwise) and the same +1.00 bar, changing nothing but the allocation: 2 ratings per site from 1 seat became 6 from 3. It returned +1.389 where it had returned +0.667, the gate passed, and a primary withheld for a whole session was released. S139's other under-measured figure resolved with it: the positive-control block's wrong-sign −0.75 on 8 ratings came back +0.500 at parity. So the note's cost is now priced in both directions — an under-powered gate does not fail safe, and repairing one is worth a session's principal unit. RS-20260809a-trajectory-ja §3. |
| (bkw) | A seat that returns an empty body after spending its whole cap on hidden reasoning can often be repaired by turning reasoning off, not by re-seating. Note (bhf)'s remedy — zero content means change the seat, not raise the cap — is right about the cap and incomplete about the seat. In E-20260808h four bodies came back with zero content at cost: two critic seats and two of the three hands, $0.273475, 42.5% of the session. Both hands were recovered on the same seat by sending "reasoning": {"enabled": false} (deepseek-v4-pro, which then cost $0.0024 against $0.0040 for the dead call) or "reasoning": {"effort": "low"} where the first form 400s (grok-4.5). Rule: on a zero-content body, try the reasoning switch on the same seat before re-seating — re-seating costs a whole new call and loses the seat the design chose; two provider-specific parameter forms exist and one of them will be accepted. Caveat measured in the same run: google/gemini-3.6-flash rejects reasoning.enabled with a 400, so the switch is attempted, not assumed, and the fallback to re-seating stands. Companion to (bhf), which keeps the prohibition on raising the cap. RS-20260808h-trajectory §9. | S139 | S140 — extended: APPLY IT BEFORE THE FAILURE, NOT AFTER. E-20260809a dispatched every seat with a dead-body record (nemotron-3-ultra, kimi-k3) with "reasoning": {"enabled": false} from the start rather than waiting for the empty body. nemotron, which spent $0.046932 on nothing as S139's registered critic, returned a ten-finding critique for $0.011126; kimi-k3, dead in four of its last six load-bearing roles, returned two clean gate calls. Session waste fell from S139's 42.5% to 7.2%. A second measurement sharpens the note against (bhq): on stage R, three of six calls died at the cap that had returned clean bodies from the same seats on the byte-identical prompt one session earlier; the cap was raised to 12,000 per (bhf)'s seat/cap distinction and two recovered, and the third recovered only on the reasoning switch, at $0.0010549. A cap raise that failed and a reasoning switch that worked, on one seat, in one stage. RS-20260809a-trajectory-ja §9. S142 — the note's own dead end is closed. Its caveat says google/gemini-3.6-flash REJECTS reasoning.enabled with a 400 and leaves re-seating as the fallback; E-20260809c lost two stage-1 bodies at that seat (one an echo of the passage, one truncated at 8 items of 14, $0.0516, 17.9% of the session) and recovered BOTH on the same seat with the other form, {"reasoning": {"effort": "low"}}, at a raised cap. Applied prophylactically to every later stage, the seat then returned 9 of 9 clean bodies. So the caveat is two-sided: that seat rejects one form and accepts the other, and re-seating was never needed. RS-20260809c-generic-voice §9. |
| (bkx) | A control is not frozen until its constituent has been shown to exist, as the control needs it, in the actual candidate material. E-20260809a froze a positive control that swapped an address title for a second-person pronoun in the tail, on the strength of a whole-novel census showing 先生 75 times and 奥さん 17 times inside dialogue. Applied to the selected sites it collapsed: most of those tokens are third-person reference, not address — the narrator talking to Sensei's wife about Sensei — and where the tail's address token was a pronoun it was the senior speaking to the junior, for whom no title exists. Only 2 of 6 sites could carry the manipulation at all, and only in one direction. The control had to be rebuilt after the design was frozen and after two critics had reviewed a version that could not be built. The census that justified it counted the right strings and answered the wrong question: presence of a token is not availability of a manipulation. Rule: before a control is frozen, run the mechanical check that the constituent it manipulates is present in the role the control assigns it, at every site the design will use, and commit the per-site count. One grep over the tails, keyed to the addressee, was the whole cost. Companion to (bkt), which governs what kind of thing a control manipulates; this one governs whether the thing is there at all. RS-20260809a-trajectory-ja §7 limit 5; E-20260809a amendment A7. | S140 | S140 |
| (bky) | A reading probe whose raters cannot quote the manipulated constituent is not measuring the manipulation, and the required quote field says so for free. Two consecutive runs on two language pairs have now produced the same shape. E-20260808h: on the pronoun block 24 of 70 quotes contain an address term and every one of them is a term the manipulation never touched, because a Russian pronoun switch leaves nothing in the English to point at. E-20260809a, on Japanese: 2 of 36 quotes on the primary contain a contraction — while the mechanical census over the same post-split English shows contractions moving +2.639 per 100 words in the manipulation's direction — and 2 of 36 quotes on the positive control name the inserted formula, the control returning −0.125 against a +0.75 bar. In both runs the raters quote events instead: "the door of my house is closed to you", "The wife's eyes filled with tears", "he burst out laughing". The hands encode the change; the probe does not see it. Rule: where a design asks readers whether something changed, make the required quote a first-class measure and check it before the primary — if the quotes do not name the manipulated constituent on the block where the manipulation is largest, the probe has not measured the manipulation, whatever its means say. The corollary is cheap and load-bearing: validate a positive control by its quote rate, not only by its effect size. RS-20260809a-trajectory-ja §5; RS-20260808h-trajectory §5. | S140 | S142 — third consecutive firing, on a third language pair, and the first time the note WITHHELD a primary rather than explaining one after the fact. E-20260809c promoted the quote to a declared gate G6 with a chance baseline computed and committed before dispatch (0.2887, the share of the manipulated arm's tokens lying inside a declared edit span) — the pre-run critic's one wholly correct finding, which caught that the original bar of 0.33 was below chance. Observed: 0.2857, at chance, so F5 fired and P1 is descriptive only. The baseline is the addition the note now carries: a quote-rate gate without one cannot be read at all. RS-20260809c-generic-voice §7. |
| (bkz) | A ladder of regimes executed by one hand cannot tell the programme costs accuracy from this translator's rule-following costs accuracy, and the arm that separates them is one extra call. RS-20260807e §5 and RS-20260808c §5 both published following either declared programme cost about 1.1 points of propositional accuracy — in two language pairs, from opposite programmes, with a contamination control and a compliance audit attached. Neither run contained an unruled independent arm: both had IND-R07 and IND-R08 and no IND-R06, so the quantity the claim is about had no lead-free measurement anywhere in the design. E-20260809b added that one arm for $0.0092 and it inverts the generality of the claim: the unbriefed hand's tax is +0.024 (P = 1.000, CI [−0.166, +0.250]) on the very passage where the lead's is +0.714, and +0.167 (ns) on the new one. The tax is real — it survives on a source no model can quote an English translation of, P1 +0.354, P = 0.039 — and it is a property of the lead translating under rules, not of rules. Rule: whenever a finding is about what a REGIME does, the design needs at least one arm of that regime executed by a hand that is not the lead, including the baseline regime — a ladder missing its own control condition in the independent half will produce a claim about programmes from evidence about a translator, and no amount of contamination control or compliance auditing inside the lead's half can catch it. Corollary measured in the same run: the published magnitude did not survive re-measurement either — the identical statistic on byte-identical texts under a byte-identical prompt came back +0.619 against +1.095, so cite such a figure for its direction and not for its size until it has been re-measured. RS-20260809b-programme-tax §3, §4.1, §4.3. | S141 | S146 — AMENDED by note (blk): the remedy stated here is sufficient for a positive finding and NOT for a null. Six hands returned +0.98 where this note's one hand returned +0.024 |
| (bla) | A distributed carrier can move without any edited site naming it, so a site-level check cannot certify that an operator left it alone. E-20260809c's AWKWARD arm was built to roughen English without touching any carrier of voice, and arms.py asserted exactly that, mechanically, against ten named CLOSE strings — all ten survived byte-identical. The arm nevertheless came back +0.476 of diction formality above CLOSE, P = 0.03125, and register is the first of the five carriers voice's own entry names: twenty-one nominalisations, agentive passives and with the exception of periphrases are individually neutral and collectively bureaucratic. The design gated the other arm's register (G2 = 0.143, passed) and never gated this one's, so a registered reverse dissociation failed with part of the failure built into the manipulation. This is the mirror of RS-20260806-same-man's finding that register and rhythm are distributed properties a site edit cannot move: it cannot move them at a site, and it can move them in aggregate. Rule: where a design asserts that an operator holds a property, the property is measured on every arm the assertion covers, not only on the arm the primary is read from — a site-list check certifies sites, and a distributed property needs a distributed measurement, which here was one more blind rating stage the run had already built for the other arm. RS-20260809c-generic-voice §5, limit 1. | S142 | S142 |
| (blb) | A control's items must be matched on every property its scale can read, not only on the property it is named for. E-20260809d's G2 asked whether raters were scoring the speaker rather than the words, and selected its four items on one criterion: turns with no address content. The two turns that fell to speaker A are his long heraldic answers — "A huge human foot d'or, in a field azure…" — and the two that fell to speaker B are his short questions. Both pairs are address-free and the pairs differ in length and diction formality, which is exactly what a ceremony scale reads. G2 came back at 0.7917 against a registered bar of 0.75 and withheld a primary of +3.500 at P = 1.08 × 10⁻⁵, and the run cannot say whether the gate caught a speaker halo or its own item selection. The pre-run critic named this defect (SERIOUS 4) and it was overruled on the ground that address-freeness was the property the gate needed — the wrong ground, because a control has to be matched on the scale's inputs, not on the hypothesis's variable. Rule: before freezing a control, list what the rating scale can respond to and check the arms are balanced on each of them; a control matched on one dimension and unmatched on three is an uninterpretable gate, and an uninterpretable gate still fires. Companion to (bkx), which governs whether a control's constituent exists at all; this one governs whether its items are comparable. RS-20260809d-forced-choice §4, §5a. | S143 | S143 |
| (blc) | Substituting the names does not blind a canonical text, and the recognition probe that says so costs one line of prompt. E-20260809d replaced every proper name and the title object in Poe's "The Cask of Amontillado" — Fortunato→Guarnieri, Amontillado→Madeira, palazzo→house — and asked each seat, after its ratings, whether it recognised the story. 11 of 12 bodies named it, with author. What survives name substitution is the situation: a masonic exchange, a chaining, a walling-up. The canonical reading of that story is the run's own hypothesis, so every rating was taken by a seat that already held it. Rule: where a design needs raters blind to a work's received reading, name substitution is not a blind; ask for recognition in the same call, report the rate whatever it is, and prefer material the panel has no reading of — for a canonical author of a canonical story there may be no such passage, and that is a materials constraint to settle before dispatch rather than a limit to write afterwards. The probe is nearly free and is the only thing that turned an assumption into a number here. RS-20260809d-forced-choice §5b. | S143 | S143 |
| (bld) | A permutation that cannot move most of its tokens is not a test, and the immovable count is what says so. E-20260809e's registered null permuted MARKED/UNMARKED labels within part × lemma — the right stratification, forced by the pre-run critic, because it controls for a lemma simply being frequent. On this corpus it was nearly vacuous: 40 of the 55 marked tokens, 72.7%, sit in strata containing no unmarked token at all, because most of Mikszáth's archaic verb forms are hapax lemmas with no ordinary-past counterpart in the same part. Immovable tokens contribute the same floor to the observed statistic and to every simulated one, so the null's spread comes from a quarter of the data and the p-value (0.1253) is substantially a power failure rather than evidence of absence. Rule: before reporting a p-value from a within-stratum permutation, report how many units the permutation can actually relocate; if it is not most of them, the test's failure is reported as underpowered and the design says so in the same sentence. A conditional null bought correctness at the price of power, and both halves have to be published. RS-20260809e-legend-layer §4a. | S144 | S144 |
| (ble) | A test run on the corpus that contains the passage which generated the hypothesis must report the same test with that passage removed, and that sensitivity is often the result. E-20260809e tested whether Mikszáth's narration-internal archaic past clusters, on a hypothesis formed from two paragraphs of Part I. On the whole novel the statistic sat at the 97.5th percentile (p = 0.1253); with Part I excluded it collapsed from 7 to 1, p = 0.8050 — six of the seven contributing paragraphs were in the part that produced the rule. The pre-run critic named exactly this (SERIOUS 9: randomising across the whole novel partly re-discovers the passage that motivated the inventory) and the sensitivity exists only because the finding was accepted; had it been overruled, a p just under the bar would have been published as a device. Rule: name the passage the hypothesis came from before freezing, and register the leave-it-out recomputation as a required output, not as a robustness check to run if time allows. Companion to (blb): that note is about accepting critic findings on controls, this one about the specific finding a hypothesis-generated-from-the-data design will always attract. RS-20260809e-legend-layer §4b. | S144 | S144 |
| (blf) | A billed API request survives the client that dispatched it, so a call started in a foreground shell with a shorter timeout than the call is money spent on nothing. S145's first loc dispatch ran in a shell the harness killed at two minutes while the request was open. The request completed server-side and was charged $0.0597891; no body was ever received, nothing was written to runs/, and the loss showed up only as a gap between the key-usage delta and the per-response sum — which is the one place it can show up, and the reason that reconciliation is worth doing every session. Rule: dispatch every API call in the background, or with a client timeout longer than the call can possibly take; never in a foreground shell that something else can kill. Corollary for the ledger: an unexplained positive gap between the key delta and the body sum is a lost body, not rounding, and it is named rather than absorbed. RS-20260809g-device-cross §10. FIRED AGAIN S171, at $0.062188150: E-20260812i's first run_hands.py was launched in the foreground, killed by the harness at 120 s, and had already billed at least one translation whose response was never written. The remedy is unchanged and was not applied. RS-20260812i-footing-channel §7.| S145 | S171 |
| (blg) | A permission is not an application, and a design that prices two permitted devices is measuring their REACH unless it says otherwise. E-20260809g crossed two register devices as may use switches and defined its primary as the difference between what each bought. The pre-run critic's first three findings were one finding: each revision arm changes the English only where its own device has something to offer, so the two quantities are measured on different units, and a device looks weaker whenever its opportunities are rarer. Measured in that run: eleven of twenty-eight ⟨hand, site⟩ device opportunities produced no change at all, and the two devices' applied sets overlapped at fewer than half the sites. The primary was split before dispatch into a policy contrast (all sites, zeros included — what a translator gets across a book) and a conditional contrast (applied sites only, reported with its own n and not as a paired test), with neither permitted to be called the other. Rule: whenever arms differ by a permission rather than by an instruction that must be executed, register reach as its own estimand and say which of the two any headline number is. RS-20260809g-device-cross §2, §4.1, §6. | S145 | S145 — and S150, where the reach the note insists on registering became the whole result. E-20260810c gave two hands the located-idiom permission at 30 Japanese narration sentences: they used it at 4 of 60 hand-sites (0.067), against 32 of 60 (0.533) for the respelling permission, with 116 of 180 revision cells byte-identical to their own baseline. The manipulation check on the located arm returned −0.0167 and 0.0000 against +0.20 — an arm that is the baseline at 93% of hand-sites cannot be more located than it — and the REG primary was withheld. Because reach had been registered as its own estimand before dispatch, the run has a licensed finding (the permission is not taken) where it would otherwise have had only a withheld one. Extension, and it is cheap: compute reach in the gap between generation and rating. Here it was computable from the frozen artifacts alone, cost $0, and named the session's result before a single rating call went out. RS-20260810c-register-reach §5. |
| (blh) | Derive the power gate from the exact test you will actually run, and state the smallest attainable P before dispatch. RS-20260808f reached P = 0.125 as the smallest value its test could return — four of eight sites were exact ties and significance was unreachable whatever the data did — and nobody knew until afterwards. E-20260809g registered the arithmetic in advance: a sign-flip permutation on k non-tied blocks bottoms out at 2/2^k, so k = 6 → 0.03125 (readable at α 0.05) and k = 5 → 0.0625 (not), and a cell with fewer than six non-tied sites is reported UNDERPOWERED rather than as a null. Both of its cells then failed that gate and the primary was withheld — which is the note working, not the note failing. Rule: a power gate stated as a round number is decoration; state it as the k at which the planned test can first reach your α, and register what a cell below it may and may not be said to show. RS-20260809g-device-cross §3, §4. | S145 | S145 — and S150, where the gate PASSED and the run was withheld anyway. E-20260810c re-derived k for a 30-site cell and registered 12 non-tied sites; it got 20, at a smallest attainable P of 1.9e-06. The power problem the note exists to catch was genuinely removed, and the comparison was still withheld — on the manipulation check. That is the note's real limit, worth carrying: a power gate answers could this design have said something, and never is the thing it measures the thing you meant. RS-20260810c-register-reach §4. |
| (blp) | A symmetric agreement rule over free-form annotator spans must fix its unit of agreement in advance; clustering the spans themselves chains. E-20260810c replaced a spine-based site vote with a symmetric one on a critic's finding, and implemented it as single-link clustering under an overlap relation. Three annotators returned 312 stretches at three different grains — one sub-clausal, two whole-sentence — and transitivity collapsed them into 26 blobs, several spanning half a paragraph, in a criterion whose own exclusion 3 forbids a site longer than one sentence. The repair, registered before any generation call and with no English in existence: compute each annotator's character coverage, split the passage into its sentences, and admit a sentence at ≥ 2 of 3. Rule: when annotators are free to choose their own spans, agreement must be scored on a unit the design fixes — a sentence, a clause, a character — never on the spans themselves, because overlaps is not transitive and single-link clustering will make it so. E-20260810c/critic.md A21. | S150 | S150 |
| (blq) | Before reusing a selection criterion on a new work, measure its admission rate on that work; a criterion built to find islands says nothing where the whole text is the island. E-20260808e's site criterion finds stretches below the neutral written register of their own language, and was built on Verga, Maupassant and Turgenev, where the low register is an island inside standard prose. Re-applied unchanged to Sōseki's 俺-narration in 「坊っちゃん」, it admitted 97 of 111 sentences — 0.874 — and 80% of the characters in the span. Nothing was wrong with the criterion; the text has no islands. The consequence had to be registered mid-run: the phrase marked narration sites was struck and the estimand narrowed to criterion-positive narration sentences of this span, because at 0.874 a "site" is not a place, it is the book. Rule: an admission rate is a property of the criterion crossed with the material, and it is one cheap count — take it before designing anything on top of the criterion, and treat a rate near 1 as a finding about the work rather than a supply of sites. RS-20260810c-register-reach §3. | S150 | S150 |
| (blu) | Two panel slugs cannot have reasoning disabled at all, so note (bkw)'s remedy has an exception and the fallback is effort: low. x-ai/grok-4.5 and google/gemini-3.6-flash both return HTTP 400, "Reasoning is mandatory for this endpoint and cannot be disabled", to a body carrying "reasoning": {"enabled": false} — measured on both slugs at S153 against openai/gpt-5.6-terra, deepseek/deepseek-v4-pro and qwen/qwen3.7-max, which all accept it. The failure is at dispatch, not a dead body, so it costs nothing and is caught immediately; but a runner that applies note (bkw) uniformly will not run at all on those two seats. Rule: apply {"enabled": false} by default and carry a named exception set dispatched at {"effort": "low"}, which is what S151 recovered on. E-20260810w-recurrence-census/run.py. | S153 | — |
| (blv) | A substring gate over an OCR witness fires on the scan, not on the seat: de-hyphenate before gating. E-20260810w's F3 required every quoted expression to be present in the passage it came from, and fired twice on two different seats quoting Garnett correctly — the Internet Archive scan carries end-of-line soft hyphens inside words ("others of still greater conse- quence") and the seats had de-hyphenated. Had the gate been applied as written, two of three cells on one hand would have been voided as fabrication, on the strongest single piece of evidence in the run. Rule: any substring or exact-match gate over an OCR witness normalises word-internal hyphen-plus-whitespace on both sides first; the repair is a materials repair, applied identically to every arm and independent of any answer, and is declared as one. RS-20260810w §6. | S153 | — |
| (blt) | A retry-on-exception path re-dispatches a request that may already have been billed, so a key-usage reconciliation should PREDICT its residual from the retry log rather than expect zero. E-20260810t reconciled to a residual of $0.0072064, not to zero, and the dispatch log explains it exactly: of seven retries, six were 429 (no body, no charge) and one was IncompleteRead(319 bytes read) — a response that began arriving, was billed upstream, never reached the client, and was re-dispatched and billed again. The figure is one source-present call at the top of the observed range. A run that reports "delta exact to zero" as a hygiene claim will read a truncated-read retry as a defect, and a run that reports the residual without attributing it has not reconciled. Fires at: every cost reconciliation on a runner that retries on exception. Remedy: count the non-429 retries and name them beside the residual. | S152 | — |
| (bmt) | A structural signal in a file is a candidate list, not a count — confirm it against the source before publishing a figure from it. Note (bmm) counted the empty speech-heads in ARM-dakghar's copy-text file and published "section ৩ has sixteen" missing speeches; it was carried forward by the arm page, the register and NEXT.md for five sessions. The whole-span refetch at S170 found fourteen. The other two are empty in the print as well — a speech-head above a centred stage direction, which the file correctly carries as its own unit — so the signal was detecting layout, not loss. The direction of the error is the benign one here, but nothing about the method guarantees that: the same test would have under-counted had the extraction dropped a speech whose head it also dropped. Fires at: any figure derived from a structural proxy in a file — empty fields, unmatched delimiters, count mismatches — that is quoted before the source has been consulted. Remedy: state such a figure as at most N candidates until the source has been read at each one, and let the confirmed count be the first thing the next session's gate produces. dakghar/collation.md §7A, ARM-dakghar S170 log. | S170 | S170 |
| (bmu) | An assertion inherited from a state page is not evidence, and a frozen artifact must not repeat one without checking it against the source. register.md §Unresolved said at span B that nobody addresses Amal as তুই in sections ১–২ of «ডাকঘর». It is false — the Headman does, at [190], in a speech this project had already translated and already argued about (note (bml) turned on the same speech). The false premise then propagated into the span-C translator's log, D40, which was written and frozen before the census that caught it, and it took an erratum (D51) to remove. The register is a hand-off convenience; it had acquired the standing of a measurement simply by being written down. Fires at: every frozen log, design or result that repeats a fact from an arm page, a register, NEXT.md or a prior result rather than from the artifact. Remedy: any inherited claim that a frozen artifact restates must be re-derived from the source in the same session, and cited to the source rather than to the state page; where the re-derivation is cheap — here it was one census over 430 units — run it before the freeze, not after. RS-20260812h-dakghar-grade §2, T-dakghar-R05-v1 D51. | S170 | S170 |
| (bmr) | A "licensed" or correspondence rate for ONE published hand must be reported per copy-text, or the page must say it was not — on this hand the same statistic is 0.8723 and 0.1034 depending only on which printed volume the transcription came from. E-20260812g censused Balmont's Poe across seven tales whose Wikisource transcriptions name four different printed volumes (Скорпион 1901 vol. I, Собр. соч. 1906 vol. VII, Скорпион 1906, Скорпион 1911 vol. III). Pooled licensed(RU) = 0.4414 — a perfectly quotable number that is an average over two populations differing by a factor of 8.4. And it is not simply lost markup: STRESS retention is 1.000 in both volumes, and an independent re-keying of one tale on az.lib.ru carries 27 of the 28 body italics the Wikisource text has, so the unlicensed marks are in both electronic witnesses. Either the translator's practice changed between volumes or their compositors differed, and a pooled figure hides the question. Fires at: every census, count or correspondence rate computed over a single translator whose texts were gathered from more than one edition or more than one electronic route — which is most of what a free corpus gives you. Remedy: record the declared copy-text per text at build time, report the statistic per copy-text as well as pooled, and where the split is large treat the pooled figure as uninterpretable rather than as the headline. Applies retrospectively as a caution, not a correction: prior single-hand rates on this shelf were not computed this way. RS-20260812g-foreign-italic §5, framework/v0.2 §7.8a. | S169 | S169 |
| (bms) | Put a same-hand-different-text reference cell in every corpus build and read it before the corpus is used — it catches page furniture that inspection does not. E-20260812g built seven Russian texts from Wikisource and looked right: sane paragraph counts, balanced delimiters, plausible span totals. The defect surfaced only when a reference cell compared the same translator on two DIFFERENT tales and returned a 64-token identical run, which no two translations of different texts can have. It was a category footer. Repairing that exposed three more classes of furniture, all of which had been counted as the translator's marks: {{right}} epigraph blocks, source-edition lines, and — worst — the editor's notes section, which glosses the source's foreign phrases IN ITALICS, i.e. exactly the spans the run was counting. One page headed that section === КОММЕНТАРИИ === in capitals and a case-sensitive rule let five of its glosses through, including one of Poe's own French sentences. Span totals moved 128 → 118 → 112 across the repairs, all before any correspondence call. Fires at: every corpus built from a wiki, an e-text library or any page that carries apparatus around the text. Remedy: build the reference cell into the corpus script, not into the analysis — one call comparing the same author or translator across two of the corpus's own texts — and require a near-zero long-run before the corpus is used; then assert that the admitted span count equals the body span count per text, which is the check that stayed green afterwards. Kin to (bmm), which says the same about silently SHORT extractions; this is the silently LONG case. RS-20260812g-foreign-italic §6.5. | S169 | S169 |
| (bmq) | Before a design classifies sites by the source's effect depends on something the target language lacks, measure whether two readers of the source agree on that classification at all — on this material they agreed at ZERO sites of fifteen. E-20260812f cut a Bulgarian comic chapter into 15 segments and asked two source-side seats, holistically and with the feature checklist struck by its pre-run critic, whether each segment's effect on a Bulgarian reader depends on something English does not have. 0 both-yes, 12 both-no, 3 split — and the three splits go in different directions, each seat naming a different segment as the untranslatable one. The segment the translator's own frozen log had spent two entries on — «Санким честна ли е? … изведнъж гут моргин», a Turkish sneer and a mangled German phrase — came back carriable from both. This is the third time this project has turned a judgment about untranslatability into a yes/no coding and the third time it has failed to converge (RS-20260804b §5 at five of seven loci; RS-20260804h's yardstick stage; here) — and the first time the category was the translator's own rather than a critic's. Fires at: any design that strata-fies sites by difficulty, untranslatability, "where the translator faced a real choice", or any near relative, from the source side. Remedy: measure the classification's inter-reader agreement in a cheap first stage and size the design to what survives; if nothing survives, the design's primary is dead before dispatch and the classification is the finding. Sibling of (bmo), which says the same about assumed floors. RS-20260812f-affect-unprompted §4. | S168 | S168 |
| (bml) | Before buying a measurement, ask what the frozen artifact already answers by inspection; if the primary's outcome is legible in the text, cancel that part of the run and report the inspection. E-20260812g reached its second frozen design with twelve translation calls and two blind-coder calls in the pre-flight. Its critic's finding 16: "the supplied frozen item text already visibly contains 'boy' at [190]; absent a very unusual coding decision, the outcome is apparent before dispatch." True — the primary was a sealed self-report about three speeches, and one of the three visibly contradicted it. The calls would have bought a majority vote on a fact anyone can read, and would have dressed a count of three as a measurement. Fires at: any design whose primary is a property of text this project already holds frozen. Remedy: before the pre-flight is written, state which predictions could be settled by reading the artifact, settle those by reading it, and let the run buy only what reading cannot give. RS-20260812c-grade-shift §6. | S165 | S165 |
| (bmm) | A text extracted into this repository can be silently SHORT, and neither the image gate nor the index lint can see it: diff the extraction against the upstream source before translating. ARM-dakghar's copy-text file was missing ten whole speeches of span B — 101 Bengali words, 4.7% of the span — one at the foot of each of ten pages, present in every case in the upstream Wikisource page text and lost by this project's own S160 extraction. The only trace was an empty speech-head, which no tool reads; the arm's declared page-image gate would have compared images against a text that had already dropped the lines. Section ৩ was forecast to have sixteen more; it had fourteen — see (bmt). Fires at: every span of every serial work whose source was extracted from a paginated upstream. Remedy: the gate runs in two stages and in this order — extraction against the upstream page text first, then the reconciled text against the page images — and the word count of the restored span is compared with what the arm page declared. RS-20260812c-grade-shift §3, dakghar/collation.md §5, §7A. | S165 | S170 |
| (bmn) | A copy-text set from a collected edition decades after first publication is a signal to go looking for the earlier witness, not a copy-text to accept — and the looking happens BEFORE translating. T-shukuhai-R06-v1 was frozen declaring one witness and asserting that no second freely reachable text existed. False: 『道理』(春陽堂, 1921), the story's first appearance in book form, is on NDL Digital Collections under a Public Domain Mark as page images, and the Aozora card names its own 底本 — a 1960 collected edition — on its face. The witness was found by a contamination search run after the translation and the whole experiment had been frozen and run. The late collation found five divergences, one of them a change of sense ([11] 興味 → 意味) landing on the exact sentence the translator's log D12 had spent an entry adjudicating between two readings of a word the author may not have written; none of the five fell on a measured site, which was luck and is not a method. Fires at: every translation limb whose copy-text is an e-text set from a modern collected edition. Remedy: read the e-text's own 底本 statement as step 1 of the copy-text gate, and search for the earliest reachable witness before the draft is begun — the same gate S160/S165 ran on 「ডাকঘর」 and S161 on 「Kaşağı」, run for the right reason and at the right time. RS-20260812d-slot-or-carrier §10, shukuhai/collation.md. | S166 | S166 |
| (bmo) | temperature 0.0 is not determinism, and a floor that assumes it is will fail — measure what one hand's own re-rendering moves before setting any detection bar. E-20260812d registered G3 at ≤ 0.60: the same hand, the same passage, two calls at temperature zero, rated 0–3 for how much the two English versions differ. It came back at 0.8333 and withheld every primary in the run. None of the six pairs was byte-identical; their similarity ran 0.0229 to 0.8511 — one pair was effectively a different translation. The same arbiters on byte-identical English scored 0 at 12 of 12 (post-hoc, unregistered), so none of the floor is arbiter false-alarm and all of it is the hand. Fires at: any design whose statistic is a rated difference between two generated texts, which is most of this project's arrival, carriage and recovery measurements. Remedy: measure the hand-variation floor in a cheap first stage, as E-20260812d already measured source-side visibility before selecting windows, and set the detection bar above the measured figure rather than at an assumed one; report every cell as a difference from that floor, not as a raw mean. Corollary: prior published arrival numbers whose floor was assumed are not thereby wrong, but their smallest reportable effect is unknown. RS-20260812d-slot-or-carrier §3. | S166 | S166 |
| (bmp) | One dispatch process per output directory, and a runner you mean to stop must be PROVEN dead before the replacement starts — a pkill that reports nothing is not a proof. E-20260812e switched a seat's reasoning configuration mid-run, killed the runner, and started a replacement. The first runner did not die. Both dispatched against the same runs/ directory for the rest of the run, racing on the same (pair, seat) files: 90 of 108 bodies on that seat came back at the OLD configuration and 18 at the new one, and the defect was invisible in the costs and in the answers — it was caught only by auditing the request bodies file by file for the parameter that had been changed. Where both processes wrote the same slot, both were billed and only the later writer's record survived: $0.407307700 with no surviving record, plus $0.074160000 of bodies deleted in the repair — 30.6% of the run's spend, buying nothing. Fires at: any session that restarts, resumes, or reconfigures a runner that writes into a directory keyed by item. Remedy, three parts: (i) confirm the process is gone (ps returning empty, not pkill returning quietly) before starting anything that writes the same directory; (ii) make the runner record its configuration in a lock or manifest file and refuse to start if a different one is live; (iii) when a configuration changes mid-run, move the superseded bodies aside — never delete them — because the raw output is the record and the accidental duplicate is a free robustness check. The 18 deleted bodies here destroyed exactly that check. RS-20260812e-dose §7.1. [FIRED S168, used correctly, and remedy (iii) needs one more clause. The runner wrote runs/_config.json on its first call and refused to start on a changed configuration; three deliberate reconfigurations went through a --reconfigure flag that preserved the previous configuration in a _history list; every kill was verified by ps returning empty rather than by pkill returning quietly. The part that failed was the move. Superseded bodies were moved into ONE FLAT directory on two separate occasions, and the second move overwrote same-named bodies from the first — 63 moved, 51 survive — destroying exactly the cost records the residual reconciliation needed. Clause: move superseded bodies into a directory named for the superseded configuration, never into one shared bin.] | S167 | S168 |
| (bmg) | When a critic amendment narrows the question a control is scored against, re-check every existing control against the narrowed question before dispatch — two amendments accepted separately in one sitting can be jointly incoherent. E-20260811f accepted the critic's SERIOUS 3 (add a control on the right axis) and its ADVISORY 6 (narrow the arbiter question from how they stand to each other to the social relationship) in the same adjudication. The narrowing made the two pre-existing controls — a change of referent and a change of tense — inapplicable, because neither changes a social relationship; G2a duly failed at 0.2500 and withheld a primary that the new control (G2b, 0.8333) had already licensed. Nothing in the disposition process compares amendments with each other. RS-20260811f-dakghar-address §4. | S160 |
| (bng) | An open-ended enumeration is the one output shape whose cap cannot be guessed: if the answer is a LIST whose length is set by the input, either bound the output form or shard the call — do not pick a number. E-20260813g's LAT call asked one seat to classify 758 word types by etymological origin and return the R-list and the O-list, on the reasoning that the lists would be short. Both attempts returned finish_reason: length at a 6,000 cap with effort pinned low, the second returning zero characters, and $0.076043850 — 34.7% of the session's spend — bought nothing. F6 struck the measure, F3 reduced the axis to one measure, and F1 then withheld half the run's registered predictions: one mis-sized cap cost an entire axis. Note (bmb)'s remedy (pin the effort and size the cap to the pinned effort's observed output) could not be applied here, because on a new task shape there is no observed output — which is exactly when this note fires. The remedies, in order of preference: (a) ask for a closed output form whose size is fixed by the input rather than by the answer (a code per item, a fixed-width table, counts), (b) shard the list across several calls with a bounded slice each, (c) if neither is possible, price a pilot on a tenth of the list before dispatching the whole. The sibling calls in the same run asked for closed-form output over the same six texts and returned 618 and 550 characters at a 4,000 cap. | S178, 2026-08-13 | fires at dispatch design, on any call whose output length is a function of the input's length. S183 — fired and NOT repaired: P4 truncated mid-finding-10 at a 6,000 cap and the verdict line never arrived; NEXT.md recorded that a second $0.11 would have bought a verdict word and it was not spent. S184 — fired, and repaired, and the repair is not the one this note prescribes for a critic: raising the cap to 8,000 did not work — the seat spent 6,085 of the 8,000 on hidden reasoning and truncated in the same place. See note (bnr), which carves the pre-run critic out of remedy (a)/(b)/(c) and prescribes a continuation call instead. |
| (bnr) | A pre-run critic is the one open-ended enumeration whose cap must NOT be raised — budget a CONTINUATION CALL instead. Note (bng)'s three remedies all assume the output can be bounded or sharded; a critic's cannot, because bounding the form is exactly what stops it finding the thing nobody anticipated. Raising the cap does not work either: S183 truncated P4 at 6,000, S184 raised it to 8,000 and truncated in the same place, the seat having spent 6,085 of the 8,000 on hidden reasoning — the visible budget was 1,915 tokens at both caps and the cap is not a bound on a reasoning seat. What worked, at $0.035: a second call carrying the design and the truncated critique, asking only for the unfinished finding, any further findings, and the verdict line. It returned finding F-G complete, five further findings and the verdict — and two of the five were accepted amendments, so the continuation was not a formality. Fires at: every pre-run critic dispatch. Remedy: size the first call at whatever is affordable, expect truncation, and hold ~20% of the critic's budget for a continuation; never let a run proceed on a critic that did not return a verdict line. RS-20260814f-carriage-elevation-2 §9. | S184 | S184 |
| (bny) | When a reasoning seat returns ZERO characters, reasoning.max_tokens is the remedy and reasoning.effort is not — and the two may not be sent together. E-20260815's pre-run critic (P4 moonshotai/kimi-k3) was dispatched at max_tokens 6,000 with reasoning: {"effort": "low"} — the pin note (bnk) prescribes — and returned finish_reason: length with content of zero characters, the whole cap spent on hidden reasoning. $0.100636, 31.3% of the run, for nothing. Note (bnr)'s remedy is a continuation call, and it was inapplicable: a continuation needs a truncated critique to continue from and there was none. The retry that worked, at $0.075334, sent reasoning: {"max_tokens": 2000} with max_tokens 9,000 and returned nine findings and the verdict line complete. Sending both keys is HTTP 400 ("Only one of reasoning.effort and reasoning.max_tokens can be specified") — free, but it costs a round trip if discovered at dispatch. Fires at: every first call to a seat that reasons, and hardest on the critic, whose output cannot be bounded by shape. Remedy, which supersedes (bnk)'s pin for this case and closes the gap (bnr) leaves: on any call whose value is the written answer, send an explicit reasoning cap — not an effort pin — sized to a fraction of a total that leaves the answer room the reasoning cannot take; keep (bnr)'s continuation budget for the case where the answer starts and stops, and use this note where the answer never starts at all. RS-20260815-register-room §10. | S188 | S188 |
| (bns) | A non-zero reconciliation residual is not evidence of an unrecorded call until a SECOND key read, with no calls in between, has been taken. S184 closed with per-request sum $0.263575750 against a key delta of $0.381083810 — an unattributed $0.117508060, 45% of the session's own spend, which on every prior session's convention reads as a defect to hunt. A second GET /api/v1/key immediately afterwards, with nothing dispatched between the two reads, returned a figure $0.001734720 higher: the key was still settling. NEXT.md had already observed the behaviour once in prose (S181 $1.109046 plus $0.265127 settling late) without a way to test for it. Fires at: every session close that reconciles. Remedy: when the residual is non-zero, take the second read before writing anything about it, and report the movement-with-no-calls figure beside the residual; CLAUDE.md's rule that per-request cost is primary and the delta is the sanity check is what is ledgered either way. The check costs one curl and distinguishes settlement lag from a call nobody recorded — which are different defects with different repairs. RS-20260814f-carriage-elevation-2 §9. | S184 | S184 |
| (bnz) | A call in flight when you kill a runner is billed and lost — note (bnx)'s append-and-resume protects everything that RETURNED, not the one in the air. E-20260815b's judging runner was stopped four times: twice by a harness wall clock and twice deliberately, to hold back a gate and protect the write-up. (bnx) worked exactly as written — 110 of 110 cells survived every interruption and not one had to be re-bought. The residual was $0.004922400, 0.7% of the spend, and note (bns)'s second key read returned a figure identical to the first, so it is not settlement lag: it is the sixth QD request, in flight at the moment of the kill, billed by the provider and never written to disk. Fires at: every session that stops a dispatch loop by signal rather than by letting it finish — which, on a long serial judging run, is most of them. Remedy: none is worth building; the figure is one call's cost and the repair (a shutdown flag the loop checks between calls) would trade a bounded, known loss for a runner that cannot be stopped promptly. What this note is for is the reconciliation: a residual of roughly one call's cost, after a kill, needs no hunt — record it as the kill's price, name the stage it belongs to, and do not spend a session looking for a defect that is a signal arriving mid-request. RS-20260815b-fluent-carriage §9, config/budget.md S189. | S189 | S189 |
| (boa) | A verifier proves the arms differ where they should and nowhere else; it cannot prove they say the same thing — put the content-parity control BEFORE the judging calls, not after them. E-20260815b built three arms whose differences were mechanically exact: 342 pre-run checks, 0 failures, 8 of 8 mutations caught, asserting that the carriage manipulation and the oddity manipulation touched disjoint paragraph sets and nothing else. Every one of those checks passed and the run's carriage half is still unreadable, because an independent seat — after catching 3 of 3 planted errors — returned SAME on 0 of 7 segments. Three of the seven declared form-only device classes turned out not to be: a suspension cannot be flattened without inventing its completion; a reduplication reads as a different claim; varying a held keyword changes the degree asserted. Two of thirty-one oddity edits were also propositional defects the author did not catch — one reversed an agency — and no mechanical check could have seen either, because no check tests meaning. The parity control was dispatched last, so it cost the whole run instead of one call. Fires at: every design whose primary contrasts two texts asserted to differ in form only — which is this project's commonest shape. Remedy: dispatch the parity control first, on the frozen arms, as a gate; a design that cannot pass it has no primary and should learn that for one call's price. Corollary that travels further than this run: R14's operator table has never been put to a parity control of this kind, and RS-20260808d's sixteen-device flattening is the figure most exposed to it. RS-20260815b-fluent-carriage §6. | S189 | S189 |
| (bob) | A "first N by rule" locus selection is only as good as the enumeration behind it, and an enumeration is not an enumeration until its exclusions are published. E-20260815's translator's log declared eight PATTERNED loci as the first eight in copy-text order under a two-part criterion, and named a ninth as the last. Both claims were false: at least two further criterion-(b) loci exist in the same 1,036 words, one of them earlier in the text than the second locus taken, so "the first eight" never described the set. Nothing was wrong with the rule; what was missing was the rejection list. The pre-run critic did not find this directly — it found that the domain pooled sound with grammar (BLOCKING 1), and acting on that finding forced a re-enumeration, which is what exposed it. The repair that works is cheap and was applied the same session: the small domain (SOUND, 4 loci) is claimed exhaustive and prints every candidate it rejected, with the reason; the large one (COGNATE, 5) drops the completeness claim and says so. A reader can now test the first and cannot be misled by the second. Fires at: every design that selects loci by "all X, take the first N" — which is this project's standard shape for census work, used at E-20260814 and E-20260814g before this. Remedy: a selection rule that claims exhaustiveness must publish its exclusion list in the frozen design; where the list would be long, drop the exhaustiveness claim instead of asserting it untested. Corollary: an author enumerating figures in his own translation is the worst-placed reader for the ones he did not notice while translating — the misses here are all in sentences the log discusses for other reasons. RS-20260815-supplied-sound-confirm §9.5, erratum E3 on T-alf-layla-R05-v1. | S190 | S190 |
| (boc) | A confirmatory primary registered as a strict inequality can PASS while the effect it stands for collapses — register a magnitude, or the replication cannot fail. E-20260815 was built to confirm RS-20260814g's Burton supplies English sound where the Arabic has none, at 4 of 8 plain loci. Its pre-run critic rightly demoted the rate threshold (0.375) as circular — it was mined from the run being confirmed — and the primary became the between-hand contrast Burton > Lane. That primary HOLDS, at 1 of 8 against 0 of 8, and the effect fell from 0.500 to 0.125. One cell satisfies a strict inequality; on n = 8 with a comparator at floor, the primary was close to unfalsifiable in the direction that mattered, and the honest reading had to be assembled afterwards from the secondary figures and from where the surviving cell sat (in the confounded speech stratum, with the unconfounded slice at 0–0). Fires at: every replication whose comparator is at or near zero, which is common here because the null hand is usually the flat one. Remedy: when a primary is a contrast against a floor, register both a direction and a magnitude — e.g. > Lane and within 0.25 of the original point estimate — and pre-commit to reporting the replication as failed if the magnitude misses even where the direction holds. The circularity the critic named is real and the answer is not to drop the magnitude but to source it from outside the run being confirmed, or to declare the study a direction-only replication in its title. RS-20260815-supplied-sound-confirm §5. | S190 | S190 |
| (boe) | A re-dispatch ladder that overwrites the failed response's usage under-reports the run's cost by exactly the attempts it replaced — accumulate every attempt, not the one that worked. E-20260816 closed with per-request costs summing to $1.022045450 against a key delta of $1.047809700, the key reading $0.025764250 high, and note (bns)'s second read returned an identical figure, so it was not settlement lag. The cause is one line of the runner: on an unusable reply it re-calls and rebinds r, so only the second body's usage is written to run.jsonl. Ten bodies were re-dispatched and every one of them had failed by truncation, which means the discarded attempt had burned its whole max_tokens and was the most expensive body of its stage, not an average one: 9 P1 at cap + 1 P2 at cap ≈ $0.029 against a measured residual of $0.0258. Fires at: every runner in this project — they all share this dispatch shape, and the defect has been invisible because a residual of two or three cents reads as settlement. Remedy, one line: keep a per-body list of attempt costs and write their sum, with the attempt count, so the ledger is the money spent rather than the money spent usefully. Distinguish from (bnz), which is about a call in flight when a runner is killed — that one returned nothing and is unrecoverable; this one returned, was billed, and was discarded on purpose by code that meant to keep the good answer. RS-20260816-answering-figure §5; config/budget.md S191. | S191 | S191 |
| (bof) | The key-usage delta is not a cross-check on this project's spend, and must not be reported as one. E-20260815c closed with per-request costs of $0.542027300 against a key delta of $1.008081890. This is not settlement lag (note (bns)) and it is not note (boe)'s runner leak, which is fixed in that runner and ledgered inside the total: two key reads twenty seconds apart, with no request from the session in between, differ by $0.005285, and over a 60.9-second idle window the figure moves $0.013748 — about $0.81 an hour while nothing is being dispatched. Something outside these sessions is drawing on the same key continuously. That retro-explains the movements S190 ($0.327797), S191 ($0.253922) and S192 ($0.306086) each opened on and each recorded as unexplained, and it means every past reconciliation that "closed exactly" did so because the drift happened to be small over that window, not because the delta is an instrument. Fires at: every §Spend section and every config/budget.md reconciliation. Remedy: per-request billed cost from "usage": {"include": true} is the ledger and is exact; a key read may be recorded as provenance but no residual against it is a defect to be explained, and no session should spend its §Spend section explaining one. Supersedes the reconciliation practice of notes (bns) and (boe) without touching either note's own finding. RS-20260815c-footing-direction §9. | S192 | S221 |
| (bog) | An instrument control for "what does the WORDING mark" must hold the narrated EVENTS flat, or it is not a control. E-20260815c's negative control removed every rank word from a scene and kept the scene: a household standing back, a younger woman curtsying, the hostess coming forward to thank the visitor. All three seats named the same person as socially highest in the unmarked version as in the marked one, and all three quoted the behaviour, not a word — "the household stood back and the hostess came forward and thanked her". The gate fired and every registered primary of the run was withheld. A second pair, same three women, no deferential behaviour in either member, separated at the ends of the scale: REL 1, 1, 1 with WHO E, E, E unmarked, against REL 7, 6, 6 with WHO B, B, B marked. Fires at: any design asking a judge what the wording does while showing them a story. The trap is that the natural way to build a deference control is to write a deference scene. Remedy: build the control on an incident in which nobody defers to anybody — people sitting and talking — and vary only the words; then, separately, run the deferential-events version to measure how much the events swamp. The two together are the measurement, and in this case they gave the run its only unwithheld finding: with events flat, four rank words are worth 5.33 scale points and a named person; with events deferential, the same four words are worth 2.33 points and change the named person not at all. RS-20260815c-footing-direction §2–§3. | S192 | S192 |
| (bph) | A content check that a response "looks right" does not catch a response that was cut off — assert finish_reason explicitly, and set the reasoning budget well below the content cap. E-20260815d set reasoning: {"max_tokens": 400–600} inside a content max_tokens of 400–500 for its short stages, so the reasoning budget consumed the whole allowance and 21 of 120 bodies came back with finish_reason: "length", cut off mid-sentence. Every one passed the run's content check, because a truncated Japanese sentence still contains four Japanese characters, and one of them (「「いや、決して間違いなどあってはならない。それは必ずやそこにある」) was missing the polite ending that would have changed its coded value. Note (bny)'s working shape is a ratio — 2,000 inside 9,000 — and this session copied the numbers without the ratio. Fires at: every dispatch with a max_tokens under about 2,000, which is every per-item stage. Remedy, both halves: (i) keep reasoning.max_tokens at roughly a fifth of the content cap and never at or above it; (ii) make finish_reason != "length" a kept-ness condition in the runner and an assertion in the verifier, so a truncated body is re-bought automatically and can never reach the analysis. Re-purchase is selected on finish_reason, never on what the body says. RS-20260815d-supplied-footing §8. | S193 | S198 — fired four more times in one run, on two seats and two stages; see (bps), which is the advance check that makes it preventable. |
| (bpv) | max_tokens must be strictly GREATER than the reasoning cap, and two seats fail in opposite directions when it is not — one with an HTTP 400, one with a silent empty body. E-20260816f sent max_tokens: 900 with reasoning: {max_tokens: 1500}. qwen/qwen3.7-max refused the request outright — HTTP 400, "max_completion_tokens [200] must be greater than thinking_budget [400]" — which is the benign failure, because it costs nothing and is impossible to miss. openai/gpt-5.6-terra accepted it, billed it, and returned finish_reason: length with an EMPTY content string, twice, at 600/1500 and again at 1400/700 on a short prompt. Fires at: every dispatch that sets a reasoning budget, which by (bph) and (bps) is every dispatch. Remedy: make the two caps a ratio and assert it in the runner before the first call — this run settled at 1,500/900 for the long bodies and 1,400/700 for the short ones, and all four seats were clean. Kin to (bph), which is the truncation seen from the content end, and (bps), which is the per-seat probe; this note is the one line of arithmetic that makes both cheap. RS-20260816f-night-seam §4, config/budget.md S200. | S200 | S200 |
| (bpx) | temperature: 0 does not make a second call a second reader, and a design that counts replicates as independent bodies inflates its n by the replicate factor. E-20260816g's v1 planned three temperature-0 replicates per seat and treated the bodies as independent observations; the pre-run critic's BLOCKING 3 named it — "repeated deterministic calls are pseudoreplicates, while model families are fixed, heterogeneous instruments rather than a random sample of readers". Served models are not bit-deterministic, so the bodies do differ; the point is that the variation is not a sample from anything the design has defined, and the four or five seats are the instruments. Fires at: every design that buys more than one body from the same seat on the same stimulus. Remedy, two lines: set temperature above 0 so replicates are genuine samples, and make the seat, not the body, the unit of the estimator — average within a seat first, then across seats, so a seat contributes once however many bodies it supplied. E-20260816g did both; its Δ is the mean over three seats of each seat's own FULL-minus-FLAT, and its two permutation nulls are over labels and loci, never over bodies. Flagged, not adjudicated: earlier runs whose exact tests permute bodies — RS-20260816f's placement P over 3,432 relabelings is the nearest — bought temperature-0 replicates, and whether that inflates their nulls is a question for whoever next needs those figures, not a claim here that they are wrong. | S201 | S201 |
| (bpi) | Two frontier models translating the same passage independently are not two observations, and the dependence check will say so if it is pointed at them. E-20260815d measured openai/gpt-5.6-terra against google/gemini-3.6-flash on 501 words rendered EN→JA from identical prompts, no shared context: 69 shared 7-grams, 35 twelve-grams, 26 fifteen-grams, longest contiguous run 33 characters — more long n-grams and a longer run than the run's known dependent pair, a 1930 published translation against its own 2021 revision (80 / 23 / 11, run 18). Both models were clean against the published hand (runs of 8 and 11). Fires at: any design whose criterion is "k of n independent model hands", which this project has written many times. Remedy: run dependence_check_cjk.py / dependence_check.py between the model hands, not only between each hand and the human comparator, and report a "2 of 2 hands" criterion as what the measurement says it is. This does not invalidate model panels used as judges of distinct objects; it bites when the models are being counted as independent producers. RS-20260815d-supplied-footing §7.1, A-mikami-silver-blaze §Contamination. | S193 | S193 |
| (bpj) | A form-only device set is a hypothesis about meaning, and only an independent reader can test it — this project has now had five of seven classes fail that test on one story. E-20260815b declared seven Class A "form-only" device classes on Makino and its parity control convicted three: a suspension cannot be flattened without inventing its completion, a reduplication asserts more than a single token, and varying a held keyword moves the degree asserted. E-20260815e dropped those three and the same control convicted a fourth — folding a standalone attribution into a dialogue tag forced the flattener to invent an addressee at one site and a speech verb at another, because the source attributes by juxtaposition. A blind carriage audit then refused a fifth: the reduplicated mimetic fails as carriage at 5 of 6 sites, against the translating log's claim that English sound-symbolic verbs do the work and "cost nothing". Only utterance-first, attribution-after transferred entire. Fires at: any design whose arms are asserted to differ in form only — this project's commonest shape — and at any R14 table. Remedy: declare the class set, then have an independent reader price it before the figures are read; report a translator's carriage census as a claim (this one was priced at 25 of 33) and never as a measurement. Note (boa) says when to run the control; this one says what it keeps finding. RS-20260815e-fluent-carriage-2 §3–§4. | S194 | S194 |
| (bpl) | When the lead is primed on a specific published rendering of the very phrase it must translate, write the primed solution down BY NAME and exclude it — an invisible contamination becomes a stated constraint, and the excluded thing stays checkable. D43 had to render the Nights' night-closing formula, a rhymed sentence repeated about a thousand times, and the lead had seen Burton's rhyme (day / say) before span A was written. Every rhyme the desk then produced had to be checked against a solution already in the room. What settled it was naming the excluded pair in the log, listing the three further candidates refused and why, and recording the Arabic end-rhyme as lost with a compensating figure at the same seam rather than reaching for the known answer. Fires at: any R05-style source-only rendering of a phrase from a heavily translated work — the formulae, the proverbs, the titles — and at any artifact carrying contamination: high. Remedy: name the primed rendering in the translator's log before choosing, exclude it explicitly, and publish the rejected candidates; a contamination: line that says only high leaves a reader unable to tell exclusion from imitation. Kin to the standing contamination rule in CLAUDE.md, which governs measurement; this governs the decision when a measurement is not what is at stake. T-alf-layla-R05-v1 D43, register.md V12. | S195 | S195 |
| (bpm) | Counting a phrase with a literal-space regex under-counts every wrapped text, and the miss is silent and directional. E-20260815f first counted Burton's night formula at 25 by grep on the raw Project Gutenberg file, and at 33 over a whitespace-normalised token stream. The eight missing instances are the ones where the line break falls inside the phrase. Twenty-five was believed long enough to be written into a working note, and it is the retention figure — so the error would have published a translator as dropping a quarter of his boundaries. Fires at: every census over PG, Wikisource, Aozora or any hard-wrapped source, which is all of them. Remedy: normalise whitespace or tokenise before counting, and make the independent verifier use a \s+-joined pattern; where a count is load-bearing, have the two paths agree. RS-20260815f-night-formula §7.5, verify.py. | S195 | S195 |
| (bpk) | A source-blind "which followed the original?" score is not evidence about carriage, and the same seats answer differently when they can see the original. E-20260815e is the first carriage manipulation here whose arms passed an independent content-parity control (SAME on 7 of 7, 5 of 5 planted errors caught). On it: an arm carrying nothing and merely badly written was chosen as the source-follower over a clean flattening in 20 of 21 cells; the cell where carriage and non-fluency point at opposite arms came back indeterminate for the second time in two builds (13/21 then 11/21, seats disagreeing both times); and the same pair with the original printed above it went to the carrying arm in 20 of 21, both reversing seats moving the same way (2/7 → 7/7, 3/7 → 7/7). Fires at: any use of perceived-source-carriage, and at any evaluation that asks a source-blind reader an attribution question. Remedy: score this sense with the source available, or report the figure as a fluency judgment; a source-blind carriage score may not be cited as evidence that anything was carried. Corollary for oddity arms: on a source whose content names its language, no English-side awkwardness could be built that an independent seat did not attribute to that language — 25 of 35 edits, and retiring the operator a critic named moved it not at all. Stilted collocation was the one operator that escaped (4 of 12 against nominalisation's 11 of 12, circumlocution's 7 of 8 and cleft's 3 of 3): if a design needs unlicensed markedness that does not leak the source language, make it lexical, not syntactic. RS-20260815e-fluent-carriage-2 §7–§8. | S194 | S194 |
| (bpn) | If the primary is "does X do anything", DERIVE the second arm from the first — an independently written contrast arm cannot carry the attribution, and a critic shown both will say so before you spend a dollar finding out. E-20260816b first wrote its two arms separately: same declared device classes, same locus set, word cost matched to ±2. The pre-run critic returned NEEDS REDESIGN with the whole point as BLOCKING 1 — two independently written arms differ in propositions, member count and rhetorical form as well as in the manipulated property, so their difference is not attributable. The rebuild made the second arm a de-sounded twin of the first: same added material, same member count, same word count at every locus, token-identical except at 15 positions across thirteen loci, one at eleven of them, and a builder check that asserts it. The primary then fired at +0.3846 and could be read, because the seats' own free-text reasons name the changed word. Fires at: every design contrasting two versions of one text — this project's commonest shape, and the shape of R14, R31/R32, R34/R35 and every arm table since S133. Remedy: build arm B by editing arm A at declared slots, never by writing it; make the token-level diff a builder assertion with a hard maximum; and record per locus the material both arms add, so a residual rule-1 violation is common to both and cannot generate the effect. Corollary the same run establishes: exact synonymy does not exist, so a one-word substitution always changes sense a little — say so as a limit rather than claiming equivalence, and let the seats' reasons show what they answered on. RS-20260816b-invented-figure §3, §10.2; E-20260816b §11. | S196 | S196 |
| (bpo) | A contrast arm denied the resource the real move uses measures the resource, not the move — and this project published such a reading and had to withdraw it one session later. RS-20260816 built its invention arm under a substitution only, no added words rule, measured it inaudible at 0 of 13, and wrote into framework/v0.2 §7.12 that ornament cannot be invented, with the length restriction named as a caveat. E-20260816b gave the same hand, the same loci and the same device classes the words an ornamentalist actually takes and the same instrument heard the source-inference at 9 of 13 — indistinguishable from genuine compensation at 8 of 13. The starved arm had measured the starvation. Fires at: every decoy, control or contrast arm built under a constraint the genuine move does not obey — length caps, substitution-only rules, single-site budgets. Remedy, in the order that works: measure the move's profile in the wild first, on continuous prose, and build the arm to that profile; here the translation limb rendered a whole fresh chapter under the permissive regime and counted 22 ornamented sites, +6.1% in length, +1.55 words per site before a single arm locus was written, and the arm was then funded above that rate so that a null would be conservative. Where the constraint cannot be lifted, the finding is about the constrained condition and the title must say so. RS-20260816b-invented-figure §7; T-kalila-nasik-R34-v1 §4. | S196 | S196 |
| (bpp) | A control carried over from an earlier design must be re-validated against the NEW prompt, not merely against the old result. E-20260816c imported the four third-party Q1 instrument controls of E-20260815 verbatim, as RS-20260816b had, and they had been unanimous for three sittings. But Q1 had been rewritten for this design to say "Look only at the span marked ⟪ ⟫", and the control passages carry no marked span. CTL.POS came back N from P1 — reason: "N; no marked span is present" — the abort gate fired, and the run stopped after 12 bodies and $0.042105 spent on an ill-formed stimulus. The control was right, the prompt was right, and the pairing was wrong. Fires at: every design that reuses an instrument while changing the prompt around it, which is what reusing an instrument means. Remedy: the prompt's presuppositions are part of the cell contract and belong in the same assertion block as the cell counts — here one line, exactly one ⟪…⟫ pair in every dispatched string, now in cells.py, would have caught it before dispatch. RS-20260816c-checked-ornament §8. | S197 | S198 — FIRED AGAIN AND GENERALISES BEYOND CONTROLS TO MECHANICAL PARSE RULES. E-20260816d carried E-20260815d's accept-rule for a Japanese body verbatim — at least 4 CJK characters — into a census whose shortest turns are \"Uncle!\" and \"Nothing!\", whose correct Japanese is three characters. It rejected a perfectly good 「おじ様!」. Anything carried from an earlier design — a control, a prompt, a parser, a threshold, an accept-rule — is re-validated against the NEW material's extremes before dispatch, not just against its typical case. |
| (bpq) | A stop-loss set EQUAL to the worst-case estimate is not a margin, it is the estimate again — and it destroys whatever the dispatch order put last. E-20260816c §11 built a worst case of $2.40 from max_tokens, correctly per note (abc), and then set the runner's ceiling to $2.40. Those being the same number, the run halts exactly when the estimate is right rather than when it is wrong, and the arms dispatched last — which carried a registered primary — are what a correct-but-tight estimate would have cost. Fires at: every run with a stop-loss. Remedy: put the stop-loss above the worst case, at whatever the day's headroom allows, because its job is to catch an estimate that is wrong — a routing surprise, a runaway cap, a re-dispatch storm. Where headroom will not allow that, dispatch the registered primary first and the descriptive arms last; and where the order is instead a gate order, as here, say in the design that the two criteria conflict and which one won. E-20260816c §10, §11. | S197 | S198 — FIRED AGAIN, FROM THE OTHER SIDE, AND THE RULE IS NOW TWO-SIDED. (bpq) said a stop-loss EQUAL to the worst case is not a margin. E-20260816d v1 over-corrected and set it BELOW the worst case ($1.25 against $1.45), and the pre-run critic's F13 caught what that means: the design licensed its own truncation — a healthy run could be halted mid-census and the halt would look like a defect in the data rather than in the plan. The correct structure is worst case < stop-loss < ceiling: the stop-loss fires only when something is actually wrong, and is still a margin. Where the true worst case will not fit under an affordable stop-loss, the design is reduced until it does — tighter caps, fewer cells — and the reduction is written down. v2 landed at $1.97 < $2.05 < $2.20. |
| (bpr) | A gate that withholds a whole family of predictions should withhold only the ones it actually threatens — check, before freezing, which of them are between-group comparisons and which are within-item contrasts. E-20260816c gated P2, P3 and P4 on a manipulation check that the panel's labelling agreed with the design's. It failed, and all three were withheld — but only P2 is a between-stratum discrimination that needs the labels to be right. P3 and P4 are within-site contrasts (the same source locus, the same marked span, with and without the English device) and are well-posed under any labelling; the gate cost them for nothing, and the rule could not be relaxed afterwards without relaxing a rule after seeing it bind. Fires at: every design with a manipulation check in front of more than one prediction. Remedy: in the design, list under each prediction whether it needs the gate and why; gate only those. RS-20260816c-checked-ornament §10.3. | S197 | S197 |
| (brt) | A max_tokens cap verified on one task shape does not transfer to another, and the failure is silent — an empty body billed at the cap. Note (brr) fixed qwen/qwen3.7-max (QR) at 3000 and recorded that at 3000 it ran 56 calls without a single failure. E-20260826c sent the same seat the same cap on a different task shape — read a 130-word passage, return a short JSON list — and 29 of its 68 calls came back with an empty content string and finish_reason: "length", 44% of the seat, while P1 and P2 at 2500 lost nothing. The seat spends the whole allowance on hidden reasoning when the visible answer is tiny, which is the opposite of the intuition that a short answer needs a small cap. Fires at: every reuse of a per-seat cap across designs, which is what a note like (brr) invites. Remedy: treat a cap as verified for the task shape it was measured on, and on any new shape either probe one call per seat before the fan-out or budget at twice the previous cap; and make finish_reason != "length" a re-dispatch condition in the runner, which (bph) already requires and which here converted a silent loss into a counted one. Kin to (bpp) — anything carried from an earlier design is re-validated against the new material. RS-20260826c-register-cost §3. | S225 | S229 |
| (bru) | A quote-what-you-would-query instrument scores asymmetrically if the stored target is longer than the phrase a reader would naturally quote — score by LOCATION, not by string. E-20260826c froze its hit rule as the returned query contains the target string. On the seeded control SC4 all six seats quoted the planted defect — "three fish", "three" — and all six scored zero, because the stored target was the longer there were three fish. A one-word target (screwdriver) is hit by any quotation containing it; a four-word target is missed by every quotation shorter than itself, so the same rule is strict or lax depending on how the analyst happened to type the target. The defect inverted the run's instrument gate: 15 of 22 blatant seeds under the frozen rule, 22 of 22 under location scoring. Fires at: every design whose outcome is a free-text quotation matched against an expected string — this project's whole family of name the place where… instruments. Remedy: locate each returned quotation in the passage it came from and score overlap with the manipulated span (all 308 quotations in this run were found verbatim, so no judgment was needed); or, where that is not possible, store the minimal distinctive string. Report both scorings when the correction is made after the run, as RS-20260826c-register-cost §5 does. | S225 | S225 || (brv) | A locus pool built by hand will violate its own admission rule, and two independent critics will both find it — build the pool with a script and publish a disposition for EVERY candidate. E-20260827-declared-play froze 23 تجنیس loci hand-picked from the Persian under a written phonemic rule. Both pre-run seats returned BLOCKING on it independently: P2 showed two admitted loci that fail the stated distance test while a refused one passes it; P1 named a block (p041) that satisfies the rule and carries no disposition at all, which falsified the page's claim that the rule was applied once, block by block. The rebuild — enumerate_loci.py emits candidates under a rule fixed before it is run, adjudicate.py records a coded disposition for every candidate and fails if one is missing or names a candidate that does not exist — cost the lead four of its own favourite loci and added six it had missed or wrongly demoted, including one the hand-built page had excluded by name. Fires at: every design whose primary is what happens at the places where the source does X — this project's commonest shape. Remedy: generate mechanically, adjudicate in the open, publish the refusals beside the admissions, and state where lead judgment enters and what it cannot see. The blind spot is then a declared property of the rule rather than an undisclosed property of the analyst — and §6 of the result page is what one blind spot cost: a real locus sat in the control pool, and the declarer's one kept promise was scored against him. RS-20260827-declared-play §2, §6; critic-response.md §2. | S226 | S226 |
| (brw) | A runner that retries a failed call must accumulate the cost of EVERY attempt, or the ledger under-counts by exactly the discarded ones. E-20260827-declared-play's run.py recorded only the last attempt's usage.cost per cell. 17 cells needed a second attempt and then a third at a repaired cap; the per-request sum came to $1.541667931 against a key-usage delta of $1.747847221, and the $0.206179290 gap is the discarded attempts, not drift — which is why this is the first session able to explain a divergence rather than record it (contrast note (bof) and the S225 row). Fires at: every runner with a retry loop, which is all of them since (bph). Remedy: sum costs across attempts inside the cell record, and keep a separate attempts count; where the defect is found after the fact, ledger the key delta, which is the larger and therefore the conservative figure. RS-20260827-declared-play §9. | S226 | S229 |
| (brx) | At temperature 0 a malformed reply is REPRODUCIBLE, so the mandated single re-dispatch buys nothing — a formatting failure needs a tolerant parser, not a retry. E-20260827b's eight void cells are one defect: openai/gpt-5.6-terra returns {"choice":"A","why":"… “quoted phrase.} — an unterminated final string, because it opens a curly quote inside why and never closes either — and every one of the eight was re-dispatched once at double the cap and came back byte-identical, which is what temperature 0 is for. The choices were all plainly recoverable from the raw text. Fires at: every design whose F4 missingness rule is re-dispatch once, which is all of them since (bph); the retry rule was written for truncation at finish_reason: length, where the cap change does something, and it does nothing for a parse failure at finish_reason: stop. Remedy: branch the re-dispatch on finish_reason — retry a length, and for a stop that will not parse either repair the string (close an unterminated final value) or re-ask with a changed prompt; and where the repair is made after the run, report the registered figure and the repaired one side by side, as RS-20260827b-shown-or-told §7 does. | S227 | S232 |
Discharged
| id | note | discharged by |
|---|---|---|
| (n) | Critic dispositions are not carried forward — three of S020's 21 blockers were regressions of previously accepted fixes. | S021: the standing 13-item critic-disposition checklist in workshop/experiments/README.md, read before freezing any design. |
| (bmd) | When a manipulation's whole risk is that it smuggles content in under a formal label, ask the control the POSITIVE question as well as the negative one, and on a second seat. E-20260811c's FLAT arm was supposed to change only form. Its equivalence control asked do these assert the same things? ignore style, punctuation and emphasis — and that instruction suppresses exactly the defect connective explicitation introduces, which is why the pre-run critic called the gate unable to see what it existed for. A second control was added on a second seat asking the opposite: what relation does B state that A leaves implicit?, with a count and a list. The two instruments converged: the equivalence seat certified 1 of 6 CARRY×FLAT pairs, and the one it certified (S3) is the one the audit scored lowest (1 addition against 10, 2, 2, 3, 4). Neither alone would have been believed; together they located the clean segment and fired the run's F3. The negative form of an equivalence question, plus an instruction to ignore form, is a control that cannot fail. RS-20260811c-source-beliefs §3. |
S158 |
| (bmc) | A manipulation that REMOVES a marking must be audited for what it INSERTS, and the natural way to write it inserts exactly the thing the primary is about. E-20260811b muted source-culture realia to ask whether the world leaks into a judgement of the English. The first build replaced Akakiy with Andrew, Kantaro with Kenneth, 趙七爺 with Mr Sharpe, Yamashiro-Ya with Bellwood's — every one an Anglophone name, in a run whose primary is does the world shift the verdict toward BRITISH or AMERICAN. The pre-run critic called it BLOCKING and was right: a positive P1 would have been unattributable, and the fix was not subtle — every name in the muted column became a role description (the singer, the clerk, the gentleman, a pawn shop, the office), so that no proper noun of any nationality enters a muted form. Same shape as note (bhy), which caught a gloss handing the seat the answer; this is the removal case. Rule: for every substitution in an ablation, name the property being removed and check the replacement against that same property — a "generic" chosen by a native speaker of the target language is not neutral, it is domestic. Corollary the run then measured: the leak that did occur ran entirely through the translator's Anglicisations — halfpennies, smock, councillor — and not through one retained foreign word, so an ablation that removed only the obviously foreign items would have removed the wrong thing. RS-20260811b-realia-channel §4, critic.md F05/F08/A8. |
S157 |