Repository path: config/archive/budget-2026-07-23-to-2026-09-04.md · rendered 2026-09-09
Page metadata (front matter)
| type | ledger |
|---|---|
| id | budget-archive-2026-07-23-to-2026-09-04 |
| status | active |
| created | 2026-07-23 |
| updated | 2026-09-03 |
Budget
Cap: USD 5.00 per calendar day (UTC), in OpenRouter billed cost, all sessions that day combined. Soft cap, self-enforced, no rollover; a ceiling, not a target — but chronic underfunding of load-bearing lines (above all jury calibration) is itself a defect. Runs that don't fit today's headroom are split, scaled down, or deferred (note the deferral in NEXT.md).
Method. Before any API run: check today's rows below, snapshot key usage (GET https://openrouter.ai/api/v1/key, field data.usage), write the pre-flight estimate. After: record the API-returned actual ("usage": {"include": true} → usage.cost per response, summed), and cross-check against the key-usage delta. The key is also used outside this project, so per-request costs are primary; the delta is a sanity bound.
One-time increases (charter §6): REQUEST TO TOM block at the top of NEXT.md plus a matching page in wiki/decisions/open/, stating task, amount, why the cap cannot accommodate it, and the consequence of declining. Active only when Tom's written approval appears in the repo; record any grant verbatim here with scope and expiry. Never spend against an unapproved request.
Grants
(none)
Key-usage snapshots
| timestamp (UTC) | all-time usage (USD) | context |
|---|---|---|
| 2026-08-26 (S223 start) | 144.547740451 | S223 session start, before any project spend. +0.009646800 above S222's closing snapshot — ordinary between-session non-project drift, note (abf). |
| 2026-08-26 (S223 end) | 144.764653501 | S223 after the operationalisation gate and the pre-run critic, nine calls. Delta 0.216913050 against a per-response sum of 0.216913050 — exact to 1e-9. Per-request costs are ledgered. |
| 2026-08-26 (S224 start) | 144.764653501 | S224 session start; identical to S223's closing snapshot — no between-session drift on this pair of sessions. |
| 2026-08-26 (S224 end) | 145.604695426 | S224 after the critic, the control and 168 judging calls. Delta 0.840041925 against a per-response sum of 0.840041925 — exact to 1e-11. |
| 2026-08-26 (S225 start) | 145.871654426 | S225 session start, before any project spend. +0.266959000 above S224's closing snapshot — non-project drift, thirty times the ordinary +0.009 of note (abf) and the largest between-session figure since S163's +1.07. Recorded, not explained; no project row accounts for it. |
| 2026-08-26 (S225 end) | 148.342551051 | S225 after the critic and the 204-call run. Delta 2.470896625 against a per-response sum of 1.969541175 — the delta is larger by 0.501355450, the non-project direction and by far the largest in-session divergence this project has recorded. Per-request costs are ledgered and are primary (config/budget.md §Method); at this magnitude the key delta is not usable as a cross-check of the session total, and it is the second figure on the same UTC day to run that way. |
| 2026-08-27 (S226 start) | 148.342551051 | S226 session start, before any project spend. Identical to S225's closing snapshot — no between-session drift. New UTC day, $0.00 of $5.00 spent. |
| 2026-08-27 (S226 end) | 150.090398272 | S226 after the critic, two pilots and the 240-call run with 17 re-dispatches. Delta 1.747847221 against a per-response sum of 1.541667931 — larger by 0.206179290, and this time the cause is identified rather than recorded: run.py stored only the last attempt's cost per cell, so the discarded first attempts of the 17 re-dispatched cells are missing from the per-request sum, note (brw). The key delta is ledgered, as the larger and therefore conservative figure. |
| 2026-08-27 (S227 start) | 150.499214772 | S227 session start, before any project spend. +0.408816500 above S226's closing snapshot — non-project spend on the same key, forty times the ordinary +0.009 of note (abf) and the second-largest between-session figure this month. Recorded, not explained; no project row accounts for it. |
| 2026-08-27 (S227 end) | 152.075441822 | S227 after the critic, the cap probe and 420 preference calls. Delta 1.576227050 against a per-response sum of 1.631771425 — the per-request sum is larger by 0.055544375, the opposite of the usual direction and the first time this project has seen it. Per-request costs are primary (§Method) and are also the larger figure here, so the per-request sum is what is ledgered: note (brw)'s conservative rule applied in the direction the numbers actually fell. run.py accumulated cost across every attempt, so the S226 defect cannot explain this one. |
| 2026-08-28 (S230 start) | 153.782128547 | S230 session start, before any project spend. Identical to S229's closing snapshot — no non-project spend on the key between the two sessions, the first zero between-session figure this month against S227's +0.408816500 and S229's +0.266938700. |
| 2026-08-28 (S230 end) | 155.314429272 | S230 after the two-seat critic (one re-dispatch), the 39-call unrhymed control arm, the cap probe and 420 judging calls. Delta 1.532300725 against a per-request sum of 1.532300725 — exact to 1e-9, note (brw)'s accumulate-across-attempts rule holding for the third session running. |
| 2026-08-29 (S231 start) | 155.667729713 | S231 session start, before any project spend. $0.353300441 above S230's close — non-project spend on the key between the two sessions, so the zero recorded at S230 was one session's luck and not a change of regime. |
| 2026-08-29 (S231 end) | 155.865436013 | S231 after two critic seats and 27 blind classification calls. Delta 0.197706300 against a per-request sum of 0.197706300 — exact. |
| 2026-08-29 (S232 start) | 156.738407873 | S232 session start, before any project spend. $0.872971860 above S231's close — the second consecutive inter-session gap of non-project spend on the key, and larger than S231's $0.353300441. |
| 2026-08-29 (S232 end) | 157.947158523 | S232 after two critic seats, 276 rating calls and 108 forced-choice calls. Delta 1.208750650 against a per-request sum of 1.208750650 — exact. |
| 2026-08-30 (S233 start) | 158.280830373 | S233 session start, before any project spend. $0.333671850 above S232's close — the third consecutive inter-session gap of non-project spend on the key. |
| 2026-08-30 (S233 end) | 159.782701273 | S233 after two critic seats, 345 adjudication calls and 10 re-buys. Delta 1.501870900 against a per-request sum of 1.398166500 — NOT exact. $0.022294000 is two truncated bodies deleted before the run restarted (billed, and ledgered); $0.081410400 is unexplained. |
| 2026-08-30 (S234 start) | 159.995500773 | S234 session start, before any project spend. $0.212799500 above S233's close — the fourth consecutive inter-session gap of non-project spend on the key. |
| 2026-08-30 (S234 end) | 162.059280998 | S234 after 151 bodies. Delta 2.063780225 against a per-request sum of 2.190483225 — NOT exact, and short by $0.126703000, the opposite sign to S233's. The larger figure, the per-request sum, is what is ledgered. |
| 2026-09-01 (S237 start) | 167.955767498 | S237 session start, before any project spend. +0.013577500 above S236's closing snapshot — ordinary between-session non-project drift, note (abf), and the first small one since S230. New UTC day, $0.00 of $5.00 spent. |
| 2026-09-01 (S237 end) | 171.717925748 | S237 after two critic seats, seven cap probes, a three-seat hand-identification probe and 88 blind coding calls. Delta 3.762158250 against a per-request sum of 3.465563850. Reported as an observation and not as a check, per note (bso): the endpoint lags, so neither agreement nor disagreement is claimed. The per-request sum is ledgered. |
| 2026-09-01 (S238 start) | 172.061750848 | S238 session start, before any project spend. +0.343825100 above S237's closing snapshot — the largest between-session drift since S232, and it is the direction S237's own lag hypothesis predicts (its delta ran $0.296594400 above its per-request sum). Recorded as an observation, not as a reconciliation of S237, per note (bso); nothing is claimed about which part is lag and which is non-project spend. |
| 2026-09-01 (S238 end) | 172.179654248 | S238 after two pre-run critic seats and nothing else — every figure of the run was computed from free, already-stored material. Delta 0.117903400 against a per-request sum of 0.117903400 — exact to 1e-9. |
| 2026-09-02 (S239 start) | 172.327928648 | S239 session start, before any project spend. +0.148274400 above S238's closing snapshot — between-session drift of the ordinary size, note (abf); recorded as an observation and not as a reconciliation of S238, per note (bso). New UTC day, $0.00 of $5.00 spent. |
| 2026-09-02 (S239 end) | 174.385329998 | S239 after two pre-run critic seats, a five-call seat-and-cap probe, 111 gloss and locate calls and 28 control calls. Delta 2.057401350 against a per-request sum of 2.057401350 — exact to 1e-9, the second exact reconciliation in a row. |
| 2026-09-03 (S242 mid) | 247.816893901 | S242 opening snapshot, taken after the two-seat critic pass (the critics ran before the snapshot, so it already carries them). +73.431563903 above S239's close — three sessions of gap (S240 and S241 each spent $0.00 and took no snapshot, so this covers all between-session non-project drift since 2026-09-02) plus whatever non-project key spend occurred; no project figure depends on it, and the within-session delta below is what bounds this run. |
| 2026-09-03 (S242 end) | 251.219193176 | S242 after the critic pass, three cap/price probes and 76 blind coding calls (5 truncated on Q2, not re-bought). Delta from the mid snapshot 3.402299275 against a per-request sum of 3.402299275 — exact to 1e-9, and exact because the mid snapshot already carried the critics. Per note (bso) the snapshot is an observation; the per-request sum is ledgered. |
| 2026-09-04 (S243 start) | 251.723042776 | S243 session start, before any project spend. +0.503849600 above S242's closing snapshot — between-session drift, larger than the ordinary size (note (abf)) and in the direction S242's own lag would predict. Recorded as an observation, not a reconciliation of S242, per note (bso). New UTC day, $0.00 of $5.00 spent. |
| 2026-09-04 (S243 end) | 251.931674426 | S243 after two pre-run critic seats, twelve blind second-coder batches and four refill batches. Delta 0.208631650 against a per-request sum of 0.208631650 — exact to 1e-9. |
| 2026-08-11 (S162 start) | 94.566276446 | S162 session start, before any project spend. +0.225104930 above S161's closing snapshot — twenty-five times the +0.009 drift of recent sessions, and the seventh consecutive session to show non-project drift. |
| 2026-08-11 (S162 end) | 95.554610396 | S162 after the critic and the 210-body run. Delta 0.988333950 against a per-response sum of 0.975093250; the delta is larger by $0.013240700, the non-project direction, breaking a run of three exact reconciliations. Per-request costs are ledgered. |
| 2026-08-11 (S163 start) | 96.627118158 | S163 session start, before any project spend. +1.072507762 above S162's closing snapshot — the eighth consecutive session to show non-project drift and by far the largest yet, roughly five times S162's +0.225104930 and nearly four times S159's +0.293605. Recorded, not explained; no project row accounts for it. At this magnitude the key delta is not usable as an independent cross-check of a session's total. |
| 2026-08-11 (S163 end) | 96.650683658 | S163 after its single pre-run critic call. Delta 0.023565500 against a per-response cost of 0.0235655 — exact, the first since S161, and it is exact because the snapshots bracket one call with nothing else in between. Per-request costs are ledgered. |
| 2026-08-11 (S157 start) | 90.464883994 | S157 session start, before any project spend. +0.008239800 above S156's closing snapshot — between-session non-project drift, the same shape as S044/S045. |
| 2026-08-11 (S157 end) | 91.482416375 | S157 after the critic, judge, re-dispatch and annotation stages. Delta 1.017532381 against a per-response sum of 0.956578567; the delta is larger by $0.060953814, i.e. the opposite direction from note (abf)'s lag. Re-read once per note (bil): identical figure. Per-request costs are ledgered. |
| 2026-08-11 (S158 start) | 91.487141875 | S158 session start, before any project spend. +0.004725500 above S157's closing snapshot — between-session non-project drift, the third consecutive session to show it. |
| 2026-08-11 (S158 end) | 93.371317150 | S158 after the critic, recognition, parity, audit, judge and three re-dispatch rounds. Delta 1.884175275 against a per-response sum of 1.846625863; the delta is larger by $0.037549412 — the same direction as S157 and the opposite of note (abf)'s lag. Per-request costs are ledgered. |
| 2026-08-11 (S159 start) | 93.664922150 | S159 session start, before any project spend. +0.293605000 above S158's closing snapshot — between-session non-project drift, the fourth consecutive session to show it and forty times the largest previous instance (S158's +0.004725500). Recorded, not explained; it is not project spend and no project row accounts for it. |
| 2026-08-11 (S159 end) | 93.822004040 | S159 after the critic and the 60-call census. Delta 0.157081890 against a per-response sum of 0.157081896 — agreement to 6e-9, the exact reconciliation S148 and S149 had and S157/S158 did not. Per-request costs are ledgered. |
| 2026-08-11 (S160 start) | 93.838473200 | S160 session start, before any project spend. +0.016469160 above S159's closing snapshot — between-session non-project drift, the fifth consecutive session to show it: four times the S157/S158 figures and a twentieth of S159's unexplained +0.293605. Recorded, not explained. |
| 2026-08-11 (S160 end) | 94.041883500 | S160 after the critic, 54 hand calls and 135 arbiter calls. Delta 0.203410300 against a per-response sum of 0.203410309 — agreement to 9e-9, the second exact reconciliation in a row after S159's. Per-request costs are ledgered. |
| 2026-08-11 (S161 start) | 94.050841400 | S161 session start, before any project spend. +0.008957900 above S160's closing snapshot — between-session non-project drift, the sixth consecutive session to show it. Recorded, not explained. |
| 2026-08-11 (S161 end) | 94.341171516 | S161 after the critic, 50 hand calls and 150 arbiter calls. Delta 0.290330116 against a per-response sum of 0.290330128 — agreement to 1.2e-8, the third exact reconciliation in a row. Per-request costs are ledgered. |
| 2026-07-23 14:47 | 8.2901842 | S001 session start, before any project spend |
| 2026-07-23 15:10 | 8.383558578 | S001 after probe + pilot; delta 0.0934 covers the probe only — key accounting lags recent calls, so per-request sums are primary (S001 total: 0.10382) |
| 2026-07-23 15:44 | 8.397545986 | S002 before ratification votes |
| 2026-07-23 15:49 | 8.460975277 | S002 after 3 ratification votes; delta 0.063429 matches the per-request sum exactly |
| 2026-07-24 13:37 | 8.819933777 | S010 before the first R01-vs-R02 comparison run (E-20260724-r01r02-selfrevise). Note: all-time figure jumped ~0.36 since S002 from non-project key use; per-request sums remain primary. |
| 2026-07-24 (post-run) | 12.946679798 | S010 after the run. Key delta from pre-run = 4.126746 vs per-request sum 4.488172; delta is smaller than the per-request sum (accounting lag — recent calls not yet reflected), so per-request sums are primary and this delta is a loose lower bound, not a contradiction. |
| 2026-07-25 00:48 | 14.532282548 | S014 session start, before any project spend (new UTC day; budget resets to $5.00). |
| 2026-07-25 (S015 start) | 16.780929839 | S015 session start. Cross-check: S014's start snapshot 14.532282548 + S014's ledgered $2.248649 = 16.780931 — matches to 2e-6, so the S014 ledger is confirmed exact. |
| 2026-07-25 (S015 end) | 17.931507092 | S015 after E-20260725-anchor-verification. Delta from S015 start 1.150577 vs per-request sum of kept responses 1.074446; the $0.076131 gap is two discarded first attempts (elicit__chatnoir__P3, elicit__vanka__P5) whose raw files the runner overwrote on retry, losing their usage.cost. The larger key-delta figure is what is ledgered below. Tool fixed (discarded attempts are now written to <path>.attemptN.json). |
| 2026-07-25 (S020 start) | 18.290925792 | S020 session start, before any project spend. Note the +0.359 drift from the S015-end snapshot is non-project key use; per-request sums remain primary. |
| 2026-07-25 (S020 end) | 19.046184168 | S020 after E-20260725-tierD-ladder (both stages). Delta from S020 start = 0.755258 against a per-request sum of 0.755259 — agreement to 1e-6, the second exact cross-check in the ledger and the first on a 60-call run. |
| 2026-07-25 (S021 start) | 19.046184168 | S021 session start — identical to the S020-end snapshot, i.e. zero non-project key drift between the two sessions; the first time that has been observed. S021's single call is cross-checked per-request only (one call, one cost). || 2026-07-28 (S045 start) | 23.482880228 | S045 session start, before any project spend. +0.005169 above S044's closing snapshot — the same between-session drift S044 attributed to concurrent non-project key use, at the same order of rate and two orders smaller. |
| 2026-07-28 (S045 end) | 23.689321227 | S045 after E-20260728b-forced-run-ru (15 probe calls) and two critic passes. Delta 0.206440999 against a per-request sum of 0.179372599, gap $0.027068. Third inexact cross-check in a row and the third with the same attribution; the per-request sum is what is ledgered, per this page's stated method. |
| 2026-07-28 (S046 start) | 23.689321227 | S046 session start — identical to S045's closing snapshot, so the between-session drift S044 and S045 both recorded did not recur. Note (abf) applied for the fourth time. |
| 2026-07-28 (S046 end) | 23.718394977 | S046 after one critic call. Delta 0.02907375 against a per-request sum of 0.02907375: exact to 1e-9, and the first exact cross-check in four sessions after three consecutive gaps attributed to concurrent non-project key use. One call, one provider, both snapshots written to disk before being read (note (bco)). |
| 2026-07-28 (S048, before its only call) | 24.382920625 | S048's opening snapshot. Identical to S047's settled closing figure, which is the whole story: S047 and S048 ran concurrently on the same key, and S048 opened after S047's spend had settled. S048 first recorded this as $0.664526 of unexplained non-project drift above S046's close and that attribution was wrong — it was S047's own $0.401902 plus the lag S047 itself documented. Corrected on the merge. |
| 2026-07-28 (S048 end) | 24.382920625 | S048 after one call billed at $0.051713 per-request. Delta exactly 0.000000, on two snapshots plus a third taken later in the session. The endpoint had not moved at all — the same settling lag S047 measured directly (its own figure took three snapshots and $0.347 to catch up). The per-request figure is what is ledgered, per this page's stated method. Note (bcx). |
| 2026-07-28 (S049 start) | 24.443194462 | S049 session start. $0.060274 above S048's close, of which S048's own $0.051713 accounts for 86% — S048 closed on a delta of exactly 0.000000 and said so, and its call had settled by the time this snapshot was taken. Residual $0.008561 unattributed, the same order as the small inter-session drifts at S042 and S045. Not the S033 shape (a delta in a session that made no call). Note (abf), sixth application. |
| 2026-07-28 (S052) | 25.048308603 | S052 opening and closing snapshot, taken either side of the session's only call. Identical, on a call billed at $0.024405625 per-request — the settling lag note (bcx) names, and the third session to see a delta of exactly zero across a real spend. The per-request figure is what is ledgered, per this page's stated method. Both snapshots written to disk before being read (note (bco)). |
| 2026-07-28 (S049 end, settled on re-read) | 24.909870691 | S049 after 14 calls. Delta from the opening snapshot $0.466676 against a per-request sum of $0.405701, gap $0.060975. The larger figure is ledgered, on the S030 precedent for a call the endpoint under-reports: the most economical explanation is the z-ai/glm-5.2 extraction attempt, which consumed 4,248 in / 5,972 out and returned usage.cost of exactly 0. Concurrent non-project key use is the alternative and this ledger has documented that at larger magnitudes; both readings are recorded rather than one being asserted. |
| 2026-07-29 (S054 start) | 25.843601589 | S054 session start. $0.008268 above S053's closing 25.835333489. S053 flagged an abandoned qwen/qwen3.7-max call that had billed nothing and warned that a late settlement would appear here; this drift is not evidence that it did — $0.008268 is the same order as the unattributed residuals at S042 ($0.005169) and S049 ($0.008561), neither of which had an abandoned call to blame, and a settled 8,000-token qwen call would be an order larger. The most economical reading is ordinary between-session drift, and the S053 warning is left standing rather than discharged. Note (abf), seventh application. |
| 2026-07-29 (S054 end) | 26.163373455 | S054 after nine billed dispatches. Delta 0.319771866 against a per-request sum of 0.319771867 — exact to 1e-9, the second consecutive session with that agreement, and this one across four providers and three labs including two calls that returned nothing. Both snapshots written to disk before being read (note (bco)). A tenth call — the four-class re-run — was dispatched and never returned; it had billed nothing at this snapshot, the same shape S053 recorded and flagged. |
| 2026-07-30 (S064 critic, either side) | 30.074394956 | S064 opening snapshot and the snapshot after its critic call, identical, on a call billed at $0.05598 per-request. The fourth session to see a delta of exactly zero across a real spend — the settling lag note (bcx) names. Both written to disk before being read (note (bco)). |
| 2026-07-30 (S064 end) | 30.177689543 | S064 after the 7-dispatch rater block. Delta from the rater-block opening 0.103294587 against a per-request sum of 0.146431981, a $0.0431 shortfall. The session-level bound holds (note (bet)): opening 30.074394956 + $0.202411981 = 30.276806937 against a closing 30.177689543, i.e. $0.0991 still unsettled at close, which is the ordinary lag direction and not the (bet) excess direction. Per-request sums are ledgered, per this page's stated method. |
| 2026-08-01 (S085 start) | 42.209298353 | S085 session start, before any project spend. |
| 2026-08-01 (S085 end) | 42.262070476 | S085 after four ratification calls. Delta 0.052772123 against a per-request sum of 0.0527721234 — exact to 1e-9, and both snapshots were written to disk before being read (note (bco)). Routing read off every response (note (x)): P1 OpenAI, P2 Google, P3 xAI, P5 Baidu — no price excursion on the seat that has produced them before. |
| 2026-08-02 (S089 open) | 45.164290590 | S089 session start, before any project spend. |
| 2026-08-02 (S089 close) | 45.896845469 | S089 after 61 calls (1 critic + 60 scoring). Delta 0.732554879 against a per-request sum of 0.578047435, gap $0.154507444. The per-request sum is what is ledgered, per this page's stated method. The gap is the same order and shape as the non-project key drift recorded on this very UTC day at S087 ($0.543269701); both readings are on the record rather than one asserted. Routing read off every response (note (x)): P1 OpenAI; P2 Google / Google AI Studio; P5 across six providers in 20 calls — Alibaba, DigitalOcean, GMICloud, Novita, StreamLake, Together — with no call billing above its declared per-call worst case. |
| 2026-08-02 (S090 open) | 46.007974247 | S090 session start, before any project spend. $0.111128778 above S089's close (45.896845469), consistent with S089's own $0.154507444 residual settling afterwards. | | 2026-08-02 (S090, after the critic) | 46.022567975 | Delta from the opening snapshot 0.014593728 — the accepted critic body's per-request cost to 1e-9. This is what establishes that the dispatch killed in flight by a foreground timeout billed exactly zero. | | 2026-08-02 (S090 close) | 46.041783975 | After the 9-call subject stage. Stage delta 0.019216000 against a per-request sum of 0.094170700 — $0.074954700 unsettled at close, the ordinary settling lag, note (bcx). Per-request sums ledgered. All three snapshots written to disk before being read (note (bco)). |
| 2026-08-07 (S126 open) | 71.490506475 | S126 session start, before any project spend. |
| 2026-08-07 (S126 close) | 71.706679770 | S126 after 14 billed bodies (2 critic passes + 12 rating calls). Delta 0.216173295 against a per-request sum of 0.216173290 — agreement to 5e-9, and the second exact cross-check on this UTC day. Both snapshots written before being read (note (bco)). Routing read off every response (note (x)): P1 OpenAI, P2 Google / Google AI Studio, P3 xAI, P5 DigitalOcean (pinned), critic 1 Together, critic 2 Alibaba — no price excursion. |
| 2026-08-07 (S127 open) | 72.239736726 | S127 session start, before any project spend. $0.533056956 above S126's close (71.706679770), and no ledger row and no landed session explains it — origin/main's last commit at S127's start was S126's. Recorded as unexplained rather than absorbed; it is the same order and shape as the non-project key drift already on the record for 2026-08-02 ($0.543269701, S087). No project figure depends on it, because this project ledgers per-request costs and uses the delta only as a cross-check. |
| 2026-08-07 (S129 open) | 74.485572001 | S129 session start, before any project spend. $0.007481000 above S128's close (74.478091001) — an ordinary settling lag (note (bcx)), not the unexplained kind recorded at S087 and S127. |
| 2026-08-07 (S129 close) | 74.837369916 | S129 after 17 billed bodies. Delta 0.351797915 against a per-request sum of 0.351797915 — exact to 0.000000000, the second exact reconciliation on this UTC day and the third this week. Both snapshots written to disk before being read (note (bco)). Routing (note (x)): J1 OpenAI, J2 Google / Google AI Studio, J3 DigitalOcean (pinned), critic BaseTen, pilot Alibaba, independent hand Mistral — no price excursion. |
| 2026-08-07 (S127 close) | 72.975808246 | S127 after 17 billed bodies. Delta 0.736071520 against a per-request sum of 0.732740290 — residual $0.003331230, which is one deepseek translation call billed server-side after a two-minute foreground shell timeout killed the runner mid-call, so its body was never written. It is ledgered as a waste row below and the reconciliation is then exact. Both snapshots written before being read (note (bco)). Routing (note (x)): P1 OpenAI, P2 Google / Google AI Studio, P3 xAI, P4 Morph / Moonshot AI / Wafer, P5 DigitalOcean (pinned), critic Together, replacement seat Alibaba — no price excursion. |
Ledger
S242 — 2026-09-03 (UTC), E-20260903b-line-end-order, ARM-line-end-order step 1 (T3, arm closed resolved)
UTC day 2026-09-03 had one prior session, S241, which spent $0.00, so the whole $5.00 was
available. Declared ceiling $4.50 (design-v2 §11), built from a freshly-read P1 price of
$2.00 / $12.00 — double the stale table value, note (bsw). Spent $3.402299275, well under both
the declared ceiling and the ~$3.50 the run's shape implied; $1.597700725 unspent.
| stage | calls | actual |
|---|---|---|
C pre-run critics, C1 + C2, cap 12000 — both NEEDS-REDESIGN, 28 findings, 8 BLOCKING |
2 | $0.089021400 |
| cap/price probes, three seats at the dispatch batch size — note (bsf) | 3 | $0.130562125 |
B blind coding, Q1 openai/gpt-5.6-terra, 447 items, 19 batches of 24, 0 dead |
19 | $1.017083600 |
B blind coding, Q2 google/gemini-3.6-flash, 447 items, 38 batches of 12, 5 dead (truncation) |
38 | $1.425220500 |
B blind coding, Q3 qwen/qwen3.7-max, 447 items, 19 batches of 24, 0 dead |
19 | $0.740411650 |
| per-request sum, and the ledgered figure | 81 | $3.402299275 |
Note (bsf) fires for the sixth time and this run honoured it and was still bitten: the cap was
probed per seat at the dispatch batch size and passed, and Q2 (gemini-3.6-flash) then truncated
on five later batches of the same shape, losing 60 items — its hidden reasoning is variable, as
S231 and S237 both recorded. Q2 failed its gate regardless, so the loss changed no verdict; the
bodies were not re-bought (note (brx)).
Note (bsw) fired for the first time: openai/gpt-5.6-terra was read at $2.00 / $12.00 from the
API before dispatch, double the $1.00 / $6.00 the table had carried since S106 — the first upward
price move recorded here, and the ceiling was set from the read value, not the cache.
Key usage 247.816893901 → 251.219193176, delta 3.402299275, exact against the per-request
sum (the opening snapshot already carried the critics). The translation limb
(T-hafez-shahed-R59-v1, غزل ۱۵ whole), the selection, both gold coding passes and all arithmetic
are the lead's own and are not ledgered (charter §3, A4).
UTC day 2026-09-03 running total after S242: $3.402299275 of $5.00, one session, $1.597700725 unspent.
S240 — 2026-09-02 (UTC), ARM-gulistan step 3, span D (T1)
$0.00 spent. No API call was made, and none was planned. The session's principal unit was a
translation limb — حکایات ۳۲–۴۰ of «گلستان» باب دوم rendered whole by the lead, which is free and
never ledgered (charter §3, A4) — and a study limb, RS-20260902-arabic-function, whose materials
are four public-domain English translations already stored in the repository from S235. Reading them
costs nothing.
No key-usage snapshot was taken, because none is needed for a session that made no request: the snapshot exists to bound a spend, and there is no spend to bound. The next session that spends will snapshot against S239's close (174.385329998) as usual, and any gap it finds is between-session drift covering two sessions rather than one — which is worth saying here so the gap is not read as this session's.
UTC day 2026-09-02 running total after S240: $2.057401350 of $5.00, two sessions, $2.942598650
unspent. A $0 session is a normal outcome (continue-prompt.md §7), and this one is $0 because the
question the unit asked did not need a model to answer it.
S239 — 2026-09-02 (UTC), E-20260902-rhyme-slot, ARM-rhyme-family step 2 (T2, arm closed resolved)
UTC day 2026-09-02 had no prior row, so the whole $5.00 was available. Declared ceiling $2.20
at design freeze; revised to $2.80 before any run call, on the measured seat probe rather than an
assumption (E-20260902-rhyme-slot/amendment-v2-1-probe.md §3, note (abc) — build the worst case
from the cap the request actually permits). Spent $2.057401350, $0.742598650 unspent against
the revised ceiling and $2.942598650 unspent on the day.
The translation limb — Hafez غزل ۳۵ rendered whole twice, monorhymed and unrhymed, under the new
regime R59 — is the lead's own and is not ledgered (charter §3, A4), as are the Ganjoor fetches,
the Leaf/Payne extraction, the line-geometry computation, all four contamination measurements, the
free gate on step 1's two-word window, and the whole analysis and verification.
| stage | calls | actual |
|---|---|---|
C pre-run critics, C1 openai/gpt-5.6-terra + C2 x-ai/grok-4.5, cap 12000 — both NEEDS-REDESIGN, 19 findings, 5 BLOCKING; neither truncated |
2 | $0.104133000 |
| seat-and-cap probe, on the lead's own poem only — note (bst), and it changed the locating seat | 5 | $0.094716750 |
G gloss, P1, cap 2500 — kept bodies including three re-buys at 6000 |
53 | $0.580438000 |
G gloss — three abandoned cap-2500 bodies, billed in full for nothing |
3 | $0.092926000 |
L locate, P3 x-ai/grok-4.5 at low reasoning effort, cap 2500 |
58 | $1.078872400 |
POSBIAS (both executions) and X-POS |
28 | $0.106315200 |
| per-request sum, and the ledgered figure | 149 | $2.057401350 |
Note (bst) is new and it cost the run its first-choice seat. google/gemini-3.6-flash located
8 of 14 senses at cap 2500 and 4 of 14 at cap 6000 on the same item — more room, more tokens
spent, a worse answer. It was dropped from the shape. x-ai/grok-4.5 at
reasoning: {"effort": "low"} did the work at $0.011 a call, which is also note (bsq)'s
counter-case: low effort failed a keyed calibration at S237 and passed a keyed positional control
here, so it remains a per-seat, per-shape measurement and not a lever.
Note (bsf) fired for the seventh session running. The three dead stage-G bodies cost $0.092926 and their re-buys $0.062950: a body that spends its whole cap and returns nothing bills for the whole cap, so the abandoned calls were the more expensive half.
Key usage 172.327928648 → 174.385329998, delta 2.057401350 against a per-request sum of 2.057401350 — exact to 1e-9. The opening snapshot sat $0.148274400 above S238's close, recorded as an observation per note (bso).
UTC day 2026-09-02 running total after S239: $2.057401350 of $5.00, one session. $2.942598650 unspent.
S238 — 2026-09-01 (UTC), E-20260901b-leaf-pattern, ARM-leaf-contract step 2 (T4, arm closed resolved)
UTC day 2026-09-01 had one prior session, S237, which spent $3.465563850, leaving $1.534436150. Declared ceiling $0.30 — the run needs no bought judgment at all, only the charter's pre-run critic pass. Spent $0.117903400; $1.416532750 of the day unspent.
| stage | calls | actual |
|---|---|---|
C pre-run critics P1 openai/gpt-5.6-terra + P3 x-ai/grok-4.5, cap 8000, both NEEDS-REDESIGN, 29 findings, 9 BLOCKING |
2 | $0.117903400 |
| per-request sum, and the ledgered figure | 2 | $0.117903400 |
This is what the budget is for and it is the cheapest thing it buys. The whole primary — 301 lines, 18 measures, a 9999-resample cluster permutation test, three controls, two coding sensitivities and a 92-check verifier — ran on stored material at $0. The $0.118 bought the two seats that replaced the primary and withdrew a registered prediction before any figure existed; had it not been spent, the session would have published a correlation that could not fail (note (bsr)).
Both seats returned clean at cap 8000 — P1 finish_reason: stop, 96.1 s, provider OpenAI,
$0.066007; P3 stop, 155.5 s, provider xAI, $0.0518964. Note (bsf) does not fire: neither
truncated, and no probe was needed because the prompt shape (one long document, a numbered-list
answer) is the shape S233 and S237 both dispatched these two seats at.
The key-usage delta is exact to 1e-9 for the first time since S232 — 172.179654248 − 172.061750848 = 0.117903400 against a per-request sum of 0.117903400.
UTC day 2026-09-01 running total after S238: $3.583467250 of $5.00, two sessions. $1.416532750 unspent.
S237 — 2026-09-01 (UTC), E-20260901-inversion-habit, ARM-inversion-price step 2 (T3, arm closed)
UTC day 2026-09-01 had no prior row, so the whole $5.00 was available. Declared ceiling
$3.50 — raised from $3.00 before the run, on both pre-run critics' finding that the design's own
worst case exceeded its ceiling with three seats (C1-A13, C2-A15). Spent $3.465563850;
$0.034436150 unspent against the ceiling, $1.534436150 against the day.
| stage | calls | actual |
|---|---|---|
C pre-run critics P1 + P3, cap 12000, both NEEDS-REDESIGN, 28 findings, 9 BLOCKING |
2 | $0.078468000 |
| cap probes at the batch size actually dispatched, all three seats — note (bsf)/(bsm) honoured | 7 | $0.297209350 |
H hand-identification leak probe, 3 seats — P1's answer lost to truncation at cap 3000 |
3 | $0.132838400 |
B blind coding, P1, 523 items, 22 batches of 24, cap 12000, 0 dead |
22 | $1.196756100 |
B blind coding, P2, 511 items, 44 batches of 12, cap 12000, 1 batch truncated, 12 items lost |
44 | $1.482960000 |
B blind coding, P3, 523 items, 22 batches of 24, effort: low, 0 dead |
22 | $0.277332000 |
| per-request sum, and the ledgered figure | 100 | $3.465563850 |
What the money bought and what it did not. The coding stage completed — 88 calls, 1 dead batch
of 88 — and then the registered calibration gate admitted one seat of a required two, so
$2.957048100 of blind coding is reported as a withholding. That is the failure criterion working:
it was written before the run and it fired against the run's own interest. The re-run is priced from
these figures in wiki/backlog.md's one live row.
Two seat-level facts for future pre-flights. P2 needs batch 12 on this task shape — at 24
it truncates at cap 12000 having answered 18 — which trebles its call count and made it the dearest
seat of the three at $1.48. And P3 at reasoning: {"effort": "low"} costs a thirteenth of full
effort ($0.0073 against $0.094 a batch) and failed both gates, calling 7 of 19 inverted keyed lines
canonical: note (bsq). Its $0.277332 bought nothing that voted.
Note (bsf) fires for the sixth session running, and this is the first run to honour it in full and
be bitten anyway — every seat was probed at the batch size it would be dispatched at, and one P2
batch in 44 still truncated at the probed cap. The note is amended rather than discharged.
The one re-buy that was declined. The dead P2 batch would have cost about $0.035 to re-dispatch
and would have taken the run to the ceiling exactly. It was left dead, the 12 items carry two seats'
codes instead of three, and this is on the result page rather than netted out.
UTC day 2026-09-01 running total after S237: $3.465563850 of $5.00, one session. $1.534436150 unspent.
S236 — 2026-08-31 (UTC), E-20260831b-radif-hands, ARM-radif-hands step 2 (T5, arm closed)
UTC day 2026-08-31 had one prior session, S235, which spent $0.00, so the whole $5.00 was available. Declared ceiling $2.20; spent $1.947485275, $0.252514725 unspent against the ceiling and $3.052514725 of the day's cap unspent.
| stage | calls | actual |
|---|---|---|
| match probes — ten held-out anchors unmasked, then masked; both 10 of 10 | 2 | $0.062963 |
C pre-run critics P1 + P3, cap 12000 — both NEEDS-REDESIGN, 22 findings, 6 BLOCKING |
2 | $0.070775 |
cap probes on P2 for the match shape — 3000 truncated, 8000 clean |
2 | $0.050556 |
M match, P1, cap 3000 then 8000 — 5 of the first 10 calls truncated and lost, $0.301 |
15 | $0.754757 |
M match, P2, cap 8000 then 12000 — 1 call lost |
13 | $0.327461 |
G blind grammatical classification, 87 radifs × 3 seats, batched 15 |
19 | $0.519773 |
K blind correspondence, 110 items × 2 seats, batched 28 |
8 | $0.161114 |
reconciliation control call, P1, cap 50 |
1 | $0.000086 |
| per-request sum, and the ledgered figure | 56 | $1.947485275 |
Note (bsf) fired for the fifth session running and this time it cost money. The cap was probed on
this task shape for P2 and not re-probed for P1, because P1 had returned a ten-item probe
of the same template clean at 1500. At twenty items per call P1 truncated on five of ten calls at
$0.060 each — $0.301 that bought no data. New note (bsm): the batch size is part of the task
shape, so probe at the batch you will dispatch.
The key-usage delta does not reconcile, for the first time, and it is the endpoint that is wrong. Opening snapshot 162.267773148, closing 167.942189998, delta $5.674416850. A control call costing $0.000086 moved the endpoint by $0.000000000, so it lags rather than counts extra, and the $3.727 residual is within $0.12 of S233's and S234's ledgered totals combined ($3.610943725). The most economical reading is that the endpoint had not caught up with those two sessions when this one opened. That is an inference, not a measurement. New note (bso): the per-request sum is the ledgered figure and the delta is reported as an observation, never as agreement.
UTC day 2026-08-31 running total after S236: $1.947485275 of $5.00, two sessions (S235 $0.00, S236 $1.947485275). $3.052514725 unspent.
S163 — 2026-08-11 (UTC), E-20260812-unlicensed-typography (ARM-source-beliefs step 2, T2, arm closed)
Pre-flight, written into the frozen design before dispatch and built from max_tokens, not from
an expected answer length (note (abc)). One call and one call only: the pre-run critic, seat
openai/gpt-5.6-terra, cap 3,000. 3,500 prompt tokens at $1.00/M plus 3,000 completion at $6.00/M =
$0.0215. Declared ceiling $0.03. Nothing else in the run costs anything: the census is
arithmetic over stored public-domain texts, and the translation is the lead's.
The seat was chosen against a documented failure mode rather than for price. P2, P4 and P5 have all
returned finish_reason: "length" with null content in this project; note (bmb)'s fix in run.py
is unmade, and the UTC day had $0.213 of headroom left. P1 has no such history.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-11 | S163 | pre-run critic over the frozen design and the built-materials manifest (openai/gpt-5.6-terra, cap 3,000, temperature 0) |
0.0215 line | 0.0235655 | per-response | NEEDS-REDESIGN, 7 findings, 7 BLOCKING, all 7 accepted. Body cut at the cap inside finding 7; not re-dispatched |
| 2026-08-11 | S163 | All lead work — Garshin «Очень коротенький роман» translated whole under R06 with a frozen registered prediction; the Gogol and Garshin copy-texts fetched, cleaned and checksummed; build_materials.py, census.py, verify.py; the result page; wiki/goodness-senses.md's consequence statement; framework/v0.2 §7.7; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.00. No dead body, no re-dispatch, no discarded call. 78.6% of the declared ceiling.
Key reconciliation: EXACT, the first since S161. Snapshots 96.627118158 → 96.650683658, delta 0.023565500, against a per-response sum of 0.0235655.
Between-session drift, and it is the largest yet recorded. S162 closed at 95.554610396; this session opened at 96.627118158 — +1.072507762 with no project call in between, against S162's own +0.225104930 and the "twenty-five times recent sessions' figure" it reported. Eighth consecutive session showing drift. Per-request costs are primary and are what is ledgered; at this magnitude the key delta is no longer usable as an independent cross-check, and a session that needs one should snapshot immediately either side of the dispatch, as this one did.
Where the estimate was wrong. By $0.0020655, 9.6% over the line and well under the ceiling —
the prompt was longer than 3,500 tokens because it carried the whole design plus the manifest. The
cap itself was the defect that cost something other than money: 3,000 tokens did not hold the
critic's answer, and findings 8+ are lost. Note (abc) prices the worst case from max_tokens; it
still says nothing about whether max_tokens is large enough to hold an answer, which is now the
second consecutive session to record that gap.
S162 — 2026-08-11 (UTC), E-20260811h-domestication-channel (ARM-realia-channel step 2, T3, arm closed)
Pre-flight, written before dispatch and printed by run.py --dry-run, built from max_tokens
and the exact input length of every pair (note (abc)). Study 64 pairs × 3 seats at cap 450
$0.747534; six duplicate pairs $0.070081; a 20-body re-dispatch contingency $0.129000;
pre-run critic $0.041000; plus $0.153746 already burned on two dead critic bodies before the
ceiling was set. Declared ceiling $1.15, revised in session from $0.85 → $0.95 → $1.10 →
$1.15 as the critic's amendments added a contrast and the dead bodies had to be paid for; every
revision is written into design.md §7 with its reason. Opening key snapshot 94.566276446; the
UTC day already carried $3.787817294 from six sessions, headroom $1.212182706.
Spent $0.975093250, 84.8% of the declared ceiling and 80.4% of the day's headroom.
| stage | bodies | cost |
|---|---|---|
pre-run critic, moonshotai/kimi-k3, cap 2,500 — null content, 2,497 reasoning tokens |
1 | $0.053388000 |
pre-run critic, moonshotai/kimi-k3, cap 6,000 — null content, 5,997 reasoning tokens |
1 | $0.100358400 |
pre-run critic, openai/gpt-5.6-terra, cap 5,000 — stop, 10 findings |
1 | $0.031341250 |
| study + duplicates, three seats, cap 450 | 210 | $0.790005300 |
| total | 213 | $0.975093250 |
Waste $0.153746400, 15.8% of spend — all of it the two dead critic bodies, and the second of them was the lead breaking note (bhf) rule (iii) by raising the cap instead of changing the seat. A further $0.058131 was billed for 14 truncated study bodies that the strict parser voided (note (bmb), new mode); those are counted as spend, not as waste, because 196 of 210 bodies parsed and the primaries were unaffected.
Key reconciliation, NOT exact. 94.566276446 → 95.554610396, delta 0.988333950 against a per-response sum of 0.975093250 — larger by $0.013240700, the non-project direction, breaking a run of three exact reconciliations.
UTC day 2026-08-11 total: $4.762910544 of $5.00 across seven sessions (S156 $0.329161470, S157 $0.956578567, S158 $1.851254924, S159 $0.157081896, S160 $0.203410309, S161 $0.290330128, S162 $0.975093250). $0.237089456 headroom left, 4.7% of the cap.
S156 — 2026-08-11 (UTC), E-20260811-floor (ARM-idiom-reach step 2, T5, arm closed)
Pre-flight, written before dispatch, built from max_tokens (note (abc)). One critic pass
(13k prompt, cap 6,000), 9 nat calls (2.3k prompt, cap 1,400), 6 swap calls (1.4k prompt, cap
1,400), plus re-dispatch headroom. Declared ceiling $0.75, raised to $0.90 at the amendment
commit when the pre-run critic's amendments took items 36 → 46, swap items 16 → 20 and calls
15 → 18. Opening key snapshot 90.140136724; the UTC day opened empty, headroom $5.00.
Spent $0.329161470, 36.6% of the declared ceiling.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-11 | S157 | pre-run critic, one pass over design.md and the built items (openai/gpt-5.6-terra, cap 12,000) |
0.10 line | 0.051103 | per-response | stop, 25 findings, 19 BLOCKING, verdict NEEDS-REDESIGN. Cap set at 12,000 from the start because S156's 6,000-cap pass truncated |
| 2026-08-11 | S157 | judge stage, 360 cells — 36 items × 2 questions × 3 seats, plus the recognition, fluency and spelling-only stages the critic forced in | 0.35 line, $1.20 ceiling | 0.861120482 | per-response sum | 336 of 360 usable on the first pass |
| 2026-08-11 | S157 | licensed re-dispatch, 25 cells at a 2,400 cap (deepseek/deepseek-v4-pro only) |
— | 0.034556086 | per-response sum | recovered 24 of 25; one cell (D-MB QR L2) never returned and is reported, not imputed |
| 2026-08-11 | S157 | independent realia annotation, 7 calls (openai/gpt-5.6-terra, not a judge) — critic finding F24 |
— | 0.009799 | per-response sum | agreement with the lead's spans 0.517 |
| 2026-08-11 | S157 | WASTE — 25 dead bodies, all deepseek/deepseek-v4-pro |
— | 0.043458086 | per-response | finish_reason: "length", null content, 700 reasoning tokens. Note (bmb) on a second slug; 4.5% of spend, against S156's 45.7% |
| 2026-08-11 | S157 | All lead work — Lu Xun 〈風波〉 translated whole under R06 with its frozen log and its contamination measurement; select.py, code.py, run.py, analyse.py, verify.py; the result page; framework/v0.2 §7.5; cultural-mediation; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
S161 — 2026-08-11 (UTC), E-20260811g-carrier (ARM-carrier step 1, T5, arm constituted)
Pre-flight, written before dispatch. Built from max_tokens and not from an expected answer
length (note (abc)): $1.456, declared as a $1.46 ceiling. The pre-run critic's SERIOUS 3
added the LEXNULL gate G4 and the ceiling did not move: the four extra variants were paid
for by measuring FLOOR on four windows instead of six, so the run stayed at 50 items and 201
calls. Spent $0.290330128, 19.9% of the ceiling.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-11 | S161 | pre-run critic, one pass over the frozen design and all built windows (qwen/qwen3.7-max, reserve slug outside both stages, cap 12,000) |
0.062 line | 0.020465625 | per-response | stop, 8 findings — 1 BLOCKING, 4 SERIOUS, 3 ADVISORY, verdict NEEDS-AMENDMENT. 5 accepted, 3 in part with the overrules written. Its SERIOUS 3 bought G4, which returned 0.0000 and is the run's most useful number |
| 2026-08-11 | S161 | Stage T, 50 translation calls — 2 hands × 25 (6 windows × 3 variants + 2 LEXNULL + 1 C-NEG + 4 A-repeats) |
0.320 line | 0.113725100 | per-response sum over 50 bodies | 50 of 50 stop, 0 empty, 0 re-dispatch |
| 2026-08-11 | S161 | Stage J, 150 arbiter calls — 3 seats × 50 same/different items | 1.074 line | 0.156139403 | per-response sum over 150 bodies | 147 stop, 3 dead at the 900 cap (deepseek/deepseek-v4-pro), all three spending the whole cap on hidden reasoning with null content |
| 2026-08-11 | S161 | WASTE — 3 dead bodies, all deepseek/deepseek-v4-pro |
— | 0.004260050 | per-response | finish_reason: length, null content, 900 reasoning tokens each. Note (bmb), on the same slug S157 recorded it on. 1.47% of spend, against S160's 0.0% and S158's 28.4%. Not re-dispatched: the primary was already withheld by G1 and a re-dispatch would have put a second cap into one seat's cells for no gain |
| 2026-08-11 | S161 | All lead work — «Kaşağı» fetched from tr.wikisource and structured into a 74-unit copy-text; a Turkish Ministry of Education school PDF decoded from its glyph ids and collated against it; the story translated whole under R06 with its frozen log D1–D18; the contamination measurement; build_items.py, run.py, analyse.py, verify.py; the design, the critic adjudication, the result page; the arm page; framework/v0.2 §8; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.004260050, 1.47%. Key reconciliation: EXACT — 94.050841400 → 94.341171516, delta 0.290330116 against a per-response sum of 0.290330128, 1.2e-8 apart, the third exact reconciliation in a row.
Where the estimate was wrong. It over-priced by 5.0×, the conservative direction, for the same reason as S160: arbiter answers are two-line JSON objects against a 900–1,200 cap. And the 900 cap killed three bodies outright — note (abc)'s complement for the fifth session running, and this time the verdict token was not recoverable, so the cells are void rather than rescued.
S161 total: $0.290330128 against a declared ceiling of $1.46. UTC day 2026-08-11: $3.787817294 of $5.00 across six sessions (S156 $0.329161470, S157 $0.956578567, S158 $1.851254924, S159 $0.157081896, S160 $0.203410309, S161 $0.290330128); $1.212182706 headroom, 24.2% of the cap unused.
S160 — 2026-08-11 (UTC), E-20260811f-dakghar-address (ARM-dakghar step 1, T1, arm constituted)
Pre-flight, written before dispatch. Built from max_tokens and not from an expected answer
length (note (abc)): $1.12, declared as a $1.25 ceiling and raised to $1.40 before
dispatch with the reason written, when the pre-run critic's SERIOUS 3 added a ninth pair (C-c),
worth +6 hand calls and +15 arbiter calls. Spent $0.203410309, 14.5% of the ceiling.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-11 | S160 | pre-run critic, one pass over the frozen design and all built pairs (qwen/qwen3.7-max, reserve slug outside both stages, cap 12,000) |
0.08 line | 0.015822325 | per-response | stop, 6 findings — 2 BLOCKING, 2 SERIOUS, 2 ADVISORY, verdict NEEDS-AMENDMENT. 4 accepted, 2 accepted in part with the overrules written. Its SERIOUS 3 added the only control that passed |
| 2026-08-11 | S160 | Stage T, 54 translation calls — 2 hands × (18 variants + 9 A-repeats) | 0.33 line | 0.094097000 | per-response sum over 54 bodies | 54 of 54 stop, 0 empty, 0 re-dispatch |
| 2026-08-11 | S160 | Stage J, 135 arbiter calls — 3 seats × 45 same/different items | 0.76 line | 0.093490984 | per-response sum over 135 bodies | 134 stop, 1 truncated at the 900 cap (J3 on FLOOR_T-a_C-c) with its verdict token intact and recovered; the cell is not in any primary |
| 2026-08-11 | S160 | WASTE | — | $0.00 | — | No dead body, no discarded call, no re-dispatch. The one truncated body was usable |
| 2026-08-11 | S160 | All lead work — «ডাকঘর» fetched, structured into a 430-unit copy-text and collated against the page images for all eleven pages of span A; span A translated whole under R05 with its frozen log D1–D23; register.md; build_items.py, run.py, analyse.py, verify.py; the design, the critic adjudication, the result page; the arm page; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.00. Key reconciliation: EXACT — 93.838473200 → 94.041883500, delta 0.203410300 against a per-response sum of 0.203410309, 9e-9 apart.
Where the estimate was wrong. Nowhere that cost money — it over-priced by 5.5×, which is the conservative direction, mostly because the arbiter answers were two-line JSON objects against a 900–1,200 cap. But the 900 cap truncated one arbiter body, which is note (abc)'s exact complement for the fourth session running: the cap prices the estimate and must also be large enough to hold the answer, and those are two different sizes.
S160 total: $0.203410309 against a declared ceiling of $1.40. UTC day 2026-08-11: $3.497487166 of $5.00 across five sessions (S156 $0.329161470, S157 $0.956578567, S158 $1.851254924, S159 $0.157081896, S160 $0.203410309); $1.502512834 headroom, 30.1% of the cap unused.
S159 — 2026-08-11 (UTC), E-20260811d-title-network (ARM-recurrence step 2, T4, arm closed)
Pre-flight: declared ceiling $1.40, built from max_tokens and the worst plausible provider
(note (abc)) — 60 calls at a 4,000 cap plus one critic pass at 12,000. Expected ≈$0.25, from the
identical task at E-20260810w ($0.182 for 48 calls). ACTUAL $0.157081896, 11.2% of the ceiling and
below the expectation.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-11 | S159 | pre-run critic, one pass over design.md and a complete built seat prompt (qwen/qwen3.7-max, reserve slug outside the jury, cap 12,000) |
0.08 line | 0.014851775 | per-response | stop, 5,887 chars, 4 findings — 1 BLOCKING, 2 SERIOUS, 1 ADVISORY, verdict NEEDS-AMENDMENT. 3 accepted, 1 accepted in part with the overrule written. Its BLOCKING 1 rewrote Q2 into a prediction that could fail, and Q2 then failed |
| 2026-08-11 | S159 | census, 60 calls — 3 seats × 20 items (4 hands × 3 regions, control on D and E) | 1.24 line | 0.142230121 | per-response sum over 60 bodies | 0 dead bodies, 60 of 60 finish_reason: stop, no re-dispatch. Call parameters byte-identical to E-20260810w's, which is why note (bmb)'s remedy needed no second application |
| 2026-08-11 | S159 | WASTE | — | $0.00 | — | No dead body, no truncation, no re-dispatch, against S158's 28.4% |
| 2026-08-11 | S159 | All lead work — Gogol «Шинель» regions D and E translated under R06 with its frozen log; the three-region cut; build_materials.py, build_items.py, run.py, analyse.py, verify.py; the design, the critic adjudication, the result page; consistency, S-berman-tendances §9, the anchor's second network; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
S159 total: $0.157081896 against a declared ceiling of $1.40. UTC day 2026-08-11: $0.329161470 (S156) + $0.956578567 (S157) + $1.851254924 (S158) + $0.157081896 (S159) = $3.294076857 of $5.00, 65.9% of the cap, $1.705923143 headroom. Waste $0.00.
Key reconciliation: EXACT. Snapshots 93.664922150 → 93.822004040, delta 0.157081890 against the per-response sum of 0.157081896 — 6e-9 apart, and the opposite of the $0.03–$0.06 excess deltas S157 and S158 both recorded.
Where the estimate was wrong. Nowhere that cost anything: the ceiling was 8.9× the actual, which
is the max_tokens arithmetic note (abc) requires and not a forecast. The one figure worth carrying
forward is that this task's real appetite is ~$0.0026 per call and has now been measured twice on
the same procedure ($0.182/48 and $0.142/60), so a third run of it can declare a tighter ceiling
without violating (abc) by citing the measurement rather than the hope.
S158 — 2026-08-11 (UTC), E-20260811c-source-beliefs (ARM-source-beliefs step 1, T2)
Pre-flight: declared ceiling $1.60, raised to $1.80 before dispatch when the critic's amendments
added 6 parity pairs and a 12-call explicitation audit; built from max_tokens and the worst
plausible provider (note (abc)). Expected ~$0.99. ACTUAL $1.846625863 — an OVERRUN of $0.046626,
2.6% above the declared ceiling, and it is the parity re-dispatch.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-11 | S158 | pre-run critic, one pass over design.md and the built items (openai/gpt-5.6-terra, cap 14,000) |
0.06 line | 0.094221 | per-response | stop, 25,077 chars, 20 findings, 19 of them severity-tagged, 17 BLOCKING, verdict NEEDS-REDESIGN. 14 accepted, 3 in part, 3 overruled |
| 2026-08-11 | S158 | G3 recognition, 3 calls on a non-judge seat (openai/gpt-5.6-terra) |
0.02 line | 0.029953 | per-response sum | 2 of 3 died at a 900 cap and were re-dispatched at 3,000 |
| 2026-08-11 | S158 | G2 parity, 23 pairs (moonshotai/kimi-k3) — 18 real + 5 planted |
0.18 line | 0.618838 | per-response sum over 33 bodies | 10 of 18 real bodies died at cap 2,000, $0.318536 of them; recovered at 4,000 with effort: low. This stage alone is a third of the run |
| 2026-08-11 | S158 | G2b explicitation audit, 12 pairs (openai/gpt-5.6-terra) |
0.14 line | 0.075273 | per-response sum | 12 of 12 stop first time |
| 2026-08-11 | S158 | judge stage, 108 cells — 3 arms × 6 segments × 2 blocks × 3 seats | 0.55 line | 1.028340 | per-response sum over 173 bodies | 39 dead, $0.192505. Three re-dispatch rounds; 2 cells never recovered |
| 2026-08-11 | S158 | WASTE — 51 dead bodies across deepseek/deepseek-v4-pro, moonshotai/kimi-k3 and google/gemini-3.6-flash |
— | 0.525140707 | per-response | all finish_reason: "length", null content, cap spent on hidden reasoning. 28.4% of spend, against S157's 4.5% and S156's 45.7%. Note (bmb), fired on two further slugs |
| 2026-08-11 | S158 | one live probe (deepseek/deepseek-v4-pro, effort: low, cap 4,000) |
— | 0.004629061 | per-response | the diagnostic that ended the truncation loop. Its body was printed and not persisted — a note (bco) violation by the lead, recorded rather than hidden; the cost is from the printed usage.cost |
| 2026-08-11 | S158 | All lead work — Ola Hansson «Sensitiva amorosa» IX translated under R06 with its frozen log; the three arms; census.py, build_items.py, run.py, analyse.py, verify.py; the design, the critic adjudication, the result page; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
S158 total: $1.851254924 — $1.846625863 in persisted bodies plus the $0.004629061 probe — against a declared ceiling of $1.80. An OVERRUN of $0.051254924, 2.8%. UTC day 2026-08-11: $0.329161470 (S156) + $0.956578567 (S157) + $1.851254924 (S158) = $3.136994961 of $5.00, 62.7% of the cap, $1.863005039 headroom. Waste $0.525140707, 28.4%, against S157's 4.5%.
Where the estimate was wrong, and it is one place. Every stage but one came in at or under its
line; the audit came in at half. Parity was estimated at $0.18 and cost $0.618838, 3.4× — not
because the per-call price was misjudged but because 11 of 18 bodies had to be bought twice.
Note (abc) prices the worst case from max_tokens; note (bfc) says declare a reserve for every
stage. Neither says what this run learned: a gate's reserve must be the full cost of dispatching
it twice, because a gate cannot be scaled down or dropped when its bodies die — a run with no
parity data is a run about damage rather than about form. Recorded as a firing of (bfc).
S157 total: $0.956578567 against a declared ceiling of $1.20. UTC day 2026-08-11: $0.329161470 (S156) + $0.956578567 (S157) = $1.285740037 of $5.00, 25.7% of the cap.
Where the estimate was wrong. The $0.35 pre-flight was built from 240 calls at an assumed ~120
output tokens. The critic's amendments added 151 calls — a recognition probe, a fluency probe, a
spelling-only manipulation check and an independent annotation, every one of them a repair the run
needed — and hidden reasoning on two seats put real output far above 120 tokens. The declared
ceiling held and the arithmetic under it did not; note (abc) prices the worst case from
max_tokens, and the lesson this row adds is that the call count is not fixed until the critic has
been adjudicated.
Key reconciliation, and it does not close in the usual direction. Snapshot 90.464883994 → 91.482416375, delta 1.017532381 against a per-response sum of 0.956578567 — the delta is larger by $0.060953814, where note (abf)'s lag makes it smaller. The key was re-read once per note (bil) and returned the identical figure, so this is not the lag. It is the non-project drift direction this ledger has recorded before (S020, S045). Per-request costs are primary and are what is ledgered.
| 2026-08-11 | S156 | pre-run critic, pass 1 (openai/gpt-5.6-terra, cap 6,000) | $0.049 | 0.05100675 | per-response | TRUNCATED — finish_reason: length. Note (abc) applied to a critic prompt: the cap was built for a shorter answer than the prompt invited |
| 2026-08-11 | S156 | pre-run critic, pass 2 (same seat, cap 12,000) | $0.085 | 0.0464793 | per-response | stop, 34,148 chars, 22 findings, 19 BLOCKING. This is the pass adjudicated. Cheaper than pass 1 despite twice the cap |
| 2026-08-11 | S156 | nat, 12 live bodies (4 blocks × L1 L2 L3) | $0.25 | 0.04768938 | per-response | 46 passages, 222 of 222 cells |
| 2026-08-11 | S156 | swap, 6 live bodies (2 forms × 3 judges) | $0.15 | 0.02132454 | per-response | 20 items each incl. 4 unchanged duplicate controls |
| 2026-08-11 | S156 | one probe (google/gemini-3.6-flash, reasoning: {"effort": "low"}) | — | 0.012189 | per-response | the diagnostic that ended the truncation loop; kept live in runs/probe-L3-lowreasoning.json |
| 2026-08-11 | S156 | WASTE — nine dead bodies, all google/gemini-3.6-flash | — | 0.150472500 | per-response | see below |
| 2026-08-11 | S156 | All lead work — Akutagawa 「蜜柑」 translated whole under R06; the corpus build; code.py, run.py, analyse.py, verify.py; the result page; framework/v0.2 §7.4; every state page | — | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.150472500 — 45.7% of the spend, the worst share this project has recorded. All of it
on one slug and one defect, now note (bmb): google/gemini-3.6-flash rejects
reasoning: {"enabled": false} with HTTP 400 and, left alone, expands hidden reasoning to fill
whatever cap it is given — 1,345 reasoning tokens at a 1,400 cap, 2,494 at a 2,600 cap —
truncating the answer both times. Raising the cap, which is note (abc)'s own remedy, made it
1.7× more expensive per body and did not fix it. reasoning: {"effort": "low"} fixed it at the
third round, at 1,057 reasoning tokens.
Key reconciliation: LAGGING BY EXACTLY ONE BODY. Snapshots 90.140136724 → 90.456644194,
delta 0.316507470 against a per-response sum of 0.329161470. The gap is 0.012654000,
which is to the cent the billed cost of nat-a-L3's first dispatch — a single unsettled call, not
a discrepancy. Note (abf), and the per-response sum is what is ledgered, per this page's method.
Where the estimate was wrong. Twice, in opposite directions, and both are recorded because they
cancel and would otherwise look like an accurate estimate. The critic line was under-built
(pass 1 truncated at a cap the pre-flight called sufficient) and the jury lines were
over-built — nat came in at $0.048 against $0.25 and swap at $0.021 against $0.15, because
the worst case was priced at max_tokens on all three seats when two of the three are terse. The
$0.150 of waste is not an estimation error at all: it is a defect the pre-flight had no way to
price, because the slug's behaviour was unknown before the run.
S119 — 2026-08-06 (UTC), E-20260806c-source-grammar (ARM-translated-register step 2, arm closed)
Pre-flight, written before dispatch. One API call only: the independent pre-run critic, seat
x-ai/grok-4.5 (P3). Worst case $0.20 — max_tokens 24,000 at the P3 list output price of
$6.00/M is $0.144, plus ~14k prompt tokens at $2.00/M ($0.028), plus routing margin. Everything else
in this session is arithmetic over public-domain text and the lead's own translation, both $0.
Opening key-usage snapshot 67.593647444. Headroom at dispatch: $3.719623473 of the UTC day's cap.
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-06 | S119 | stage 0 — independent pre-run critic, one pass (x-ai/grok-4.5, P3) |
0.20 from max_tokens 24,000 with routing margin |
0.041870400 | per-response usage.cost |
NEEDS-AMENDMENT, 10 findings, 6 BLOCKING, all ten accepted, finish_reason: stop, 7,031 prompt / 4,671 completion tokens of which 2,898 reasoning. Nine amendments applied before measure.py was written. F1 stripped the arm-closing power from the two predictions the translation limb had generated; F3 added FC4b, the gate that catches an effect living inside quoted speech that the share decomposition would miss; F5 named the genre confound that forced the run's first post-hoc probe |
| 2026-08-06 | S119 | Storm «Immensee» ch. «Elisabeth» rendered twice (R06 draft, R04 close) with a frozen census in the log; the Spanish comparator census; the nine-cell corpus, the measurement and its verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S119 total, one billed body: $0.041870400, against a declared worst case of $0.20 — 21%.
One dispatch, one accepted body, zero retries, zero finish_reason: length. The key-usage
cross-check CLOSES to 0.000000000 — opening snapshot 67.593647444, closing 67.635517844, delta
0.041870400 against a per-request sum of 0.0418704.
The critic was again the whole ledger and again the best line on it. It cost 100% of the
session's spend and it is the reason P1 can be believed: without F3's FC4b the run would have
reported a replication without ever checking whether the effect survives with quoted speech removed,
and without F1 it would have closed the arm on two predictions derived from a translator-coded 2×2
over two unmatched loci.
2026-08-06 day total: $1.322246927 of $5.00 — three sessions (S117 $0.614681164, S118 $0.665695363, S119 $0.041870400). $3.677753073 headroom, 74% of the cap.
S118 — 2026-08-06 (UTC), E-20260806b-berman-occurrence (ARM-berman-occurrence step 1, arm constituted)
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-06 | S118 | stage 0 — independent pre-run critic, one pass (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.30 from max_tokens 16,000 with routing margin |
0.021447400 | per-response usage.cost |
NEEDS-REDESIGN, 15 findings, 6 BLOCKING. 257.8s, plus one non-JSON keep-alive body retried per note (bgc) at $0. Eleven amendments accepted, one rejected with evidence, three accepted as declared limits. It replaced an incoherent gate (F1, which would have withheld the primary exactly when the hypothesis was supported), and forced the anti-circularity control that is the run's most useful positive result |
| 2026-08-06 | S118 | stage 1 — arm P, the unbriefed 2026 rendering | 0.15 | 0.052536960 | per-response usage.cost |
The accepted body was deepseek/deepseek-v4-pro, $0.00270396, 32.9s via StreamLake. The other $0.049833 is a moonshotai/kimi-k3 attempt that returned finish_reason: length at max_tokens 3,000 — note (bhf) again, on a plain translation payload, and the seat was changed rather than the cap raised a second time |
| 2026-08-06 | S118 | stage 2 — the coding run, 12 sites × 3 seats (P1, P2, P3) | 1.42 from max_tokens 2,500/6,000 with routing margin |
0.591711003 | per-response usage.cost |
36 of 36 seats returned, zero retries, zero seat failures, 558 codes. Two P2 bodies were lost to finish_reason: length at 2,500 before amendment A12 raised that seat alone to 6,000. The same patch moved the stage reserve off deepseek-v4-pro — the model that had produced arm P — which a fall-through would have made code its own translation, against charter §5 |
| 2026-08-06 | S118 | Verga «Cavalleria rusticana» ¶50–81 rendered three ways (R06 draft, R04 close, R08 resistancy) with frozen logs; two published renderings recovered and two-scan verified; the design, the contamination gate, the machine limb, scoring and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Three of the run's five arms cost nothing |
S118 total, summed per-request over all 41 stored billed bodies: $0.665695363, against a declared worst case of $1.71 — 39%.
The key-usage cross-check does NOT close, and it is out in the unusual direction. Opening
snapshot 66.534097841 (identical to S117's close, i.e. no inter-session drift), closing
67.553677814, delta 1.019579973 against a per-request sum of 0.665695363 — a gap of
$0.353884610, with the delta larger than the sum rather than smaller. The per-request sum is
what is ledgered, per this page's stated method. Two readings are on the record rather than one
asserted: (i) a moonshotai/kimi-k3 dispatch at max_tokens 8,000 was killed in flight by a
foreground timeout, so its body was never written and its usage.cost never read — kimi bills
$3/$15 per M and a full 8,000-token completion would be of this order; and (ii) ordinary non-project
key use, which this ledger has documented at up to $0.543 on a single day (S087). The first is
named first because it is known to have happened. 42 raw bodies are on disk; 41 carry a cost, and
the 42nd is the critic's keep-alive whitespace body, which billed nothing.
The cheapest line was again the most valuable. The $0.021 critic pass is 3% of the spend and is the reason the run has a defensible positive result at all: its finding 4 (site selection is circular) produced the three systematically-sampled sites, and those sites scored as high as or higher than the nine hand-picked ones — which is the only thing standing between this run's surviving number and the objection that it was arranged.
2026-08-06 day total: $1.280376527 of $5.00 — two sessions (S117 $0.614681164, S118 $0.665695363). $3.719623473 headroom, 74% of the cap.
S117 — 2026-08-06 (UTC), E-20260806-carrier-or-knowledge (ARM-voice-persona step 2, arm closed)
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-06 | S117 | stage 0 — independent pre-run critic, one pass (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.30 from max_tokens 16,000 with routing margin |
0.047795400 | per-response usage.cost |
NEEDS-AMENDMENT, 6 findings, 4 BLOCKING, all six accepted. 186.5s. It rebuilt a screen that would have passed the critical arm without looking at it, doubled the manipulation check, raised the primary from 3 seats to 4, added F7, and forced a fourth rendering |
| 2026-08-06 | S117 | stage 1 — N, the unbriefed paraphrase (mistralai/mistral-medium-3-5) |
0.05 | 0.004833 | per-response usage.cost |
1 of 1. The body carried an assistant preamble and closing note; the strip removed only the terminator — note (bjk) |
| 2026-08-06 | S117 | stage 2 — manipulation check, 5 texts × 3 seats (glm-5.2, qwen3.7-max, mistral-medium-3-5) |
0.58 from max_tokens 6,000 with routing margin |
0.083650455 | per-response usage.cost |
14 of 14 on first dispatch, zero retries. mistral-medium-3-5 returned K6 = 6 on all four texts it rated — a degenerate seat on the decisive scale, recorded in the result's limits |
| 2026-08-06 | S117 | stage 3 — content-parity screen, 3 pairs × 2 seats | 0.30 from max_tokens 6,000 with routing margin |
0.262543750 | per-response usage.cost |
Four bodies wasted on the C–A cell — two qwen3.7-max and two reserve kimi-k3, all finish_reason: length or non-JSON. Note (bhf)'s ninth firing, on the 70%-rewrite pair where enumeration is longest. Declared amendment A7 raised that cell's cap to 16,000 and the fifth attempt returned 61 characters |
| 2026-08-06 | S117 | stage 4 — the primary, 4 pairs × 4 seats × 2 orders (P1, P2, P3, P5) | 1.76 from max_tokens 4,500 with routing margin |
0.180203469 | per-response usage.cost |
32 of 32 on first dispatch, zero retries, zero seat failures |
| 2026-08-06 | S117 | declared post-hoc probe — the repaired floor, 4 seats × 2 orders | inside stage 4's margin | 0.035655089 | per-response usage.cost |
8 of 8, all returning 0.000. Labelled a probe everywhere; the primary was NOT re-decided on it |
| 2026-08-06 | S117 | one Gogol paragraph rendered four ways with four frozen logs, R19 minted, the design, the contamination gate, scoring and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S117 total, summed per-request over all 65 stored billed bodies: $0.614681164, against a declared worst case of $2.90 — 21%. The stage subtotals above were hand-estimated on first writing and three were wrong; they are now the figures summed from the stored bodies, and the correction is recorded rather than made silently — note (bey)'s shape, on a ledger rather than a log.
The most expensive stage was not the primary. The 32-call primary cost $0.180203469; the 6-call parity screen cost $0.262543750, because four wasted bodies and one 16,000-token re-dispatch sit inside it. 43% of the run's spend bought six lines of screen output — note (bhf) again, priced. The worst case was itself revised upward mid-run, from $2.40, by the critic's amendments; the fourth rendering they forced cost $0.
The key-usage cross-check closes to 2e-9 — opening snapshot 65.919416679, closing 66.534097841, delta 0.614681162 against a per-request sum of 0.614681164. Note (bil) was respected: the closing snapshot was taken after a settle, not at the moment the last body returned.
2026-08-06 day total: $0.614681164 of $5.00 — one session (S117 $0.614681164). $4.385318836 headroom, 88% of the cap.
| date (UTC) | session | item | est. USD | actual USD | measured how | notes |
|---|---|---|---|---|---|---|
| 2026-07-23 | S001 | panel bootstrap probe (9 models × 1 call) | 0.30 | 0.093374 | per-response usage.cost, summed |
raw: config/probes/2026-07-23/; all 9 responded |
| 2026-07-23 | S001 | pilot: 蜘蛛の糸 §一, R01 vs R02, translator P5 | 0.05 | 0.010446 | per-response usage.cost, summed |
raw: workshop/experiments/E-20260723-pilot-kumonoito/runs/ |
| 2026-07-23 | S002 | ratification votes: 1 non-Anthropic vote each for D-01/D-02/D-03 | 0.10 | 0.063429 | per-response usage.cost, summed; key-delta cross-check exact |
raw: wiki/decisions/votes/2026-07-23/; gpt-5.6-terra 0.030396, grok-4.5 0.022928, deepseek-v4-pro 0.010105. Canon verification pass was network-only (no API spend). |
| 2026-07-24 | S010 | first R01-vs-R02 comparison E-20260724-r01r02-selfrevise: 36 translation arms + 96 jury calls (2 works × 2 translators × 3 reps × 3 arms; 2 jurors × 48 payloads) | 3.00 (fund) | 4.488172 | per-response usage.cost, summed over all raw incl. failed attempts | est. was $1.9–3.0; overran to $4.49. Cause: reasoning-model max_tokens starvation truncated 26 calls (wasted $0.577 on failed deepseek/gemini attempts) + grok-4.5's heavy jury reasoning tokens (~$0.04/call vs ~$0.015 assumed). Breakdown: translation $1.116137 (incl. $0.088852 failed), jury $3.372183 (incl. $0.488225 failed). Raw: workshop/experiments/E-20260724-r01r02-selfrevise/runs/. |
2026-07-23 day total: $0.167249 of $5.00 (S001 $0.103820 + S002 ratification votes $0.063429).
| 2026-07-25 | S014 | ratification vote D-20260724-04: 1 non-Anthropic vote (openai/gpt-5.6-terra) | 0.02 | 0.016985 | per-response usage.cost | raw: wiki/decisions/votes/2026-07-25/D-20260724-04__gpt-5.6-terra.json; in 2335 / out 646 tok |
| 2026-07-25 | S014 | jury calibration Case A judge pass: 60 scoring calls (5 jurors × 6 items × 2 orderings) | 0.90 | 1.832000 | per-response usage.cost, summed over 60 kept files ($1.816490) + 1 discarded parse-fail attempt ($0.015510) | raw quarantined private-texts/experiments/E-20260723-calibration-v1/runs/jury/. Overran est: moonshotai/kimi-k3 ran ~$0.06–0.11/call (heavy reasoning); deepseek/gpt cheap. |
| 2026-07-25 | S014 | jury calibration Case A probe pass: 45 register+identity calls (5 jurors × 9 candidate-spans) | 0.25 | 0.399664 | per-response usage.cost, summed | raw quarantined .../runs/probe/. Also over est (kimi reasoning tax). |
| 2026-07-25 | S015 | second-reader verification of the 3 precedent anchors E-20260725-anchor-verification: 12 competence-screen + 9 blind-elicitation + 12 adjudication calls | 1.00–1.60 (cap 2.00) | 1.150577 | key-usage delta (conservative); per-request sum of kept responses 1.074446 | screen 0.051657 · elicit 0.343673 · adjudicate 0.679117 = 1.074446 kept, + ~0.076 in two discarded retry attempts. Raw: workshop/experiments/E-20260725-anchor-verification/runs/ (public tree — all six texts are PD, no quarantine). First run in this project to land inside its estimate; the pre-dispatch worst-case guard never fired. |
| 2026-07-25 | S020 | Tier D pilot E-20260725-tierD-ladder stage 1 — the sham arm, which is also the 1–7 scale pilot: 2 items × 2 orderings × 3 jurors | 0.166 (worst case 0.259) | 0.152631 | per-response usage.cost, summed; recomputed independently by tools/verify_tierD.py | 12/12 parsed first attempt, 0 discards. Raw: workshop/experiments/E-20260725-tierD-ladder/runs/ (public tree — both passages PD or lead-authored, no quarantine). |
| 2026-07-25 | S020 | same experiment, stage 2 — four perturbation operators × 2 passages × 2 orderings × 3 jurors | 0.663 (worst case 1.035) | 0.602628 | same | 48/48 parsed first attempt, 0 discards. Gate between stages was pre-registered and cleared on both conditions (scale usage, headroom). |
| 2026-07-25 | S021 | independent pre-run critic pass on the held-out-pair qualification design (E-20260725-heldout-pair): 1 call, x-ai/grok-4.5 | 0.04 (worst case 0.12) | 0.1150644 | per-response usage.cost | in 12,527 tok · out 15,038 tok of which 9,742 reasoning. Over the central estimate by 2.9×, inside the worst case. Cause is exactly NEXT.md note (b): the estimate priced the central case at ~4k reasoning+output and the model spent 9.7k on reasoning alone. Raw: workshop/experiments/E-20260725-heldout-pair/runs/critic__grok-4.5.json. The run it gated cost $0.00 (no API calls); the pass returned NEEDS-REDESIGN and stopped a false Tier D unblocking. |
2026-07-24 day total: $4.488172 of $5.00 (S010 first R01-vs-R02 comparison). Headroom nearly exhausted — no further API today. Lesson recorded: reasoning models (deepseek-v4-pro, gemini-3.6-flash, grok-4.5) spend 1.7k–11k tokens reasoning before output; size max_tokens for reasoning + output, and budget jury calls on reasoning-heavy models at ~2–3× the naive output-token estimate.
| 2026-07-25 | S022 | D-20260725-06 ratification: independent adversarial review (openai/gpt-5.6-terra) + review vote (google/gemini-3.6-flash) | 0.09–0.21 | 0.117062 | per-response usage.cost | review 0.08549175 (in 12,180 / out 5,496) + vote 0.0315705. Raw: wiki/decisions/votes/2026-07-25/. Verdict Q-A; both rejected the lead's stated preference. |
| 2026-07-25 | S022 | D-20260725-05 ratification: independent adversarial review (deepseek/deepseek-v4-pro) + review vote (openai/gpt-5.6-terra) | 0.04–0.11 | 0.098220 | per-response usage.cost | review 0.044837307 + vote 0.0533825. P5 cost 3.8× its config/models.md list price — OpenRouter routed the call to provider Venice at ~$1.65/$3.30 per M against the listed $0.435/$0.87. See the note below. |
| 2026-07-25 | S022 | E-20260725-slatef-verification: blind second reading of three Slate F sources, mandated by the D-05 vote (google/gemini-3.6-flash) | 0.09–0.15 | 0.282159 | per-response usage.cost, summed over all seven calls incl. four wasted | Usable $0.115349 (3 calls). Wasted $0.190665 (4 calls) — see the waste row below. Raw: workshop/experiments/E-20260725-slatef-verification/runs/. |
| 2026-07-25 | S022 | (of the row above) waste, itemised | — | 0.190665 | same | (a) $0.087833 — a shell loop wrote all three outputs to one file ($n__gemini parsed as one variable name); two paid responses overwritten. (b) $0.102833 — max_tokens sized for output but not for hidden reasoning; P2 spent 5,758 of 5,996 completion tokens reasoning and truncated mid-answer. This is standing note (b) verbatim, not applied. |
| 2026-07-25 | S022 | E-20260725-published-audit (the paired unit's study limb) | 0.00 | 0.000000 | — | No API call. The audit instrument, the translation and the analysis are all lead work. |
| 2026-07-25 | S025 | D-20260725-07 ratification: independent adversarial review (openai/gpt-5.6-terra) + review vote (google/gemini-3.6-flash) | 0.107 (worst case 0.119) | 0.071062 | per-response usage.cost; key-usage delta exact | review 0.04170125 (in 8,785 / out 950, provider OpenAI) + vote 0.029361 (in 9,914 / out 1,932, provider Google). Raw: wiki/decisions/votes/2026-07-25/. Verdict: review A, vote C (governs). Third run to land inside its estimate, and 34% under it. What worked: both prompts opened with a hard brevity instruction and a numbered output order chosen so the last item is the one the project can most afford to lose — NEXT.md notes (b) and (v), applied on purpose rather than after the fact. Neither call truncated. Provider read off both responses per note (x): no surprise routing. |
| 2026-07-25 | S026 | E-20260725c-contamination-sweep (the paired unit, both limbs) | 0.00 | 0.000000 | — | No API call. The «Припадок» translation, the sweep tool, the self-test gate, the bracket/asymmetry/centrality diagnostics and the write-up are all lead work. Lead translation is free and is never ledgered (charter §3, A4); the measurement over it is arithmetic. Day headroom unchanged at ~$0.16. |
| 2026-07-26 | S027 | E-20260726-period-control (the paired unit, both limbs) | 0.00 | 0.000000 | — | No API call. The «Роза» translation, the extraction, the 42-passage Garnett~Hapgood reference distribution, the four cell analyses, the rank statistic and the 77-check independent verification are all lead work. Lead translation is free and is never ledgered (charter §3, A4); the measurement over it is arithmetic. First row of a fresh UTC day; full $5.00 headroom remains. |
| 2026-07-26 | S028 | E-20260726-genealogie-period pre-run critic (openai/gpt-5.6-terra, P1) on the frozen design | 0.04 (worst case 0.14) | 0.080616 | per-response usage.cost; key-usage delta exact (19.938089975 → 20.018705600 = 0.080616) | in 4,452 / out 4,447, provider OpenAI — list price honoured, no surprise routing (note (x)). Landed inside its worst case, above its central estimate: the worst case was built from per-call maxima per note (m), and the call used its full output allowance because the critique was long, not because it truncated. The $0.08 bought five defects, eleven confound mechanisms, eight under-specified phrases and two predictions ruled unfalsifiable as written — all applied as dated amendments before any measurement existed (critic.md). Everything else this session — the 1,135-word Nietzsche translation, the 78-section extraction, the four-variant analysis, the reference distribution and the 76-check independent verification — cost $0.00. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-26 | S029 | E-20260726-ovid-period-form pre-run critic (openai/gpt-5.6-terra, P1) on the frozen design | 0.090 (worst case 0.113) | 0.070634 | per-response usage.cost; key-usage delta exact (20.0263443 -> 20.09697805 = 0.07063375) | in 6,821 / out 3,288, finish_reason: stop, provider OpenAI — list price honoured, no surprise routing (note (x)). Landed 22% under the central estimate and inside the worst case, the fourth run in this ledger to land inside its estimate. The estimate was a range with a worst case built from per-call maxima (note (m)) and the prompt opened with an explicit brevity instruction plus a task order whose last item is the one the project could most afford to lose (notes (b), (v)). The $0.07 bought 26 findings: 9 falsifiability rulings against named predictions and failure criteria, 11 confound mechanisms, 10 extraction defects and 5 over-claims — including the ruling that the design's central claim about its own lead statistic ("translator-specific vocabulary, form, register and text length all cancel") was simply false. All applied as dated amendments before any measurement existed (critic.md, design.md §10). Everything else this session — the 120-line Ovid translation, the five-text 15-book extraction, 90 published-pair measurements plus 90 mismatched controls, the dependence diagnostics and the 173-check independent verification — cost $0.00. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-26 | S030 | E-20260726b-forced-or-borrowed pre-run critic (openai/gpt-5.6-terra, P1) on the frozen design | 0.077 (worst case 0.078) | 0.081006 | key-usage delta (larger than the per-request sum — see below) | The successful call: in 6,015 / out 4,274, finish_reason: stop, provider OpenAI, usage.cost $0.0656205 — inside the estimate, the fifth consecutive run to land inside one. It bought 46 findings and 22 dated amendments, including that this design's own §3 claim about what its control controlled for was false (the third session running that pointing the critic at the statistic definitions produced the best single finding — note (rr)), a scaling rule cited by four predictions that had never been written down, and a selection bug that would have built target strings occurring nowhere in More. |
| 2026-07-26 | S030 | (of the row above) waste, itemised | — | 0.015386 | key-usage delta | A first attempt at the same call returned 1,034 bytes of keep-alive whitespace and no JSON body at all and was billed anyway. Measured at $0.0096395 immediately after; the key-usage delta at session end is $0.0057464 higher than the per-request sum, which is most plausibly that same call's billing settling after the first snapshot. Ledgered at the larger, conservative figure. Mitigation applied in run_critic.py: retry up to three times, detect a body with no {, and write every attempt to its own file so a second response can never overwrite a first (S022 warning 2, which was about exactly this). |
| 2026-07-26 | S030 | E-20260726b-baseline-dependence (the retro-check) and everything else | 0.00 | 0.000000 | — | No API call. The 140-line Ovid translation, the eleven-pair dependence retro-check over 124 units, the locus selection, the ten-window analysis, the name_tokens diagnosis and the 218-check independent verification are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-26 | S031 | E-20260726c-forced-or-borrowed-ru pre-run critic (openai/gpt-5.6-terra, P1) on the frozen design | 0.088 (stated worst case 0.090) | 0.091925 | per-response usage.cost; key-usage delta exact (20.54597125 → 20.63789625 = 0.091925) | in 7,399 / out 4,587, finish_reason: stop, provider OpenAI, list price honoured. This is the first run in this ledger to land OUTSIDE its stated worst case, by $0.0019, and the reason is a defect in how the worst case was built, not in the routing: note (m) says build the worst case from per-call maxima, and the estimate used an assumed output of ≤4,500 tokens while max_tokens actually permitted 7,000. The output came in at 4,587. A worst case built from a guess about output length is not a worst case; it must be built from the cap the request actually allows — new note (abc). Against that correct bound the worst case was ~$0.124 and the call landed well inside it. What the $0.09 bought: 16 dated amendments before any locus existed, including — for the fourth session running, note (rr) — that the design's own central claim about its own control was false. §3 asserted that carrying S030's thresholds over verbatim made this a replication; it does not, because the control region had been changed from ~91 tokens to a whole poem of 146–623, which changes what r_i = 0.5 means in an indeterminate direction. A1 bracketed the geometry in response. Also caught: a selection confound on poem length that §7 had missed, that G1's monotonicity can be satisfied by a globally shifted mapping, and that a verifier "not importing select.py" is worthless while both call tools/ngram_overlap.py — which is exactly how S030's 218-check pass ran over a broken name_tokens. |
| 2026-07-26 | S031 | the name_tokens repair, its re-runs, and the whole translation limb | 0.00 | 0.000000 | — | No API call. The three-defect repair with ten fixtures, the four gated re-runs over S026/S029/S030, the 42-unit prose-poem probe, the eight-poem 1,871-word Turgenev translation, the two-witness provenance check that found eleven OCR errors, the locus selection, the four-geometry analysis and the 172-check independent verification are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-26 | S033 | Unattributed key-usage delta, ledgered conservatively | — | 0.013955 | key-usage delta (20.63789625 → 20.65185145) | S031's closing snapshot to S033's opening one, across S032 — a session that made no API call and recorded $0.00. S031's own cross-check was exact at the time it was taken, so this is most plausibly late settling, the same mechanism S030 recorded and in the same direction. Charged to the day rather than written off, and the practice that found it is new: snapshot at session start as well as end — note (abf). One occurrence is not a pattern; a third would be. |
| 2026-07-26 | S033 | D-20260726-08 ratification — independent adversarial review (openai/gpt-5.6-terra, P1) | 0.068 (worst case) | 0.037055 | per-response usage.cost | in 1,591 / out 2,139, finish_reason: stop, provider OpenAI, list price honoured (note (x)). Returned E and, with the vote, rejected the lead's provisional default A. Its best line was against the default's own defence — "no repository figure uses a wider set… establishes only that widening has no immediate backward-compatibility cost. Indeed, that may be the least costly time to decide the policy." |
| 2026-07-26 | S033 | D-20260726-08 ratifying vote (google/gemini-3.6-flash, P2) | 0.030 (worst case) | 0.013969 | per-response usage.cost | in 2,183 / out 1,426, finish_reason: stop, provider Google. Returned E, agreeing with the review — the first unanimous ratification in this project; D-05, D-06 and D-07 all split. Recorded as agreement, not independent corroboration: the vote's reason restates the review's. |
| 2026-07-26 | S033 | E-20260726d-tierD-heldout pre-run critic (openai/gpt-5.6-terra, P1) on the frozen design | 0.130 (worst case from max_tokens) | 0.091200 | per-response usage.cost | in 9,015 / out 4,202, finish_reason: stop, provider OpenAI. Landed 30% inside a worst case built the way note (abc) requires — from the max_tokens the request actually permits (7,000), not from an assumed output length. The sixth consecutive critic pass to land inside its estimate. What the $0.09 bought: 22 findings and a NEEDS-REDESIGN verdict, 17 implemented before any perturbation existed. Two were defects in what the design said about its own rules — "the same rule at the same threshold" (false in proportion, sidedness, null probability and power at once) and a specificity "iff" that contradicted its own naturalness veto, leaving an outcome with no assigned reading. Note (rr) has now fired five sessions running, and five times on the same shape. It also caught that the worst case excluded retries and that a running-margin abort could truncate mid-stage despite a stage-boundary promise. |
| 2026-07-26 | S033 | the translation limb, the contamination gate, the design and everything else | 0.00 | 0.000000 | — | No API call. The 270-word Korolenko translation with its twelve-point frozen log, the blind download of Fell 1916, the contamination selection gate, the two-half split with its measured landmark check, the exact null-probability and power arithmetic, the never_names parameterisation with three new fixtures, and the v1→v2 amendment of eleven sections are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-26 | S034 | E-20260726d-tierD-heldout — the full four-stage Tier D run, 60 calls, P1/P2/P5 | worst case 3.55, central 0.755, reserved per stage | 0.725719 | per-response usage.cost; key-usage delta exact (20.79407595 → 21.519795134 = 0.725719) | 60/60 parsed, 0 failures, 0 retries, every finish_reason: stop. Landed 4% inside the central estimate and at 20% of the worst case — the third run to land inside its estimate, and the first built on per-juror max_tokens (P1 3,000 / P2 6,000 / P5 14,000) rather than one global cap sized for the heaviest juror. Stage actuals: sham $0.162415, held-out $0.192735, targeted heavy $0.196686, targeted light $0.173883. The full-stage reservation rule worked as designed (critic F2): each stage's whole worst case incl. its 10% retry allowance was reserved against the day's headroom before the stage was entered — $0.73 / $0.73 / $1.04 / $1.04 — and no stage was refused. No call on any juror billed above its declared per-call worst case (P1 max $0.01918 of $0.0538; P2 $0.02416 of $0.0503; P5 $0.02032 of $0.0520), so the tripwire on the price table never fired. Routing note (x): P1 all OpenAI; P2 Google and Google AI Studio; P5 across seven providers in 20 calls — Alibaba, Baidu, BaseTen, GMICloud, Ionstream, Novita, StreamLake — with no price excursion, which is the first time P5's routing spread has been this wide without a Venice-style overcharge. |
| 2026-07-26 | S034 | the build, the scoring, the verification and everything else | 0.00 | 0.000000 | — | No API call. The six-part §3.3 build audit, the 24 O4 perturbation sites with their Russian bases, the two count-neutral sham sets, the scorer, and the 71-check independent verifier — which recomputes the null probabilities by exhaustive enumeration of 3⁹ outcomes and the contamination gate with its own tokeniser — are all lead work. |
| 2026-07-26 | S035 | ARM-framework step 1 — the traceability inventory, the coverage measure and the FR→EN translation limb | 0.00 | 0.000000 | — | No API call. The 542-word George Sand translation with its twenty-one-point frozen log, the three-way contamination measurement against two published English translations, the sort of 17 result pages / 8 anchors / 10 sources / 2 theory pages into twelve candidate recommendations by evidence class, the nine claim pages and the 21-decision coverage mapping are all lead work. Lead translation is free and is never ledgered (charter §3, A4). The session had ~$3.79 of headroom and needed none of it: what it did was sorting and reading, and the one thing it could not do — give a jury verdict evidential weight — no amount of budget can currently buy. |
| 2026-07-27 | S036 | ARM-longwork step 1 — choosing the long work under the contamination gate, and span 1 of Verga | 0.00 | 0.000000 | — | No API call. The 2,288-word Italian→English span with its 19-point frozen log, the binding register, the R05 regime, the four-candidate survey and the three-cell contamination measurement (dependence_check.py, run locally) are all lead work. Lead translation is free and is never ledgered (charter §3, A4). Nothing this unit needed could be bought: the one thing that would have improved it — D. H. Lawrence's 1928 English, the comparator that matters — is a reachability problem, not a budget one. |
| 2026-07-27 | S037 | ARM-typology-logs step 4 — blind independent re-coding of 60 log decisions, two panel roles | 0.28 | 0.108360 | P1 openai/gpt-5.6-terra (provider OpenAI), P2 google/gemini-3.6-flash (providers Google, then Google AI Studio) | Two attempts, and the first was wasted — recorded because the estimate method is the thing that failed. Attempt 1 at max_tokens: 2000 cost $0.057517 and returned nothing usable: both models spent the whole cap on reasoning tokens that are billed and not returned, and came back finish_reason: length (P1 zero content, P2 a truncated JSON object). Attempt 2 at max_tokens: 10000 with reasoning: {"effort": "low"} cost $0.050843 and both completed finish_reason: stop. Note (abc) says build the worst case from the cap actually sent, which was done. The estimate was still wrong, and not for a new reason: note (b) — "reasoning models spend 1.7k-11k tokens before output; size max_tokens for reasoning plus output" — has been on the page since S010 and was not applied. The failure is a note that did not fire, not a note the project lacked; (b) is marked fired at S037 and (abc)'s wording is sharpened to say what "the cap allows" has to include. Raw JSON for both attempts preserved under E-20260727-log-decision-coding/run/. Actual came in at 39% of the $0.28 worst case even with the wasted attempt included. Key-usage delta unavailable as a cross-check — GET /api/v1/key returned an unchanged all-time figure (21.915886734) before, immediately after, and several minutes after the calls; per-request usage.include costs are the record and the missing sanity check is stated rather than assumed to agree. The session's translation limb (689 words of Bécquer, Spanish→English) and the whole reading of 31,500 words of translator's log cost $0 and are not ledgered (charter §3, A4). |
| 2026-07-27 | S038 | D-20260727-08 ratification — independent adversarial review (openai/gpt-5.6-terra, P1) | 0.195 (worst case from max_tokens) | 0.031278125 | per-response usage.cost | in 5,560 / out 927, finish_reason: stop, provider OpenAI, list price honoured (note (x), sixth run running). Returned B-with-amendment, rejecting S037's default of E and option A. Landed at 16% of the worst case. |
| 2026-07-27 | S038 | D-20260727-08 ratifying vote (google/gemini-3.6-flash, P2) | 0.086 (worst case from max_tokens) | 0.036393 | per-response usage.cost | in 6,267 / out 3,599 of which 2,943 reasoning, finish_reason: stop, provider Google AI Studio. Also B-with-amendment — the second unanimous ratification in this project — and it added the voice clause the review missed. The reasoning figure is the point: 2,943 unreturned-but-billed tokens is why S037's max_tokens: 2000 cap destroyed $0.057517, and why notes (b) and (abc) were applied here as a matter of course rather than as a repair. |
| 2026-07-27 | S038 | the whole principal unit — translation, anchor, instrument | 0.00 | 0.000000 | — | No API call. The 880-word English→Japanese translation with its 24-point frozen log, the sixth Tier 1 anchor with 51 machine-verified quotations, tools/dependence_check_cjk.py with eleven Japanese fixtures, the four-cell contamination measurement, the CJK decision-probe test and the null-probability recomputation that downgraded S037's headline are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-27 | S038 | S037 settlement, ledgered conservatively | — | 0.022149 | key-usage delta (21.915886734 → 22.046395584) | S037 recorded that GET /api/v1/key returned an unchanged all-time figure before, immediately after, and minutes after its calls, and said so rather than assuming agreement. S038's opening snapshot is 22.046395584, i.e. $0.130509 above S037's last reading against a per-request sum of $0.108360. The difference is most plausibly S037's own spend settling late — the same mechanism as S030 and S033, in the same direction — and is charged to the day rather than written off. Note (abf) has now fired twice and the practice that finds it is the opening snapshot. Two occurrences is not yet a pattern in the sense that matters (both are explicable by late settling of a known call); a delta in a session that made no call at all, as at S033, is the shape to watch. |
| 2026-07-27 | S041 | E-20260727b independent pre-run critic pass (openai/gpt-5.6-terra, P1) | 0.191 (worst case from max_tokens) | 0.065876875 | per-response usage.cost, and an exact key-usage delta | in 4,406 / out 3,474, finish_reason: stop, latency 48.2 s, temperature 0.2, max_tokens 12,000. Worst case built the way note (abc) requires — from the cap actually sent (12,000 × $15.00/M = $0.180) plus input (4.5k × $2.50/M = $0.011) — and sized for reasoning plus output per note (b). Landed at 34% of the worst case. It bought six mandatory fixes, one of which was a real bug in the frozen design: [.!?]+(?=\s\|$) had a markdown table-escape inside it and required a literal pipe, so the frozen sentence-counter did not count sentences. Standing disposition 9 exists for exactly that and the pass is what caught it. |
| 2026-07-27 | S041 | the whole principal unit — two translations, two regimes, the analysis, the verifier, the contamination measurement | 0.00 | 0.000000 | — | No API call. The paired JA→EN translation of Sōseki (1,210 + 1,206 words, 36 logged decisions across two frozen logs), regimes R06 and R04 v1.0, the arm-identifiability analysis over 36 stored translations and 96 stored jury payloads, the 112-check independent verifier and the three-cell contamination measurement are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-27 | S042 | ARM-longwork step 3 independent pre-run critic pass on the span-3 pre-registration (openai/gpt-5.6-terra, P1) | 0.206 (worst case from max_tokens) | 0.0642815625 | per-response usage.cost, and an exact key-usage delta | in 9,744 / out 6,541 of which 4,457 reasoning, finish_reason: stop, latency 80.4 s, temperature 0.2, max_tokens 12,000, provider OpenAI (list price honoured, note (x), seventh run running). Worst case built as note (abc) requires — from the cap actually sent (12,000 × $15.00/M = $0.180) plus input (9.7k × $2.50/M = $0.026) — and sized for reasoning plus output per note (b). Landed at 31% of the worst case. Raw: workshop/translations/jeli-il-pastore/R05-v1/critic/. It bought a redesign, not a polish: sixteen findings, fourteen mandatory, ten accepted and acted on before the freeze, including the withdrawal of a whole criterion that both this pre-registration and the frozen S039 one had been using to produce counts (note (bcl)). |
| 2026-07-27 | S042 | the whole principal unit — span 3, the recount, the verifier, the contamination re-run | 0.00 | 0.000000 | — | No API call. The 2,093-word Italian→English span with its 16-point frozen log (D39–D54), the widened criterion and its retrospective recount over 4,301 words of frozen prose, fid_sites_verify.py and its 237-candidate auditable enumeration, and the four-cell midpoint contamination measurement are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-27 | S043 | E-20260727c independent pre-run critic pass (openai/gpt-5.6-terra, P1) | 0.214 (worst case from max_tokens) | 0.051595 | per-response usage.cost | in 13,615 / out 4,043 of which 1,034 reasoning, finish_reason: stop, temperature 0.2, max_tokens 12,000, provider OpenAI (list price honoured, note (x), eighth run running). Worst case built as note (abc) requires — from the cap actually sent (12,000 × $15.00/M = $0.180) plus input (13.6k × $2.50/M = $0.034). Landed at 24% of the worst case. It bought a control the design did not have: it identified that condition B was not a manipulation of the rule alone but of rule-plus-authority-plus-length, and the sham condition added in response is what makes the result reportable. Nine MANDATORY findings, six accepted and acted on before dispatch, three declined in writing (critic/dispositions.md). Instrument caution: it fabricated seventeen item ids that do not exist, so its structural findings held on re-check and several of its examples did not. |
| 2026-07-27 | S043 | E-20260727c rating run — 3 raters × 4 conditions, 53 items each | 1.536 (worst case from max_tokens) | 0.234404 | per-response usage.cost, and an exact key-usage delta | 12 calls, all finish_reason: stop, temperature 0, max_tokens 12,000. P1 openai/gpt-5.6-terra (provider OpenAI), P2 google/gemini-3.6-flash (Google, Google AI Studio), P3 x-ai/grok-4.5 (xAI). Worst case built per note (abc) from the cap actually sent across all twelve calls at each model's out-price, plus input; landed at 15%. Key-usage cross-check exact to 1e-7 — 22.522355146 → 22.75675907, delta $0.2344039 against a per-request sum of $0.234404 — but only on re-read: the endpoint returned the unchanged opening figure immediately after the twelfth call, which is the S037 lag recurring. Method miss worth recording: critic.py printed its key snapshots instead of writing them to disk, so the critic call's own delta was lost when the terminal output was truncated and that call is on per-request cost alone. Snapshots belong in a file — note (bco). |
| 2026-07-27 | S043 | the whole principal unit — translation, item build, analysis, verifier | 0.00 | 0.000000 | — | No API call. The 2,281-character Japanese→English translation of Ōgai §4 with its 35-decision frozen log, its single-pass R06 draft frozen first, the 53-item build with its mechanical leak check, analyse.py and the 46-check independent verify.py are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
2026-07-27 day total: $0.614338 of $5.00 (S036 $0.000000; S037 $0.108360; S038 $0.067671 + $0.022149 settlement; S039 $0.000000; S040 $0.000000; S041 $0.065877; S042 $0.064282; S043 $0.285999). $4.385662 headroom remains today. Eighth session of this UTC day, and the largest single-session spend since S034 — bought 13 calls, of which the one that changed the design most was the cheapest.
S042 cross-check: EXACT, and the eighth exact one in this ledger. Opening snapshot 22.392956584, closing 22.457238146, delta 0.064281562 against a per-request sum of 0.0642815625 — agreement to 1e-9 on a single call. Both snapshots taken by rule (note (abf)).
S042 also records a third small inter-session drift, and it is a different animal from the S033 one. S041 closed at 22.385547484 and S042 opened at 22.392956584 — +$0.0074091 across a session boundary. It cannot be S041's spend settling late, because S041's own cross-check was already exact to 1e-9. It is therefore non-project key use, which this ledger has documented before at far larger magnitudes (+$0.359 between S015 and S020, +$0.36 between S002 and S010), and it is not charged to the day: per-request costs are primary and the delta is a sanity bound. The backlog row to watch is still the S033 shape — a delta inside a session that made no call at all — and that has occurred once. Recorded here so the distinction does not get lost by a later session reading three numbers and calling them a pattern.
S041 cross-check: EXACT, and the seventh exact one in this ledger. Opening snapshot 22.319670609, closing 22.385547484, delta 0.065876875 against a per-request sum of 0.065876875 — agreement to 1e-9 on a single call. Both snapshots were taken by rule (note (abf)); the opening one also shows no unattributed drift since S038's close plus S040's zero-spend session.
S039 spent nothing, and no snapshot was taken because no request was made. The session's principal unit was ARM-longwork span 2: a lead translation (free, never ledgered — charter §3, A4) wired to a study limb that read a public-domain Italian preface over plain HTTPS. There was nothing the API could do that the session needed. A $0 session is a normal outcome (continue-prompt.md §7), and this is the third in the ledger.
S038 cross-check: EXACT, and the sixth exact one in this ledger. Opening snapshot 22.046395584, closing 22.114066709, delta 0.067671125 against a per-request sum of 0.067671125 — agreement to 1e-9 across two calls to two providers. Taken at session end per note (zz) and S037's lagging-endpoint experience. It also settles the S037 row above: the endpoint tracks exactly when read late, so the $0.022149 gap is late settling of S037's own calls and not an unattributed charge — which is why the row is written as a settlement rather than as a repeat of S033.
2026-07-26 day total: $1.206080 of $5.00 (S027 $0.00 + S028 $0.080616 + S029 $0.070634 + S030 $0.081006 + S031 $0.091925 + S032 $0.00 + S033 $0.142225 + $0.013955 unattributed + S034 $0.725719). ~$3.79 headroom remains today. Nine sessions in one UTC day, of which three (S027, S032, S035) spent $0.00 — a $0 session is a normal outcome, not a shortfall.
S034 cross-check: exact, and the fifth exact one in this ledger. Session-start snapshot 20.79407595 — identical to S033's closing snapshot, so the unattributed drift S033 recorded did not recur. End snapshot 21.519795134, delta 0.725719, against a per-request sum of 0.725719: agreement to 1e-6 across 60 calls to three labs and nine distinct providers. Note (abf) — snapshot at session start as well as end — has now been applied twice and is what would have caught a second occurrence. One occurrence remains one occurrence; a third would be a pattern.
S033 cross-check: exact, and the fourth exact one in this ledger. Session-start snapshot 20.65185145, end 20.79407595, delta 0.142225, against a per-request sum of 0.142225 — agreement to 1e-6 across three calls to two providers. The unattributed row above is outside that check by construction: it is the gap between the previous session's closing snapshot and this session's opening one, and taking an opening snapshot is what made it visible at all.
S030 cross-check note — the first inexact one, and the reason is known. Per-request usage.cost sums to $0.075260 (0.0656205 successful + 0.0096395 measured for the lost call); the session key-usage delta is $0.081006 (20.45679045 → 20.53779685), $0.005746 higher. The gap is attributed to the lost call, whose cost was snapshotted while OpenRouter was still holding the connection open. The ledger takes the delta, being the larger. Standing note (x) held on the successful call: provider read off the response, OpenAI, list price. New: a snapshot taken immediately after a call that returned no body can under-report that call.
2026-07-25 day total: $4.838052 of $5.00 (S014 $2.248649 + S015 $1.150577 + S020 $0.755259 + S021 $0.115064 + S022 $0.497441 + S025 $0.071062; S016–S019, S023, S024 and S026 $0.00 each). ~$0.16 headroom left today. Lesson reinforced (S010): moonshotai/kimi-k3 is the priciest juror by far (~$0.06–0.11/reasoning-heavy call, ~10–20× deepseek); a 60-call jury pass on kimi alone ran ~$1.0. Budget kimi-inclusive jury runs at ~2–3× the naive estimate, or drop kimi from high-volume passes. S020 dropped kimi and P3 from a 60-call pass for exactly this reason, and recorded the exclusion as a power limitation rather than a finding about those models. Key-usage cross-check for S020: performed and exact (delta 0.755258 vs per-request sum 0.755259); the S014/S015 cross-checks remain as recorded above.
S020 estimate note, worth keeping. The Tier D pilot came in 9% under its central estimate and 42% under its stated worst case ($0.755 against $0.829 / $1.293) — the second run in this project to land inside its estimate, after S015. What changed: the estimate was stated as a range with a worst case built from S014's per-call maxima, not means, at the independent critic's insistence (critic.md E27), and the run was sized so the worst case fitted inside headroom rather than the point estimate. The pre-dispatch guard never fired. The two runs that overran (S010, S014) both sized against means only.
S022 spend notes — two of them are warnings
Key-usage cross-check: exact. Session start snapshot 19.171857368, end 19.669298425, delta 0.497441 against a per-request sum of 0.497441 — agreement to 1e-6, the third exact cross-check in this ledger. It also settles a loose end: one review call died in the client while parsing the response (OpenRouter interleaves keep-alive comment lines on slow generations, which tools/ratify_vote.py did not strip). The delta shows that call was never billed. The tool now strips comment lines and preserves every raw body to <out>.raw.
Warning 1 — the list price in config/models.md is not what gets billed. The deepseek/deepseek-v4-pro review cost $0.044837 on 13,556 in / 6,807 out. At the listed $0.435/$0.87 per M that should have been ~$0.012. The response's provider field reads Venice: OpenRouter routed the slug to a provider charging roughly $1.65/$3.30 per M, 3.8× the figure the panel was partly selected on ("near-frontier quality at ~1/10 frontier price — the high-volume workhorse", config/models.md). P5 is not reliably the cheap juror. Any estimate built from the table in config/models.md can be wrong by ~4× through routing alone, and the provider field should be read off every response and recorded.
Warning 2 — 38% of this session's spend bought nothing. $0.190665 of $0.497441. Neither cause was subtle: a shell variable-expansion bug, and max_tokens sized without the reasoning budget that this ledger has warned about since S010. The first is new; the second is the fourth session to be bitten by it. The mitigation that finally worked cost one sentence of prompt — "Be brief. Answer Task B FIRST, then Task A. Quote only the decisive words" — and cut reasoning from ~5,800 tokens to ~1,800, halving cost while making the answers sharper. Two lessons worth more than the money: (i) instruct brevity explicitly on reasoning models; (ii) order the tasks so that what truncates first is what you can afford to lose.
| 2026-07-28 | S047 | E-20260728d-jeli-marks pre-run critic — openai/gpt-5.6-terra, succeeded first call | 0.131 worst case from max_tokens 8000 at list out-price plus the prompt at list in-price (note (abc)) | 0.036104375 | per-response usage.cost, and an exact key-usage delta: 23.981018277 + 0.036104375 = 24.017122652, the panel run's opening snapshot, to 1e-9 | in 4,493 / out 3,878, finish_reason: stop, provider OpenAI (list price honoured, note (x), ninth run running), reasoning: {"effort":"low"}. 28% of worst case, the fifth session running to land near a quarter or fifth of a max_tokens-built estimate. No fall-through needed. Verdict NEEDS-REDESIGN, ten findings, ten accepted, four by withdrawing a claim the design had made — and the one option it forced into the answer scheme is what falsified the lead's headline prediction. |
| 2026-07-28 | S047 | E-20260728d-jeli-marks — the guillemet census: 10 stimuli × 3 models, 30 calls | 0.783 worst case from max_tokens 1500 across all thirty at each model's out-price plus prompts at list in-price (note (abc)) | 0.365797975 | per-response usage.cost summed in runs/cost.json, cross-checked against the call count by verify.py, and a settled key-usage delta of 0.365797973 — residual −2e-9 | 30 of 30 finish_reason: stop, 30/30 parsed, temperature 0, reasoning: {"effort":"low"}. P1 openai/gpt-5.6-terra (OpenAI), P2 google/gemini-3.6-flash, P3 x-ai/grok-4.5. 47% of worst case — the highest fraction in five sessions, and it is the input side: the largest stimulus carries an 11,661-word prefix, so max_tokens is a small part of the bill. The cross-check took three reads to settle and the intermediate residual was exactly the thirtieth call's billed cost — note (bco)'s lag, for the first time diagnosed to a named call rather than left as a gap. |
| 2026-07-28 | S047 | the whole translation limb and both censuses | 0.00 | 0.000000 | — | No API call. Verga span 5 (2,806 IT → 3,202 EN, log D70–D84), the completion of an 11,474-word novella, the guillemet census, the 24-site reduplication census that falsified register row V10, build_stimuli.py, and verify.py's 41 checks are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-28 | S046 | E-20260728c-length-matching pre-run critic — openai/gpt-5.6-terra, succeeded first call | 0.14 worst case from max_tokens 8000 at list out-price plus the prompt at list in-price (note (abc)) | 0.029074 | per-response usage.cost; opening and closing key snapshots persisted to runs/snap/ (note (bco)) | in 4,995 / out 2,836, finish_reason: stop, provider OpenAI, reasoning: {"effort":"low"} sent and honoured. 21% of worst case, the fourth session running to land near a fifth of a max_tokens-built estimate. No fall-through was needed — the runner chains to qwen/qwen3.7-max on an empty or length return rather than retrying the same slug (note (b), S045's own lesson) and the chain stopped at the first call. Verdict NEEDS-REDESIGN, nine findings, all accepted in substance before the run. Standing note (x) holds again: provider OpenAI, list price. |
| 2026-07-28 | S046 | the whole principal unit besides that critic pass | 0.00 | 0.000000 | — | No API call. Both limbs of E-20260728c are local computation and lead translation: the S010 re-analysis (permutation nulls, cluster bootstraps, 199 verification checks), the project's first German→English translation rendered three times with 38 logged decisions, tools/metric_a.py, and the two-way recount that retired a backlog row. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-28 | S045 | E-20260728b-forced-run-ru pre-run critic — deepseek/deepseek-v4-pro, FAILED | 0.027 (worst case from max_tokens 6000 at the S022 routed price) | 0.019293024 | per-response usage.cost | in 4,518 / out 6,000, finish_reason: length, content: null — the entire budget spent on unreturned reasoning. Provider Novita. Standing note (b), sixth session bitten. Dropped rather than retried, per S044's own lesson; the substitute's first attempt overwrote its raw body, a defect in this session's runner, recorded on the result page §7. |
| 2026-07-28 | S045 | same critic pass — qwen/qwen3.7-max (declared first reserve), succeeded | 0.042 | 0.029884975 | per-response usage.cost | in 4,775 / out 5,162, finish_reason: stop, 94s, provider Alibaba, reasoning: {"effort":"low"} sent and evidently honoured. P1/P2/P3 were subjects in this design and could not critique it; P4 is off the call list. Returned nine accepted findings and one rejected; three of them changed the scoring code before dispatch. |
| 2026-07-28 | S045 | E-20260728b-forced-run-ru — the probe: 5 conditions × 3 models, 15 calls | 0.371 worst case for conditions A–D from max_tokens 3000 at each model's out-price (note (abc)) | 0.130195 | per-response usage.cost, summed in runs/cost.json and cross-checked against the call count by verify.py | A–D $0.084 = 23% of worst case; condition E (post-hoc) $0.046. 14 of 15 stop; E/P2 truncated at length, which can only remove answers. Two cost rows were initially missing — a dispatch the harness timed out had written the response files, and probe.py's cache branch then skipped them without ledgering; recovered from the stored raw bodies, and the verifier's call-count check is what caught it. Without it the session would have under-reported by $0.0287. |
| 2026-07-28 | S045 | the whole translation limb — Verga span 4, its log, the register revision, and the Dole gate | 0.00 | 0.000000 | — | No API call. 2,274 IT → 2,572 EN, fifteen logged decisions D55–D69, register erratum E5, and the contamination measurement against Dole 1896 are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-28 | S044 | D-20260727-09 ratification — independent adversarial review (deepseek/deepseek-v4-pro) | 0.009–0.049 (worst case from max_tokens 8000 at the S022 routed price) | 0.017041908 | per-response usage.cost | in 11,918 / out 1,575, finish_reason: stop. Provider Ionstream — not the list price: at the table's $0.435/$0.87 this should have been ~$0.0066, so the routed price is ~2.6× list. Standing warning (S022 warning 1) holds for the fourth time; the estimate's worst case was built on the routed price and the run landed at 35% of it. Verdict D, rejecting the lead's default of B. |
| 2026-07-28 | S044 | E-20260728-forced-run — the forced-vs-recalled control, 3 calls × 4 Chinese sentences | 0.02 (worst case 0.091 from max_tokens 3000 across the three at each model's out-price) | 0.0166679 | per-response usage.cost, summed; snapshots written to runs/cost.json per note (bco) | P1 openai/gpt-5.6-terra (OpenAI) 0.002735 · P2 google/gemini-3.6-flash (Google AI Studio) 0.0094305 · P3 x-ai/grok-4.5 (xAI) 0.0045024. All finish_reason: stop, temperature 0. 18% of worst case. Cheapest call of the session and the one that changed an artifact's declaration: contamination: high, RS-20260728-forced-run. |
| 2026-07-28 | S044 | D-20260727-09 ratifying vote — moonshotai/kimi-k3, attempt 1, FAILED | 0.131 (worst case from max_tokens 6000) | 0.130665 | per-response usage.cost | in 13,555 / out 6,000, finish_reason: length, content: null. The entire output budget was consumed by unreturned reasoning. Standing note (b), fifth session bitten, and the note is seventeen sessions old. Raw preserved as …__vote__kimi-k3.attempt1.json. |
| 2026-07-28 | S044 | same vote — moonshotai/kimi-k3, attempt 2 with reasoning: {"effort":"low"}, ALSO FAILED | 0.161 (worst case from max_tokens 8000) | 0.160917 | per-response usage.cost | in 13,639 / out 8,000, again finish_reason: length, content: null, 909s. The reasoning.effort parameter was not honoured by provider DigitalOcean for this slug — the remedy note (b) prescribes did not work, and the call landed at 100% of its worst case, the only such row in this ledger. tools/ratify_vote.py gained --reasoning-effort for this attempt and keeps it; the flag is right and this slug is not to be trusted with it. |
| 2026-07-28 | S044 | same vote — qwen/qwen3.7-max (first reserve, config/models.md), succeeded | 0.05 | 0.03295445 | per-response usage.cost | in 11,554 / out 3,596, finish_reason: stop, 69s. Prompt trimmed by the frozen experiment design (~3k tokens) which the vote did not need. Verdict C, disagreeing in writing with the reviewer's D and with the lead's default B. |
| 2026-07-28 | S044 | the whole principal unit — the Chinese primaries, the source page, the translation, the theory amendment | 0.00 | 0.000000 | — | No API call. Reading Lu Xun 1930 whole and his 1931 reply in Chinese, the 1,999-character Chinese→English translation of §五 with its thirty-three-point frozen log and its separately frozen R06 draft, the four Venuti access checks, S-luxun-yingyi, the TH-20260724 C1 scope condition, and score.py are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-28 | S050 | E-20260728g-scale-usage independent pre-run critic — x-ai/grok-4.5 (P3), succeeded first call | 0.058 worst case from max_tokens 8000 at list out-price plus the prompt at list in-price (note (abc)) | 0.0197964 | per-response usage.cost, and an exact key-usage delta: opening 24.914190991, closing 24.933987391, difference 0.019796400 — agreement to 1e-9 | in 5,411 / out 1,532, finish_reason: stop, 35s, provider xAI, temperature: 0.2, reasoning: {"effort":"low"}. 34% of worst case. No fall-through needed (reserve qwen/qwen3.7-max unused). P3 rather than P1 because P1, P2 and P5 are the three jurors whose stored scores are the entire dataset — a subject cannot critique the analysis of itself; P4 is off the call list (note (b)). Verdict NEEDS-REDESIGN, nine findings, all nine accepted, five by withdrawing or narrowing a claim the design had made — including its central prediction. Note (x) checked on P3 for the first time: list price honoured. |
| 2026-07-28 | S050 | the whole translation limb, the contamination gates, the fault variant, and every analysis | 0.00 | 0.000000 | — | No API call. Andreyev «Баргамот и Гараська» rendered from the Russian (486 RU → 688 EN, 29-decision frozen log, R06 draft frozen in two separate commits), both contamination gates and their 47,453-word null control, analyse.py, the 332-check independent verify.py, and build_variant.py are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-28 | S052 | E-20260728j-classb-marking independent pre-run critic — openai/gpt-5.6-terra (P1), succeeded first call | 0.140 worst case, stated with the arithmetic at the critic's own insistence (finding E): prompt capped at 8,000 tok × $2.50/M = $0.020 plus max_tokens 8,000 × $15.00/M = $0.120; reserve qwen/qwen3.7-max $0.060; total cap $0.200 | 0.024405625 | per-response usage.cost; opening and closing key snapshots persisted to runs/snap/ (note (bco)) — delta 0.000000, the endpoint had not settled, note (bcx) | in 7,657 / out 1,659, finish_reason: stop, 21s, provider OpenAI (list price honoured — note (x), twelfth run running), temperature: 0.2, reasoning: {"effort":"low"}. 17% of worst case. No fall-through. No panel model is a subject in this design — it is lead translation plus a public-domain comparator — so P1 was free to be chosen. Verdict NEEDS-REDESIGN, ten findings, nine accepted, one accepted with a reasoned substitution, one clause declined in writing, and four of the amendments withdraw or narrow a claim the design had made, including the sentence its closing use rested on. Note (rr), tenth consecutive session. |
| 2026-07-28 | S052 | the whole translation limb, the comparator read, and every analysis | 0.00 | 0.000000 | — | No API call. The forced re-translation of the nine Class B sites and the continuous forced ¶42, the Dole 1896 fetch and read, analyse.py, the 23-pair paragraph-bounded recount, and the independent verify.py (91 checks, 0 failures) are all lead work, a public-domain HTTPS GET, or local arithmetic. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-28 | S051 | E-20260728i-coverage-replication independent pre-run critic — google/gemini-3.6-flash (P2), succeeded first call | 0.072 worst case from max_tokens 8000 at list out-price plus the prompt at list in-price (note (abc)) | 0.0214695 | per-response usage.cost; opening and closing key snapshots persisted to runs/snap/key-usage.json (note (bco)) | in 7,153 / out 1,432, finish_reason: stop, 10s, provider Google, temperature: 0.2, reasoning: {"effort":"low"}. 30% of worst case. No fall-through (reserve qwen/qwen3.7-max unused). P2 rather than P1 or P3 because those two are the readers this design creates — an instrument cannot critique itself; P4 and P5 are off the call list (note (b)). Verdict NEEDS-REDESIGN, nine findings, seven accepted, one accepted-as-limitation, one declined in writing, and two of the accepted ones withdraw a claim the design had made — including its central reading. Note (rr), ninth consecutive session. |
| 2026-07-28 | S051 | E-20260728i the two independent readers — openai/gpt-5.6-terra (P1) and x-ai/grok-4.5 (P3), 63 classifications each, one call apiece | 0.293 worst case from max_tokens 12000 across both at each model's out-price plus prompts at list in-price (note (abc)) | 0.0678982125 | per-response usage.cost, both cross-checked against analysis/results.json by verify.py | P1 $0.0345078125 (provider OpenAI, list price honoured — note (x), eleventh run running), in 8,828 / out 2,762. P3 $0.0333904 (provider xAI, list honoured), in 8,803 / out 2,667. Both finish_reason: stop; both returned all 63 ids exactly once with no label outside the scheme. Temperature 0, reasoning: {"effort":"low"}. 23% of worst case. |
| 2026-07-28 | S051 | the whole translation limb, the instrument gate, the contamination measurement, and every analysis | 0.00 | 0.000000 | — | No API call. 高瀬舟 spans A and B rendered from the Japanese (1,506 source characters → 895 draft / 905 revised English words, 42-decision log across two separately frozen artifacts), the fetch_aozora.py gaiji repair with its 12-check fixture file, the 22-deletion audit across four canon sources, the contamination gate with its three null controls, build_materials.py, analyse.py and the 34-check independent verify.py are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-29 | S053 | E-20260729-drift-window-verify independent pre-run critic — google/gemini-3.6-flash (P2), succeeded first call | 0.075 worst case from max_tokens 8000 at list out-price plus the prompt at list in-price (note (abc)) | 0.030753 | per-response usage.cost; opening/closing key snapshots persisted to runs/snap/ (note (bco)) | in 7,657 / out 2,557, finish_reason: stop, 14s, provider Google AI Studio, temperature: 0.2, reasoning: {"effort":"low"}. 41% of worst case. P2 because it is the only panel member that was a subject in nothing at design time. Verdict NEEDS-REDESIGN, five findings, all five accepted, four by changing the procedure and one by withdrawing a claim the design made about itself — a temperature: 0 repeat is a test of backend determinism, not of judgment variance. Note (rr), eleventh consecutive session. |
| 2026-07-29 | S053 | E-20260729 drift classification — 52 lemmas, no renderings shown, P1 and P3 | 0.132 worst case from max_tokens 6000 across both at each model's out-price plus prompts at list in-price | 0.051231 | per-response usage.cost, summed in runs/cost.json, cross-checked against the call count by verify.py | P1 openai/gpt-5.6-terra (OpenAI, list price honoured — note (x), thirteenth run running) $0.0338946875; P3 x-ai/grok-4.5 (xAI, list honoured) $0.0173364. Both stop, all 52 ids returned exactly once, no label outside the scheme. 39% of worst case. This is the pair of calls that produced the session's largest finding: four-class agreement 0.654. |
| 2026-07-29 | S053 | E-20260729 reflex scoring — 4 calls (2 passages × 2 raters), 152 cells, every judgement carrying an attesting quotation | 0.540 worst case from max_tokens 12–16k across the four | 0.142415 | per-response usage.cost, summed; every cell re-verified from the raw bodies by verify.py | P1 $0.032089 + $0.040330; P3 $0.023018 + $0.046978. 4 of 4 stop; 152 of 152 cells returned, 0 missing, and 152 of 152 quotations attest verbatim against the stored files. 26% of worst case. |
| 2026-07-29 | S053 | E-20260729 presentation-order repeat — P1, same 23 sites in a permuted order | 0.210 worst case from max_tokens 16000 | 0.039701 | per-response usage.cost | in 6,617 / out 2,215, stop, provider OpenAI. 19% of worst case. 91 of 92 cells identical to the first pass. What it does not buy is stated on the result page: the pre-run critic established before dispatch that a byte-identical temperature: 0 repeat measures backend determinism, so this measures order-invariance and this project still has no estimate of rater sampling variance. |
| 2026-07-29 | S053 | E-20260729 independent site census — qwen/qwen3.7-max, ABANDONED | 0.100 worst case from max_tokens 8000 | 0.000000 | key-usage delta, re-read after the run and after the client was killed | Dispatched, no response body after ~20 minutes, client killed, nothing billed at the closing snapshot. This is not note (b)'s shape — that is finish_reason: length with empty content, which does bill. A new shape: a call that hangs open. ⚠ If it settles later it will appear as an unattributed key-usage delta in a session that did not make it, which is the exact shape the retired S033 backlog row revives on. Recorded here so the next session reading a non-zero opening drift knows where to look first. |
| 2026-07-29 | S053 | E-20260729 same census, reserve google/gemini-3.6-flash, succeeded first call | 0.075 | 0.055638 | per-response usage.cost | in 2,217 / out 6,975, stop, 31s, Google AI Studio. 74% of worst case — the highest fraction in this ledger since S048's 49%, and the reason is the same: the output is the whole product (32 enumerated sites with senses and classes) on a short prompt. Declared role collision: the reserve for the census was P2, which had already made the pre-run critic call, so the model that told the design its census was one reader's is the model that produced the alternative census. The design named this fall-through in advance; the result page reports it as a limit on §7. |
| 2026-07-29 | S053 | the whole translation limb, the site census, the glossary probe, the corrections and every analysis | 0.00 | 0.000000 | — | No API call. Beowulf ll. 1800–1866 rendered from the Old English (67 lines → ~600 English words, frozen decision log, census frozen and committed before translating), the frozen reflex census, glossary_probe.py, analyse.py, the 102-check independent verify.py and its mutation test, and the hand-verification of the three corrections against the stored public-domain files are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-29 | S059 | E-20260729g-graded-senses independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call | 0.28 worst case from max_tokens 16000 at list out-price plus the prompt at list in-price (note (abc)) | 0.137502 | per-response usage.cost; opening/closing key snapshots persisted to runs/snap-critic-*.json (note (bco)) | in 4,276 / out 5,256, stop, 93s, provider Fireworks. 49% of worst case — the highest fraction in this ledger, and for the same reason as S057's 35%: P4 returns long. P4 because P1/P2/P3 are the raters and P5 is the site extractor, so no model both critiques the instrument and is it. Verdict NEEDS-AMENDMENT, nine findings — three BLOCKING, four MANDATORY, two ADVISORY — all nine accepted. Note (rr), seventeenth consecutive session, and its most structural return yet: it demoted the design's headline statistic out of the disposition gate on a metric-mismatch argument, rebuilt the load-bearing control as a four-way derived label, and forced the held-out material's generative procedure to be frozen in its own commit. Both of the two checks it added fired on the data. |
| 2026-07-29 | S059 | E-20260729g KO site extraction — P5 deepseek/deepseek-v4-pro, Korean source alone, brief frozen verbatim beforehand | 0.09 worst case from max_tokens 16000 | 0.020847502 | per-response usage.cost | in 3,124 / out 15,108, stop, 250s, provider Baidu — a seventh provider for this slug. 23% of worst case, and note (b) did NOT fire at 16,000, which is note (bdl)'s remedy holding for the third session. Returned 23 sites, 22 verbatim in the source, one dropped for not matching character-for-character (고구라 양복, which the source prints as <고구라> 양복). |
| 2026-07-29 | S059 | E-20260729g the rating pass — 3 conditions × 3 raters × 2 byte-identical replicates = 18 calls, 62 items each | 1.73 worst case from max_tokens 10,000 ×12 and 6,000 ×6 | 0.4052433375 | per-response usage.cost, summed in runs/run-costs.json, recomputed from the 18 bodies by verify.py | P1 openai/gpt-5.6-terra (OpenAI ×5, Azure ×1) $0.1445684; P2 google/gemini-3.6-flash (Google ×3, Google AI Studio ×3) $0.1531605; P3 x-ai/grok-4.5 (xAI ×6) $0.1075144. 18 of 18 stop first time, 62 of 62 items on every one, no reserve fired — note (b) did not fire, the second session running. 23% of worst case. The replication is the point: nine of these eighteen calls exist to discharge the defect ARM-sense-boundary returned to the backlog on closing, and they corrected that arm's closing sentence. |
| 2026-07-29 | S059 | the whole translation limb, the contamination gate, the item build and every analysis | 0.00 | 0.000000 | — | No API call. 현진건 「운수 좋은 날」 rendered from the Korean (931 → 1,803 words, the project's first Korean), the span fixed by a mechanical rule frozen before the text was opened, the comparator English stored unread and opened only by tools/dependence_check.py (10-token run, clean, no shared 12-gram), and build_items.py, analyse.py, the 84-check independent verify.py and its mutation test are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
2026-07-29 day total: $0.319738262 of $5.00 (S053 alone; first session of this UTC day). $4.680262 headroom remains.
S053 spend note — nine dispatches, eight bodies, and the cross-check is exact for the third session in four. Opening snapshot 25.515595228, closing 25.835333489, delta 0.319738261 against a per-request sum of 0.319738262 — a residual of 1e-9, and the same agreement re-read after the abandoned client was killed, which is what establishes that the hung qwen call has not billed. Three things are worth keeping. (i) Note (b) did not fire and a new failure shape did: not an empty billed response but a call that never returns at all, and the declared reserve chain handled it in 31 seconds. (ii) The worst-case fractions were 19–41% on the output-dominated calls and 74% on the one whose output is the deliverable, which is the S048 structural pattern rather than a missed estimate. (iii) $0.030753 — under a tenth of the session's spend — bought five accepted findings, one of which withdrew the design's own claim about its repeat condition before the repeat was run. Note (rr), eleventh consecutive session.
2026-07-28 day total: $1.620554 of $5.00 (S044 $0.358246 + S045 $0.179373 + S046 $0.029074 + S047 $0.401902350 + S048 $0.051713 + S049 $0.466676 + S050 $0.0197964 + S051 $0.0893677125 + S052 $0.024405625; nine sessions in this UTC day, two of them concurrent). $3.379446 headroom remains. S052 made one call and the session's entire substance — a translation limb, a first comparator read, a 23-pair recount and a 91-check verifier — cost nothing.
S050 spend note — the cheapest session with an API call since S046, and the first exact cross-check since S047. One call, $0.0197964, 0% wasted: note (b) did not fire, the fall-through chain was never entered, and finish_reason was stop on the first attempt. Two things are worth keeping. (i) The key-usage delta and the per-request cost agree to 1e-9 — opening and closing snapshots taken 90 minutes apart, no settling lag to wait out, which is what note (bco) predicts for a single short call and had not previously been demonstrated on one. (ii) The call bought a refutation of the session's own headline prediction before it was tested, which is the best available return on a critic pass and the eighth consecutive session on which note (rr) has fired.
S049 spend note — $0.039651 of $0.405701 (10%) bought nothing, and the shape is now eight sessions old. deepseek/deepseek-v4-pro returned finish_reason: length with no content on a 12,000-token cap with reasoning: {"effort":"low"} set — the parameter was sent and did not save it, exactly as at S044 on a different lab. What worked, for the third session running, was falling through to a declared reserve rather than retrying. Two other things are worth keeping. (i) The ledgered figure is the key delta ($0.466676), not the per-request sum ($0.405701), because one call reported a cost of exactly zero after consuming 5,972 output tokens; the S030 precedent is followed and both readings of the $0.060975 gap are written down rather than one being asserted. (ii) The rating run landed at 15% of a max_tokens-built worst case, identical to S043's fraction on the same experiment shape — which is the first time this ledger has had two runs similar enough for that comparison to mean anything.
S047 spend note — the largest single-session spend since S034, and every cent of it bought a measurement that came out against the lead. $0.401902350 across 31 calls, 31 of 31 finish_reason: stop — the first session in four with no wasted call at all, note (b) not firing once. Two things are worth keeping. (i) The worst-case fraction jumped from ~21–28% to 47%, and the reason is structural rather than a miss: this run's cost is dominated by input, not output — ten stimuli whose prefixes run from 896 to 11,661 words — so a worst case built from max_tokens (note (abc)) is a much tighter bound here than on the short-prompt runs that produced the 15–31% band. The note still works; the band was never a target. (ii) The key-usage cross-check settled only on the third read, and the intermediate residual was exactly one named call. Note (bco) has said since S043 that the endpoint lags and that a completion-instant snapshot is worthless; this is the first run that watched it converge and could name what was missing.
S045 spend note — the two failures are the same failure, sixteen sessions apart. $0.019293 of $0.179373 (11%) bought nothing, from deepseek/deepseek-v4-pro returning finish_reason: length with no content: standing note (b), sixth session, and the same shape as S044's $0.291582 kimi loss. The pattern across the two sessions is now specific enough to state as a rule: a reasoning-heavy model given a long prompt and a max_tokens in the low thousands returns nothing at all, and reasoning.effort is honoured by some providers and silently ignored by others. What worked both times was the declared first reserve — qwen/qwen3.7-max, $0.033 at S044 and $0.030 here, stop both times. The probe itself landed at 23% of its worst case, which was built from max_tokens per note (abc) for the third session running.
S044's original line, kept: $0.358246 of $5.00 for S044 alone. $0.291582 of it — 81% — bought nothing, in two calls to one model that returned no content. That is the largest waste in this ledger by a wide margin, and it is not a new failure mode: it is standing note (b), unapplied at attempt 1 and applied and ineffective at attempt 2. The lesson is narrower than "size for reasoning": a provider may silently ignore reasoning.effort, so the parameter is not a guarantee and a model that has once returned finish_reason: length on a long prompt should be dropped rather than retried. The successful vote, from the declared first reserve, cost $0.033 and came back in 69 seconds.
S044 cross-check: NOT exact, and the gap is attributed rather than absorbed. Opening snapshot 23.07324277 (00:50 UTC), closing 23.477710928, re-read after 30s and unchanged — delta $0.404468158 against a per-request sum of $0.358246258, a gap of $0.046222. The per-request sum is what is ledgered, per this page's stated method, and the gap is attributed to concurrent non-project use of the key. The evidence for that is on this page rather than assumed: the opening snapshot was already $0.31648 above S043's closing snapshot, and S043's own cross-check was exact to 1e-7, so that drift cannot be S043 settling late. A key that drifted $0.316 between sessions drifting a further $0.046 across a 42-minute session is the same phenomenon at the same order of rate. This is not the S033 shape — that was a delta in a session which made no call at all, and the retired backlog row revives only on that shape. Recorded so a later session reading three inexact cross-checks does not have to rediscover why.
| 2026-07-28 | S048 | E-20260728e-venuti-specifiability independent coverage classifier — openai/gpt-5.6-terra (P1), succeeded first call | 0.105 worst case from max_tokens 6000 at list out-price plus the prompt at list in-price (note (abc)) | 0.051713 | per-response usage.cost; opening and closing key snapshots persisted to runs/snap/ (note (bco)) — delta 0.000000, and see the snapshot rows: the endpoint had not settled | in 6,726 / out 5,494 (4,840 reasoning), finish_reason: stop, provider OpenAI, temperature: 0, reasoning: {"effort":"low"}. 49% of worst case, after four sessions near 20% — the gap is reasoning tokens. No fall-through needed (reserve x-ai/grok-4.5 unused). The call fired the design's pre-registered reliability failure criterion: agreement with the lead's coding 0.483 against a 0.60 threshold. Standing note (x) holds: provider OpenAI, list price. |
| 2026-07-28 | S048 | the rest of the principal unit | 0.00 | 0.000000 | — | No API call. Reading a 366-page book, two lead translations of one passage under two frozen regimes with 89 logged decisions, analyse.py (46 checks, 0 failures), and the rebuild of S-venuti-invisibility on the primary. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-28 | S049 | E-20260728f-nonlead-items site extraction — one call to a model that is neither a rater nor the critic, so the lead does not choose the sites. Three attempts, all preserved | 0.034 worst case from max_tokens 12000 at the S022 routed price plus the prompt at list in-price (note (abc)) | 0.099778 | per-response usage.cost, summed over all three attempts in extract/cost-all.json | deepseek/deepseek-v4-pro (provider Alibaba): finish_reason: length, no content at all, 4,000 in / 12,001 out, the entire cap spent on unreturned reasoning despite reasoning: {"effort":"low"} — $0.039651 for nothing, standing note (b), seventh session bitten. z-ai/glm-5.2 (Alibaba): finish_reason: error, body truncated mid-object at 4,776 characters, billed $0.000000 on 5,972 output tokens — see the snapshot rows. qwen/qwen3.7-max (Alibaba, stop, 4,050 in / 12,238 out): $0.060127, accepted. The runner's own acceptance test was defective and would have taken the truncated body; fixed to require stop and a closed array before the accepted call was made — new note (bdb). Over its worst case, because the worst case was priced for one call and three were made; the per-call figure landed inside it. |
| 2026-07-28 | S049 | E-20260728f independent pre-run critic pass — openai/gpt-5.6-terra (P1), succeeded first call | 0.191 worst case from max_tokens 10000 at list out-price plus the prompt at list in-price (note (abc)) | 0.082018 | per-response usage.cost; snapshots persisted to runs/snap/ (note (bco)) | in 16,391 / out 7,521 of which 1,552 reasoning, finish_reason: stop, provider OpenAI, list price honoured (note (x), tenth run running), temperature: 0.2, reasoning: {"effort":"low"}. 43% of worst case. Ten findings accepted, three declined in writing (critic/dispositions.md). The best of them — note (rr), seventh consecutive session — was that an amendment the session had written that hour asserted something false about its own statistic: A1 called Q3 evaluable when Q3 is stated on the stratum that came back empty. Instrument caution, recurring from S043: it cited 27 item id / position pairs and 12 were wrong; every structural claim was re-derived before being accepted and the structure held. The critic is also rater P1 — declared, with the reason (every non-rater model available had just failed this task) and with the mitigation reported (all three rater pairs separately; no overlap signature visible). |
| 2026-07-28 | S049 | E-20260728f rating run — 3 raters × 4 conditions, 40 items each, 12 calls | 1.52 worst case from max_tokens 12000 across all twelve at each model's out-price plus prompts at list in-price (note (abc)) | 0.223905 | per-response usage.cost summed in runs/cost.json, cross-checked against the call count by verify.py | 12 of 12 finish_reason: stop, 480 of 480 judgments inside the five-way scheme, 0 missing. Temperature 0, reasoning: {"effort":"low"}. P1 openai/gpt-5.6-terra (OpenAI), P2 google/gemini-3.6-flash (Google, Google AI Studio), P3 x-ai/grok-4.5 (xAI). 15% of worst case — the same fraction as S043's run on the same shape. Cost is dominated by output here, not input, which is why the band is back at S043's level after S047's structural 47%. |
| 2026-07-28 | S049 | the whole translation limb, the source collation, both contamination gates, and every analysis | 0.00 | 0.000000 | — | No API call. 蒲松齡〈王六郎〉 rendered whole from classical Chinese (1,533 characters, 31-point frozen log, R06 draft frozen separately first), the two-witness source collation with its two emendations, the paragraph-1 selection gate and the whole-work contamination measurement, build_items.py, analyse.py, the 143-check independent verify.py, and s043_repeat_check.py — which recomputed S043's own repeat condition from S043's stored bodies and found the figure that closed the arm — are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
S051 spend note — three calls, $0.0893677125, nothing wasted, and the cross-check is exact for the second session running. Opening snapshot 24.953626791, closing 25.042994503, delta 0.089367712 against a per-request sum of 0.089367713 — a residual of −1e-9, the same agreement S050 demonstrated on a single call and the first time this ledger has had it across three calls to three different labs. Note (b) did not fire; no fall-through was entered; all three returned finish_reason: stop on the first attempt. Note (x) held on all three providers — Google, OpenAI, xAI, list price each. The worst-case fractions were 30% (critic) and 23% (readers), inside the 15–34% band every output-dominated run in this ledger has landed in, and the one call whose value is hardest to price bought the withdrawal of the design's own headline reading before a number existed.
| 2026-07-29 | S054 | E-20260729b-graded-drift independent pre-run critic — google/gemini-3.6-flash (P2), succeeded first call | 0.075 worst case from max_tokens 8000 at list out-price plus the prompt at list in-price (note (abc)) | 0.0405165 | per-response usage.cost; opening/closing key snapshots persisted to runs/snap/ (note (bco)) | in 8,601 / out 3,682, stop, 20.5s, provider Google AI Studio. 54% of worst case. P2 because it is a subject in nothing here — the declared fix for S053's role collision. Verdict NEEDS-REDESIGN, six findings, all six accepted, and the critical one replaced the crux test: a likelihood-ratio test against a noisy class label cannot distinguish "the boundary is imposed" from "the boundary is real and the label is noisy", so the crux became a direct test of latent structure. Note (rr), twelfth consecutive session. |
| 2026-07-29 | S054 | E-20260729b graded rating — 6 calls (3 raters × 2 axes, separate calls per axis after critic finding 5), 132 items each | 0.665 worst case from max_tokens 12000 across six at each model's out-price plus prompts at list in-price | 0.258138 | per-response usage.cost, summed in runs/cost.json, cross-checked against the call count by verify.py | 792 of 792 scores returned, every value an integer 0–100, every row carrying a reason. P1 openai/gpt-5.6-terra (OpenAI, list honoured — note (x), fourteenth run) $0.037869 + $0.028430; P3 x-ai/grok-4.5 (xAI, list honoured) $0.029682 + $0.026862; P5 deepseek/deepseek-v4-pro $0.049828 wasted on an empty body at Cloudflare, then $0.068283 on the re-run at Together, plus $0.017012 at GMICloud. 39% of worst case. |
| 2026-07-29 | S054 | E-20260729b third categorical reader — deepseek/deepseek-v4-pro, two attempts, no body | 0.010 worst case from max_tokens 4000 | 0.021288 | per-response usage.cost; the second attempt had billed nothing at the closing snapshot | Attempt 1 (DigitalOcean): finish_reason: length, empty content, 6,000 output tokens, $0.021288 for nothing. Attempt 2, re-run at max_tokens 32,000 per F2: dispatched, never returned, nothing billed — the S053 hanging-call shape, recorded again. Over its worst case, because the worst case priced one call and the cap was raised on the retry. The reader is dropped as registered; the four-class α rests on S053's two readers and the primary comparison was computed rater-count-matched, so the drop costs a secondary figure and not the result. Notes (b) — eighth session — and new (bdl). |
| 2026-07-29 | S054 | the whole translation limb, the source collation, the contamination gate, and every analysis | 0.00 | 0.000000 | — | No API call. Alfred's Preface to the Pastoral Care rendered whole from Old English (874 words, R06 draft frozen separately first, 80-site reflex census and the lead's own graded scores frozen before translating), the two-witness Bright/Sweet collation with its three emendations, the contamination gate via tools/dependence_check.py (20-token run), analyse.py with a from-scratch Hartigan dip test and Gaussian-mixture EM, and the 77-check independent verify.py and its mutation test are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-29 | S055 | E-20260729c-neutral-summary independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call | 0.29 worst case from max_tokens 16000 at list out-price plus the prompt at list in-price (note (abc)) | 0.067524 | per-response usage.cost; opening/closing key snapshots persisted to runs/snap/ (note (bco)) | in 8,938 / out 2,714, stop, 132.8s, provider Moonshot AI. 23% of worst case. P4 because it is a subject in nothing here — P1 and P2 are the two voices being re-run, P3 and P5 are the summarisers. Its one prior failure in this project was at max_tokens 6,000 (note (b), S044, $0.130665 wasted); note (bdl) says the remedy that worked was raising the cap, so it ran at 16,000 and returned first time. Verdict NEEDS-AMENDMENT, five findings, all five accepted, and the critical one withdrew a sentence the design had written about itself an hour earlier — "a one-cell question and one cell answers it" — before any call was dispatched. Note (rr), thirteenth consecutive session. |
| 2026-07-29 | S055 | E-20260729c the two non-lead summaries — byte-identical prompts, two labs, whole re-derived primary in, no options shown | 0.15 worst case across both | 0.040235840 | per-response usage.cost | P3 x-ai/grok-4.5 (xAI, list honoured) $0.0263924, in 7,527 / out 1,926; P5 deepseek/deepseek-v4-pro (GMICloud) $0.01384344, in 7,492 / out 6,454. Both stop first time, and note (b) did not fire on P5 — at max_tokens 32,000 per note (bdl), the ninth session of watching this model. 27% of worst case. The second summary exists because the pre-run critic's F1 said one summary cannot separate summary authorship from vote variance. |
| 2026-07-29 | S055 | E-20260729c stage 2 — the ratification protocol, run across three evidence blocks (2 review calls, 3 vote calls) | 0.52 worst case from max_tokens 10,000 ×2 and 8,000 ×3 | 0.088146 | per-response usage.cost, summed in analysis/results.json, recomputed from the .raw bodies by verify.py | P1 openai/gpt-5.6-terra (OpenAI, list honoured — note (x), fifteenth run running) $0.01680375 + $0.0180278125; P2 google/gemini-3.6-flash (Google AI Studio ×2, Google ×1) $0.015582 + $0.0180195 + $0.019713. 5 of 5 stop, every one returning a parseable option letter on its first line. 17% of worst case. |
| 2026-07-29 | S055 | E-20260729c Part B — the two independent readings of the translation, both blind to the predictions | 0.20 worst case | 0.064439688 | per-response usage.cost | Asymmetry recovery, English only, no Russian and no census: google/gemini-3.6-flash (Google) $0.024975. Device classification over the 31 frozen sites: openai/gpt-5.6-terra (OpenAI) $0.0394646875, in ~6,400 / out substantial, all 31 of 31 sites classified exactly once. 32% of worst case. These two calls exist because the pre-run critic's F3 refused to let the translator classify its own devices. |
| 2026-07-29 | S055 | the whole translation limb, the primary re-derivation, both contamination gates, the frozen gates and every analysis | 0.00 | 0.000000 | — | No API call. Saltykov-Shchedrin's «Повесть о том, как один мужик двух генералов прокормил» rendered whole from the Russian (2,059 → 2,844 words), the 31-site address census frozen and committed before a word was written, the re-derivation of The Nation pp. 93–94 from the scan geometry via tools/deinterleave_djvu.py, the unit-A selection gate and the whole-work contamination measurement via tools/dependence_check.py, analysis/gates.py, analyse.py and the 50-check independent verify.py are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-29 | S056 | E-20260729d-decision-grain independent pre-run critic — google/gemini-3.6-flash (P2), succeeded first call | 0.119 worst case from max_tokens 14,000 at list out-price plus the prompt at list in-price (note (abc)) | 0.0334875 | per-response usage.cost; opening/closing key snapshots persisted to runs/snap/ (note (bco)) | in 9,535 / out 2,558, stop, 15s, provider Google AI Studio. 28% of worst case. P2 because P1 and P3 are the readers and cannot critique the instrument they are about to be. Verdict NEEDS-REDESIGN, five findings, four accepted and one declined in writing, eight amendments before dispatch. The decisive one added a positive control the design did not have: the session rested entirely on reading a zero and had no evidence the DECIDES label could be returned at all. Note (rr), fourteenth consecutive session, and its best return yet — the control fired and made both RS-20260728i's zero and this session's readable. |
| 2026-07-29 | S056 | E-20260729d classification pass — 2 calls, 63 pre-existing decisions × 17 candidates | 0.295 worst case from max_tokens 12,000 ×2 | 0.0708899 | per-response usage.cost, recomputed from the .raw bodies by verify.py | P1 openai/gpt-5.6-terra (OpenAI, list honoured — note (x), sixteenth run running) $0.0356975, in 9,599 / out 2,760; P3 x-ai/grok-4.5 (xAI, list honoured) $0.0351924, in 9,581 / out 2,708. Both stop first time, 63 of 63 lines each. 24% of worst case. The pass's registered reliability floor F1 fired (κ 0.066), so its counts are not reportable as estimates — the money bought a null about the instrument rather than a number about the framework, and that was a registered possible outcome. |
| 2026-07-29 | S056 | E-20260729d applicability pass — 4 calls (2 readers × 2 rules, separate calls so neither reader sees both) | 0.41 worst case from max_tokens 9,000 ×4 | 0.08617011 | per-response usage.cost, summed in analysis/results.json, recomputed from the raw bodies by verify.py | P1 $0.02870625 + $0.0195090625; P3 $0.0104464 + $0.0275084. 4 of 4 stop, 23 of 23 sites each, zero NORULE, zero UNSURE. 21% of worst case. This is the pass that carried the session's result. |
| 2026-07-29 | S056 | the whole translation limb, the census, the contamination gate and every analysis | 0.00 | 0.000000 | — | No API call. Gogol's «Ночь перед Рождеством» ¶¶93–104 rendered whole from the Russian (962 → 1,295 words), the 23-type / 50-occurrence realia census frozen and committed before a word was written, the contamination measurement via tools/dependence_check.py (9-token run, clean), and call.py, build_materials.py, build_prescription.py, analyse.py and the 61-check independent verify.py are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-29 | S057 | E-20260729e-revision-pass independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call | 0.29 worst case from max_tokens 16,000 at list out-price plus the prompt at list in-price (note (abc)) | 0.101034 | per-response usage.cost; opening snapshot persisted to runs/snap-open.json (note (bco)) | in 4,287 / out 3,633, stop, 43s, provider Fireworks. 35% of worst case. P4 because P1/P3/P5 are the readers and a reader may not critique the instrument it is about to be. Verdict NEEDS-AMENDMENT, seven findings — four BLOCKING, three MANDATORY — all seven accepted, one sub-clause of finding 5 declined in writing on a measurement. Four of the five registered predictions were changed by it before dispatch: one withdrawn as already-known, one withdrawn for want of a denominator, one made one-directional, and the design's central matching rule rebuilt. Note (rr), fifteenth consecutive session, and its most expensive-to-ignore return yet: the control it added is the gate that ended up withholding the session's headline. |
| 2026-07-29 | S057 | E-20260729e reader pass — 3 readers × 2 graded axes, 221 items each (205 edits + 16 undisclosed controls) | 0.83 worst case from max_tokens 12,000 ×6 | 0.2566791742 | per-response usage.cost over all seven bodies including the rejected one, summed in runs/reader-costs.json, recomputed from the .raw bodies by verify.py | 6 of 6 accepted, 221 of 221 parsed on every one. P1 openai/gpt-5.6-terra (OpenAI, list honoured — note (x), seventeenth run running) $0.0319775 + $0.0317321875; P3 x-ai/grok-4.5 (xAI) $0.0717884 + $0.0367064; P5 deepseek/deepseek-v4-pro failed on the M axis at StreamLake — 12,000 output tokens, finish_reason: length, empty body, $0.0260222655 wasted — and the declared reserve P2 google/gemini-3.6-flash served it first time at $0.0348765; P5's E-axis call returned cleanly at GMICloud, $0.0235759212. 31% of worst case. Note (b), tenth session. Consequence declared rather than buried: the M axis's third reader is P2 and the E axis's third reader is P5, so α_M and α_E are not computed over the same panel. |
| 2026-07-29 | S057 | the whole translation limb, the contamination gate, the extraction, the diff, the log matching and every analysis | 0.00 | 0.000000 | — | No API call. Machado de Assis's «O enfermeiro» ¶1–15 rendered from the Portuguese (758 → 955 words, the project's first Portuguese), the R06 draft frozen in two commits with the contamination gate run on Unit A before Unit B was drafted, the gate itself via tools/dependence_check.py, and extract.py, build_edits.py, match_logs.py, build_items.py, analyse.py and the 72-check independent verify.py are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-29 | S058 | E-20260729f-inheritance-census independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call | 0.27 worst case from max_tokens 16,000 at list out-price plus the prompt at list in-price (note (abc)) | 0.0875115 | per-response usage.cost; opening/closing key snapshots persisted to runs/snap-*.json (note (bco)) | in 5,852 / out 2,719, stop, 37s, provider Fireworks. 32% of worst case. P4 because P1 and P3 are the censusers and scorers, and a reader may not critique the instrument it is about to be — the declared fix for S053's role collision, in which the critic's own model produced the alternative census. Verdict NEEDS-AMENDMENT, eleven findings — five BLOCKING, four MANDATORY, two ADVISORY — all eleven accepted, one sub-clause of finding 8 corrected on a fact. Note (rr), sixteenth consecutive session. Two of its amendments decided what the session could say: it forced the aligned Yosano span to be located and frozen before scoring (which fired and found nothing), and it replaced a two-way fork whose unnamed third outcome was the one that exonerates the anchor. |
| 2026-07-29 | S058 | E-20260729f census pass — 4 calls, 2 models × 2 passages, temperature 0 | 0.36 worst case from max_tokens 8,000 ×4 | 0.0538302375 | per-response usage.cost, summed in runs/census-cost.json, recomputed from the .raw bodies by verify.py | P1 openai/gpt-5.6-terra (OpenAI ×2, list honoured — note (x), eighteenth run running) $0.014715 + $0.0133684375; P3 x-ai/grok-4.5 (xAI ×2) $0.0115044 + $0.0142424. 4 of 4 stop first time, no reserve fired. 15% of worst case. The registered enumeration-compliance check inside these calls is what failed on the Malory passage (0.667 on a determinate list) and withheld that arm under F1 — note (bdy). |
| 2026-07-29 | S058 | E-20260729f scoring pass — 2 calls, 26-item union, scored against the frozen aligned span | 0.33 worst case from max_tokens 12,000 ×2 | 0.0411719 | per-response usage.cost, summed in runs/score-cost.json, recomputed from the .raw bodies by verify.py | P1 (Azure this time, list honoured) $0.0270775; P3 (xAI) $0.0140944. 26 of 26 items returned by both, 52 of 52 attesting quotations verified by exact string match, 0 cells discarded under F2. 12% of worst case. This is the pass that carried the session's result. |
| 2026-07-29 | S058 | the whole translation limb, the aligned span, the contamination gate, Arm A and every analysis | 0.00 | 0.000000 | — | No API call. Malory Le Morte Darthur XVIII.xxiv rendered whole from the Middle English (821 → 851 words, the project's first intralingual translation into English), the 20-item census frozen and committed before a word was drafted, the R06 draft frozen separately first, the contamination measurement via tools/dependence_check.py, the mechanical attestation of all 34 bracketed quotations on the anchor, and build_materials.py, build_items.py, analyse.py and the 68-check independent verify.py and its mutation test are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
2026-07-29 day total: $2.256499882 of $5.00 (S053 $0.319738262 + S054 $0.319771867 + S055 $0.260345590 + S056 $0.190547512 + S057 $0.357713174 + S058 $0.182513637 + S059 $0.563592840 + S060 $0.062277000 — eight sessions in this UTC day). $2.743500118 headroom remains. S060 is the cheapest non-zero session since S052: one call, and the whole of the rest of the unit — a six-poem translation limb, the re-fetch and re-alignment of two Gutenberg volumes, the restatement of every published rank on three pages, and a 1,738-check verifier — cost nothing.
| date | session | what | pre-flight (USD) | actual (USD) | how measured | notes |
|---|---|---|---|---|---|---|
| 2026-07-29 | S060 | E-20260729h-baseline-restate independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call |
0.2565 worst case from max_tokens 16,000 at list out-price plus the prompt at list in-price (note (abc)); 0.385 with the declared P2 reserve |
0.062277 | per-response usage.cost; opening and closing key snapshots persisted to runs/snap-open.json / snap-final.json (note (bco)) |
in 5,159 / out 3,120, stop, 114s, provider Moonshot AI. 24% of worst case — the lowest fraction since S055's 23%, and P4 returned short for once. Role choice declared and it is a first for this ledger: no model is a subject in this design, every measurement being arithmetic over stored texts plus the lead's own translation, so the role-collision rule that picked the critic in S053–S059 did not bind and P4 was chosen for depth instead. Verdict NEEDS-AMENDMENT, six findings — three BLOCKING, three MANDATORY, all six accepted, one on the first of its two offered fixes with the second declined in writing. Note (rr), eighteenth consecutive session, and its most arithmetical return: two of nine registered predictions could not fail and the central one was confounded with the selection rule. Note (b) did not fire, third session running; no reserve was entered. |
| 2026-07-29 | S060 | the whole translation limb, the comparator re-fetch and re-alignment, every rank restatement and the verifier | 0.00 | 0.000000 | — | No API call. Six Senilia prose poems rendered from the Russian (898 → 1,202 words, R06 draft frozen in its own commit before the revision, two-witness source collation identical on all six), the re-run of S027's pipeline behind a reproduction gate, the completion of the 42-unit dependence split, the clean-subset restatement of every published rank on three pages, and extract_units.py, analyse.py and the 1,738-check independent verify.py with its mutation test are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
S060 cross-check — exact, and note (bco) did NOT fire. Opening snapshot 28.089972505, closing 28.152249505, delta 0.062277000 against a per-request sum of 0.062277 — agreement to the digit. Worth one line because S059 recorded this note at its largest (0.145 short): the delta read 0.000000000 immediately after the call and settled exactly by the end of the session, so (bco) is a settling lag and not a discrepancy, and a snapshot taken at hand-off rather than at dispatch is the remedy. verify.py asserts the sum, not the delta, which remains the right choice.
S059 spend note — twenty calls, twenty bodies, nothing wasted, and the key-usage delta is the worst settling lag this ledger has recorded. Rating-stage delta 0.259874787 against a per-request sum over the same eighteen bodies of 0.4052433375 — a shortfall of 0.145, note (bco) at its largest, and the reason is visible: the closing snapshot was taken seconds after the last of eighteen calls returned. verify.py asserts the delta as a bound (positive, not exceeding the sum) rather than as an approximate equality, and prints the shortfall, which is the honest form of a check that cannot be an equation. Three other things are worth keeping. (i) Note (b) did not fire on any of the four labs used, the second session running, and P5 returned cleanly at max_tokens 16,000 from a seventh provider (Baidu) — (bdl)'s raise-the-cap remedy holding. (ii) Worst-case fractions were 23%, 23% and 49%, and the outlier is the critic line for the second time in three sessions, always for the same reason: P4 answers at length. The band for output-dominated rating work is unmoved. (iii) The session's most expensive discretionary choice was doubling the rating pass to replicate every condition, +$0.20, and it bought the correction to ARM-sense-boundary's closing sentence.
S058 cross-check: key-usage delta 0.182513637 against a per-call sum of 0.1825136375 — residual 5e-10, the tightest agreement in this ledger. Opening snapshot 27.324518129, closing 27.507031766, both persisted. Seven dispatched, seven accepted first time, no reserve fired, no empty body — note (b) did not fire, the first session in three. Worst-case fractions 12–32%, the lowest band since S050.
S057 spend note — eight calls, seven bodies, one wasted, and the key delta settles exactly one call behind. Opening snapshot at the critic 26.960834756. Reader-stage delta 0.233103252 against a per-request sum over all seven reader-stage bodies of 0.2566791742 — and the difference is the final call to a residual of 1e-9, which is note (bco)'s settling lag caught cleanly rather than mistaken for a discrepancy; verify.py asserts it as an equation rather than as an approximation. Three things are worth keeping. (i) Note (b) fired for the tenth session and the declared reserve was the remedy — which is the S044/S045/S049 answer, not S054's raise-the-cap answer (note bdl); at max_tokens 12,000 StreamLake spent the whole budget and returned nothing, and P2 answered the identical prompt in 7 seconds. One slug is not one instrument and the provider decided the outcome again. (ii) The wasted $0.0260222655 bought something anyway, because falling through changed the panel composition mid-run and that is now a declared limit on the result rather than an invisible one. (iii) Worst-case fractions were 35% and 31%, at the top of the 15–34% band this ledger has always landed in and just outside it on the critic — the first time the critic line has exceeded the band, and the reason is that P4 returned 3,633 output tokens where a short structured verdict was expected.
S056 spend note — seven calls, seven bodies, nothing wasted, and the cross-check is exact for the fourth session running. Opening snapshot 26.427677644 — identical to S055's closing figure to the digit, which is the second time this ledger has had that — closing 26.618225156, delta 0.190547512 against a per-request sum of 0.1905475125, a residual of −5e-10. Three things are worth keeping. (i) The three intra-session snapshots each read a delta of exactly 0.0 and only the session-level pair settled, which is note (bco) demonstrated three times inside one session: a snapshot taken at the instant a call returns is worthless, and taking more of them does not help. (ii) Note (b) did not fire on any of the five labs used, and no fall-through was entered. (iii) The design's own worst-case figure was wrong and the spend was not. §8 of the frozen design said ≈$0.50; rebuilt from the max_tokens actually sent it is ≈$0.82, and the actual $0.190547512 is 23% of the corrected figure — inside the 15–34% band every output-dominated run here has landed in. Note (abc) fired on the estimate rather than on the spend, which is a new place for it to fire and is recorded as one.
S055 spend note — ten calls, ten bodies, nothing wasted, and the cross-check is exact for the third session running. Opening snapshot 26.167332055, closing 26.427677644, delta 0.260345589 against a per-request sum of 0.260345590 — a residual of −1e-9, after S050 (one call), S051 (three calls) and now ten calls across five labs. Three things are worth keeping. (i) Note (b) did not fire, and the model it usually fires on returned cleanly: deepseek/deepseek-v4-pro at max_tokens 32,000 produced 6,454 tokens at GMICloud on the first attempt, which is note (bdl)'s remedy — raise the cap, do not fall through — confirmed on the session after it was written. (ii) moonshotai/kimi-k3 returned cleanly at 16,000 after failing at 6,000 in S044, so the same remedy holds on the other lab this note was born on; P4 has now been usable twice in a row when sized properly, and it was the cheapest useful thing in the session at 26% of the spend for five accepted findings. (iii) The worst-case fractions were 17–32%, inside the 15–34% band every output-dominated run in this ledger has landed in, with no structural outlier — the first session in four with nothing to explain about its own estimate.
UTC day 2026-07-30
| date | session | what | pre-flight (USD) | actual (USD) | how measured | notes |
|---|---|---|---|---|---|---|
| 2026-07-30 | S061 | E-20260730-grain-clause independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call |
0.45 worst case for the chain from max_tokens 16,000 at list out-price plus the prompt at list in-price (note (abc)), P2 reserve included |
0.139212 | per-response usage.cost; opening key snapshot persisted to runs/snap/snap-open.json (note (bco)) |
in 28,259 / out 3,629, stop, 113s, provider Moonshot AI. 31% of the chain worst case. P4 because P1, P3 and P5 are the readers and a reader may not critique the instrument it is about to be — the S053 role-collision fix, ninth session running. Verdict NEEDS-AMENDMENT, eleven findings — one BLOCKING, seven MANDATORY, three ADVISORY, all eleven accepted, one narrowed in writing. Note (rr), nineteenth consecutive session, and its second-largest return: it struck a failure criterion that could not fire, a prediction that could not fail, and rebuilt the arm's own decisive measurement — the warrant statistic was a Jaccard between rules that share most of their text. |
| 2026-07-30 | S061 | E-20260730-grain-clause applicability passes — 18 accepted calls across three arms (23 Gogol sites ×4 conditions, 17 anchor sites ×3, 34 Japanese sites ×2), P1 and P3 |
1.66 worst case from max_tokens 9,000 / 6,000 per call |
0.4125394 | per-response usage.cost, summed in analysis/results.json, recomputed from the .raw bytes by verify.py |
P1 openai/gpt-5.6-terra (OpenAI throughout, list honoured — note (x)) and P3 x-ai/grok-4.5 (xAI throughout). 18 of 18 accepted first time; 442 site lines, zero NORULE, zero UNSURE, zero test-to-handling incoherence. 25% of worst case. This is the pass that carried the session's result, and the two conditions that carried the headline are byte-identical to S056's, asserted at 16,154 bytes. |
| 2026-07-30 | S061 | E-20260730-grain-clause the third reader, abandoned — P5 deepseek/deepseek-v4-pro, then the declared reserve P2 |
0.11 | 0.0146570814 wasted | per-response usage.cost on the rejected body; recorded in results.json as a wasted call and asserted by verify.py |
Note (b), eleventh session, and note (beh) is new. P5 held sockets open past seven minutes without returning on two launches and was abandoned both times — a fall-through chain protects against a call that returns a failure and does nothing against one that hangs. On the third launch it returned finish_reason: length, 9,000 output tokens, empty body, at GMICloud. The reserve P2 was dispatched, did not return, and was abandoned. The dispatch order was changed mid-session so all 18 primary calls completed before the optional seat was attempted, which is why the primary is intact. RS-20260729d revision trigger 1 remains open. |
| 2026-07-30 | S061 | the whole translation limb, the census, the contamination gate and every analysis | 0.00 | 0.000000 | — | No API call. Akutagawa 「煙管」 sections 一 and 二 rendered from the Japanese (1,732 characters → 1,015 words), the 34-site census frozen and committed before a word was written, the R06 draft frozen in its own commit, the contamination measurement via tools/dependence_check.py (13-token run against Shaw 1930 — suspected, and a finding), and build_census.py, build_prompts.py, build_anchor.py, analyse.py and the 170-check independent verify.py with a mutation test are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-30 | S062 | E-20260730b-c16-redraw the gate S061 named as its own defect — the frozen C16 prompt re-sent verbatim, P1 + P3 | 0.134 worst case from max_tokens 9,000 at list out-price plus the prompt at list in-price (note (abc)) | 0.0362414625 | per-response usage.cost; opening/closing key snapshots persisted to runs/snap/key-usage.json (note (bco)) | 27.0% of worst case, inside the 15–34% band. P1 in 3,352 / out 2,879 at OpenAI; P3 in 3,546 / out 423 at xAI — prompt token counts identical to the digit on both days, same providers. Both accepted first call, no reserve fired, note (b) did not fire. Two calls that turned an unbounded retraction into a bounded one: outcomes B and C both fire and nothing of RS-20260729d §1's trade survives. |
| 2026-07-30 | S062 | E-20260730c-revision-close independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call | 0.270 worst case from max_tokens 16,000 | 0.0876726 | per-response usage.cost; opening snapshot persisted | in 5,398 / out 4,771, stop, 122 s, provider Together. 32% of worst case. P4 because P1/P3/P5 are the readers and P2 is their declared reserve — the S053 role-collision fix, tenth session running. Verdict NEEDS-REDESIGN, the first this project has taken since S021: two BLOCKING, four MANDATORY, one ADVISORY, all seven accepted, none declined. Note (rr), twentieth consecutive session, and its most expensive return — it withdrew eight already-frozen control items unscored and converted an attribution gate that could not fail into a comparison that can. |
| 2026-07-30 | S062 | E-20260730c-revision-close blind control builder — qwen/qwen3.7-max, a call added by amendment A2 and not in the frozen estimate | (not estimated; the amendment created it) | 0.0406038 | per-response usage.cost | in 822 / out 8,902, stop, 166 s, provider Alibaba at max_tokens 4,000 — the output figure exceeds the cap because it is dominated by hidden reasoning tokens. A materials-construction role, not a judging role: not a reader, not the readers' reserve, not the critic. It matched the corpus's span statistics to within 0.083 of a word without being shown a single score. Notes (ben), new, and (abc) in a new place. |
| 2026-07-30 | S062 | E-20260730c-revision-close the reader pass — 6 calls, 3 seats × 2 axes, P1 / P3 / P5 | 0.552 worst case (+0.226 if two reserves fire) from max_tokens 12,000, P5 priced at the worst plausible provider per the S022 routing caution | 0.2400870038 | per-response usage.cost, summed in runs/reader-costs.json, recomputed from the .raw bytes by analysis/verify.py | 43% of the readers-only worst case, 31% including the reserve allowance. P1 OpenAI ×2 ($0.0672); P3 xAI ×2 ($0.1002); P5 StreamLake then SiliconFlow ($0.0727). All six accepted first time; no reserve fired; note (b) did not fire — and S057's firing on this same batch is what the analysis had to correct for (note (beo)). Key-usage delta 0.232554393 against a per-request sum of 0.2400870038, a 0.0075 shortfall asserted as a bound. |
| 2026-07-30 | S062 | the Polish translation, both logs, seven contamination measurements, the item rebuild, item (b) entire, every analysis and two verifiers | 0.00 | 0.000000 | — | No API call. Sienkiewicz «Niewola tatarska» I opening rendered from the Polish (985 → 1,329 words), the R06 draft and its 30-point log frozen at 14c8989, the R04 revision and its 21-item pass-separated log frozen at 73ba25c, contamination appended at 1bf2831, and drift.py's six-pair measurement plus the 93-check and 167-check verifiers are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-30 | S063 | E-20260730d-register-centre blind feature enumeration — 7 accepted calls, qwen/qwen3.7-max, a materials-construction seat | 0.303 worst case from max_tokens 3,000 priced at 3× the cap (the S062-observed reasoning-token ratio), against a cap-literal $0.118 — both stated, the larger declared (note (abc)) | 0.14302 | per-response usage.cost summed from the stored bodies by analyse.py and re-summed independently by analysis/verify.py | in ~2,150–2,320 / out 3,310–4,649 per call, stop, provider Alibaba on all seven. 47% of the declared worst case. One transport failure and no reserve was used: on passage D the slug returned 1,892 bytes of SSE keep-alive padding and no JSON payload, which is not note (b)'s failure mode — nothing was served — and the seat could not be replaced without putting enumerator identity inside the study's own statistic, so it was re-issued to the same slug and accepted first time. No billing for the failed call is visible in the delta. New note (beq). |
| 2026-07-30 | S063 | E-20260730d-register-centre independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call | 0.261 worst case from max_tokens 16,000 | 0.1188702 | per-response usage.cost; opening snapshot persisted to runs/snap/critic-open.json (note (bco)) | in 5,351 / out 6,866, stop, 191 s, provider Together. 46% of worst case. P4 because P1/P3 are the raters, P2 is their declared reserve and qwen is the materials seat — the S053 role-collision fix, eleventh session running. Verdict NEEDS-REDESIGN, the second consecutive session to take one: two BLOCKING, four MANDATORY, all six accepted, one narrowed with a stated reason. Note (rr), twenty-first consecutive session, and its most valuable return yet — Finding 2 established that the design's frozen primary statistic was an artifact of how the item pool was built, and it was withdrawn and replaced before the measurement ran rather than after. |
| 2026-07-30 | S063 | E-20260730d-register-centre the cross-census — 14 cells, 7 passages × 2 raters, P1 and P3 | 0.647 worst case from max_tokens 6,000, rebuilt from the actual prompt sizes after the pool shrank 56 → 38 classes | 0.2239167 | per-response usage.cost, summed in runs/census-costs.json, recomputed from the .raw bytes by analysis/verify.py | P1 openai/gpt-5.6-terra ×7 ($0.1236 — six at OpenAI, one at Azure, note (x)); P3 x-ai/grok-4.5 ×7 ($0.1004, xAI throughout). All 14 accepted first time, no reserve fired, note (b) did not fire. 266 cells parsed complete, UNCLEAR available on every one and used zero times — note (ber). |
| 2026-07-30 | S063 | E-20260730d-register-centre the repeat control — 4 byte-identical cells, P1 | 0.198 worst case from max_tokens 6,000 | 0.0669630 | per-response usage.cost; analysis/verify.py asserts the prompt token counts identical to the digit | P1 at OpenAI ×4, all accepted first time. Prompt token counts identical to the digit on all four (2,770 / 2,735 / 2,763 / 2,781), which is what makes "byte-identical" a checked claim rather than an assertion. 11 of 152 PRESENT/ABSENT judgements flipped, 7.2%, and the registered decision held under every substitution. |
| 2026-07-30 | S063 | the whole translation limb, the contamination gate, all materials construction and every analysis | 0.00 | 0.000000 | — | No API call. Tagore «পোস্ট্মাস্টার» rendered complete from the Bengali (1,652 tokens → 2,466 words) under R07, the log frozen at 001405b before the design document existed, the contamination gate run at 73e7307 (11-token run, 0 twelve-grams), and build_passages.py, build_pool.py, analyse.py, analysis/sensitivity.py and the 207-check analysis/verify.py with four mutation tests are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
2026-07-30 day total: $1.5237830577 of $5.00 — three sessions (S061 $0.5664084814, S062 $0.4046048663, S063 $0.5527697100). $3.4762169423 headroom. S063: 18 billed calls, 18 accepted first time, nothing wasted, no reserve fired, note (b) did not fire.
S063 cross-check — exact at session level, and the per-stage check had to be rewritten because it failed in a direction this ledger has not seen. Session opening 29.516047048, closing 30.068816756, delta 0.552769708 against a per-request sum of 0.5527697100 — residual −2e-9 — an exact reconciliation, alongside S055's −1e-9 and S056's −5e-10, after two sessions (S061, S062) that could only be asserted as bounds. But the repeat stage's delta (0.10137) EXCEEDS its own per-request sum (0.06696), because the census stage's outstanding billing settled inside the repeat window: the census delta (0.12254) was short of its sum (0.22392) by 0.10138, which is the repeat window's entire delta to five decimal places. Note (bco) has only ever been seen in the shortfall direction; a per-stage upper bound on the delta is therefore not a valid check and the verifier now asserts monotonicity per stage and the bound at session level. New note (bet). Intra-session snapshots persisted at runs/snap/: enum-open, critic-open, census-open/close, repeat-open/close, session-close.
And the pre-flight fraction landed outside the band for a stated structural reason. Declared worst case after amendment $1.261; actual $0.5528 = 43.8%, against the 15–34% band every output-dominated run in this ledger has occupied. The reason is that this run is prompt-dominated, not output-dominated: the census calls sent ~2,800 input tokens and returned 200–2,300 against a 6,000 cap, so the cap — which is what note (abc) requires the worst case to be built from — over-states an output the design never wanted. That is the same structural exception (abc) recorded at S047 at 47%, and the band should be read as applying to output-dominated runs only. Note (abc) fired here on the enumeration row as well, in the form it was written for: the cap-literal figure ($0.118) and the reasoning-inflated one ($0.303) are both stated and the larger is the one declared.
(S062's day-total line, superseded by the S063 line above and kept for the audit trail:) 2026-07-30 day total $0.9710133477 of $5.00 — two sessions (S061 $0.5664084814, S062 $0.4046048663), $4.0289866523 headroom.
S062 cross-check — two settling lags, both asserted as bounds rather than equations. Gate: opening 29.107243184, closing 29.111442184, delta 0.004199 against a per-request sum of 0.0362414625. Readers: opening 29.235356246, closing 29.467910639, delta 0.232554393 against a sum of 0.2400870038 — a 0.0075 shortfall, smaller than S061's 0.021 and S059's 0.145. Intra-session snapshots at the critic (29.147683646) and the builder (29.235356246) are persisted; the builder's opening figure equals the readers' opening figure to the digit, which is note (bco) demonstrated again — a snapshot taken at the instant a call returns is worthless.
And note (abc) fired in a place it has not fired before: on the pre-flight fraction, not the estimate. E-20260730c's frozen worst case was $1.05 and its actual is $0.3683634038 — 35.1%, just outside the 15–34% band this ledger has always landed in. The reason is that amendment A2 added a call the estimate could not contain. Excluding the added call the run is at 31.2% and inside the band, and both numbers are stated rather than the flattering one. The operative lesson: an amendment that adds a call invalidates the pre-flight fraction, and re-baselining it silently would hide the amendment.
S061 cross-check — the delta is short of the sum and the check is a bound, which is the right form. Opening snapshot at the critic 28.158531505; reader-stage opening 28.337017655, closing 28.743226484, delta 0.406208829 against a per-request sum over all 19 stored reader-stage bodies of 0.427196481. The shortfall is 0.021 — note (bco)'s settling lag, and smaller than S059's 0.145. verify.py asserts the delta as positive and not exceeding the sum including the wasted call, which is the honest form and which would also have caught a billed call with no stored body as an excess. No excess appeared, so the three abandoned sockets are not visible in the billing at hand-off — evidence, not proof, since the lag runs the other way.
And a figure in config/models.md was wrong for eight days. P1's list price reads $1.25 / $7.50 from GET /api/v1/models, half the $2.50 / $15.00 the panel table has carried since selection on 2026-07-23. Every pre-flight estimate built from that table has over-priced P1 by 2× — the conservative direction, which is why nobody noticed. Corrected in place; the revisit trigger "pricing shifts that break the cost structure" does not fire, since a halving on the frontier seat does not break the structure it was selected for. The lesson is the shape rather than the number: a price table is an input to every estimate this project makes and nothing was re-reading it.
| 2026-07-30 | S064 | E-20260730e-rule-coverage independent pre-run critic — moonshotai/kimi-k3 (P4), succeeded first call | 0.264 worst case from max_tokens 16,000 at list out-price plus the prompt at list in-price (note (abc)) | 0.05598 | per-response usage.cost; opening snapshot persisted to runs/snap/critic-open.json (note (bco)) | in 4,540 / out 2,824, stop, 52 s, provider Fireworks. 21% of worst case. P4 because P1/P3/P5 are the raters and P2 is their declared reserve — the S053 role-collision fix, twelfth session running. Verdict NEEDS-AMENDMENT, six findings — two BLOCKING, two MANDATORY, two ADVISORY, all six accepted, one discharged by withdrawal rather than amendment. Note (rr), twenty-second consecutive session. Its Finding 1 showed the C2 statistic counted convergence and called it silence — unanimous FLATTEN would have scored as the prediction failing — and §4 of the result page is readable only because that was repaired before dispatch. |
| 2026-07-30 | S064 | E-20260730e-rule-coverage the coverage-code pass (C1) and its byte-identical repeat (C1R) — 4 accepted dispatches, P1 / P3 / P2(reserve) / P1 | 0.285 worst case incl. the reserve allowance, from max_tokens 8,000 | 0.1103 | per-response usage.cost, summed in runs/rater-costs.json, recomputed from the .raw bytes by analysis/verify.py | P1 OpenAI ×2 ($0.0770, list honoured — note (x)); P3 xAI ($0.0081); P2 Google AI Studio ($0.0252) as the declared reserve. Prompt tokens 2,698 and 2,698 on C1-P1 and C1R-P1, identical to the digit — which is what makes "byte-identical" a checked claim. 9 of 37 codes flipped, 24.3%, self-agreement 0.757 against a pre-registered 0.80 floor: the control FAILED and the registered void condition is applied. Note (bev). |
| 2026-07-30 | S064 | E-20260730e-rule-coverage the F10-scope pass (C2) — 3 accepted calls, P1 / P3 / P5 | 0.074 worst case from max_tokens 4,000 | 0.0234 | per-response usage.cost | P1 OpenAI ($0.0157); P3 xAI ($0.0048); P5 StreamLake ($0.0029). All three accepted first call. 8 of 12 cells FLATTEN, 4 UNDECIDED, 0 KEEP — and the KEEP zero is worth nothing, because amendment A2 made the label near-unreachable (note (bex)). |
| 2026-07-30 | S064 | the wasted call — P5 deepseek/deepseek-v4-pro on C1 | (inside the reserve allowance above) | 0.012720357 wasted | per-response usage.cost on the rejected body; asserted by analysis/verify.py | Note (b), twelfth session. finish_reason: length, empty body at 8,000 tokens, GMICloud, 196 s. The declared reserve P2 took the seat and was accepted first call. Consequence for the result, stated rather than buried: C1's raters are P1/P3/P2 and C2's are P1/P3/P5, so the two conditions do not share a rater set. Note (bdt) is live rather than latent for the first time in this project — C1-P5.raw holds the rejected body and the verifier walks the attempt chain to reach C1-P5-reserve1.raw. |
| 2026-07-30 | S064 | the whole translation limb, the contamination gate, the item pool and every analysis | 0.00 | 0.000000 | — | No API call. Baudelaire's four prose poems rendered from the French (843 → 894 words) under R07, the coverage prediction frozen at e983b84 before the source was read, the log frozen at 1fc7a87, the contamination gate run at 4cfabe6 (16-token per-unit run; the whole-work 21 is a splice across two separate poems), and build_items.py, build_prompts.py, analyse.py and the 248-check analysis/verify.py with three mutation tests are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-30 | S065 | E-20260730f-recall-floor independent pre-run critic — P4 moonshotai/kimi-k3 wasted, P5 deepseek/deepseek-v4-pro (declared reserve) accepted | 0.264 worst case from max_tokens 16,000 at list out-price plus the prompt at list in-price (note (abc)) | 0.4382305347 | per-response usage.cost on both bodies; opening snapshot runs/snap/critic-open.json (note (bco)) | Note (b), THIRTEENTH session and the most expensive firing in this ledger. P4 in 12,211 / out 16,000, finish_reason: length, content: null, provider Fireworks, $0.4149495 for nothing — 41% of the session's whole spend and 157% of the stage's declared worst case on its own. The rejected body is preserved at critic.attempt0.raw and was not overwritten (the S045 runner defect, not repeated). Reserve P5 StreamLake, in 11,669 / out 11,542, stop, $0.0232810347, accepted first call. P1/P2/P3 are subjects and could not critique; P4 is the standing critic seat (S053 fix). Verdict NEEDS-AMENDMENT, seven findings — two BLOCKING, two MANDATORY, three ADVISORY, all seven accepted. Note (rr), twenty-third consecutive session. Its BLOCKING finding 2 (transliteration cues) would have decided the session in the non-conservative direction. |
| 2026-07-30 | S065 | E-20260730f-recall-floor the three choice conditions — G ×3 models ×2 orders, N ×3 ×2, P ×3 ×1 | 0.65 worst case from max_tokens 4,000 across 18 calls | 0.4080356875 (15 calls) | per-response usage.cost, runs/cost.json, recomputed from the stored bodies by analysis/verify.py | All 15 accepted first call, no truncation, no empty body. Providers: P1 OpenAI ×4 and Azure ×1; P2 Google / Google AI Studio; P3 xAI (the most expensive seat here — $0.1097724 on one Gb call, 5,843 completion tokens). P: 18/18. G: 0.467, p = 0.856. N: 0.042 — 46 of 48 judgments for the lead's own translation. |
| 2026-07-30 | S065 | E-20260730f-recall-floor the identification condition — 3 calls | (inside the 0.65 above) | 0.15424515 | per-response usage.cost | P1 OpenAI ($0.0206), P2 Google AI Studio ($0.0327, finish_reason: length — truncated, its figures are floors and nothing is inferred from it), P3 xAI ($0.1009, 16,179 completion tokens). Both models that answered named the author at 18 of 18 and the individual work at 7/10 Turgenev and 8/8 Korolenko. |
| 2026-07-30 | S065 | the whole translation limb, the span selection, the item build, the contamination measurements and every analysis | 0.00 | 0.000000 | — | No API call. 939 Russian words rendered blind from Turgenev and Korolenko and frozen at ff64960 before any comparator was opened; select_spans.py, finalise_spans.py, build_items.py, covariates.py, score.py, score_id.py and the 452-check analysis/verify.py are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
2026-07-30 day total: $2.7267064109 of $5.00 — five sessions (S061 $0.5664084814, S062 $0.4046048663, S063 $0.5527697100, S064 $0.2024119810, S065 $1.0005113722). $2.2732935891 headroom. S065: 20 dispatches, 19 accepted, 1 wasted to note (b), one declared reserve fired and was accepted first call.
S065 overran its own pre-flight, and the overrun is entirely the wasted call. Declared worst case $0.96; actual $1.0005113722 = 104% of it. Excluding the wasted $0.4149495 the session ran at $0.5856 — 61% of the declared figure and still above the 15–34% band, because like S063 and S064 this run is prompt-heavy relative to its output cap on the choice calls and reasoning-heavy on P3's, which the cap-literal worst case under-prices rather than over-prices. Both numbers are stated rather than the flattering one. Note (abc) fired on the estimate; note (b) fired on the call that broke it. The operative lesson is narrower than "the estimate was wrong": a stage whose declared worst case is $0.26 cannot absorb a single $0.41 failure, and the reserve mechanism protects the result but not the budget.
S065 cross-check — the delta is short of the sum, which is the ordinary direction, and it is asserted as a bound. Session opening (critic stage) 30.284178736, closing 31.183755706, delta 0.899576970 against a per-request sum of 1.0005113722 — a shortfall of 0.1009, note (bco)'s settling lag, and the same order as S063's within-session lags. Main-stage opening snapshot 30.72240927 is persisted at runs/snap/main-open.json; against the closing figure the main stage's delta is 0.461346436 against its per-request sum of 0.5622808375, shortfall 0.1009 to four decimals — i.e. the entire session shortfall is carried in the main stage and the critic stage settled fully, which is note (bet)'s per-stage monotonicity holding and its upper bound not being a valid check. Snapshots at runs/snap/.
(S064's day-total line, superseded by the S065 line above and kept for the audit trail:) 2026-07-30 day total $1.7261950387 of $5.00 — four sessions (S061 $0.5664084814, S062 $0.4046048663, S063 $0.5527697100, S064 $0.2024119810), $3.2738049613 headroom.
2026-07-30 day total: $1.7261950387 of $5.00 — four sessions (S061 $0.5664084814, S062 $0.4046048663, S063 $0.5527697100, S064 $0.2024119810). $3.2738049613 headroom. S064: 8 billed dispatches, 7 accepted, 1 wasted to note (b), one reserve fired. Declared worst case $1.16 (cap-literal $0.623); actual 17.4% of it.
(S063's day-total line, superseded by the S064 line above and kept for the audit trail:) 2026-07-30 day total $1.5237830577 of $5.00 — three sessions (S061 $0.5664084814, S062 $0.4046048663, S063 $0.5527697100), $3.4762169423 headroom.
| 2026-07-30 (S066 start) | 31.475270106 | S066 session start, before any project spend. $0.291514400 above S065's closing 31.183755706, of which S065's own $0.1009 declared unsettled at close accounts for 35%; residual $0.190614 unattributed, an order larger than the $0.005–0.009 inter-session drifts at S042/S049/S054. Not the S033 shape (a delta in a session that made no call) — S065 made twenty. Note (abf), eighth application, and the largest unattributed opening drift this ledger has recorded. |
| 2026-07-30 (S066 end) | 31.743668402 | S066 after ten dispatches. Delta from the opening snapshot 0.268398296 against a per-request sum of 0.2755366964, short by $0.0071384 — which is the final call's billed cost to the digit (D-rate-P3, $0.0071384). The first time in this ledger that note (bco)'s settling lag has been attributable to a named call rather than bounded. Both snapshots persisted to runs/snap/ before being read. |
| 2026-07-30 | S066 | E-20260730g-nonlead-decisions independent pre-run critic — P4 moonshotai/kimi-k3, accepted first call | 0.33 worst case from max_tokens 12,000 at list out-price plus the prompt at list in-price, times a 1.5 provider premium because the identical seat billed 1.5× list at S065 (notes (abc), (x), (b)) | 0.0934182 | per-response usage.cost; opening snapshot runs/snap/session-open.json (note (bco)) | in 11,617 / out 3,916, stop, provider Together — not Fireworks, which is where S065's $0.41 empty body came from. 28% of worst case. P1/P3/P5 are condition-A readers and P2 is a condition-D reader, so none could critique; P4 is the standing critic seat (S053 fix). Verdict NEEDS-AMENDMENT, nine findings — three BLOCKING, five MANDATORY, one ADVISORY, all nine accepted. Note (rr), twenty-fourth consecutive session. Its finding 7 showed both registered predictions would pass under any no-effort null and replaced the primary statistic with a permutation null; its finding 2 added the unforced second translation limb the session's actual result came from. |
| 2026-07-30 | S066 | E-20260730g condition A, the three classification readers — 5 dispatches, P1 / P3 / P3-reissue / P5 / P5-reserve | 0.20 worst case from max_tokens 8,000 across 3 calls | 0.0917849272 | per-response usage.cost, runs/reader-costs.json, recomputed from the stored bodies by analysis/verify.py | P1 OpenAI ($0.015425), P3 xAI ($0.0130304 on the re-issue), P5's reserve moonshotai/kimi-k3 at Modal ($0.0353022). Two wasted bodies, $0.0280 total — see the row below. All 37 units answered by all three readers; FT returns 0 DECISION unanimously and YF returns exactly 1, also unanimously, with positive controls firing in all three languages. 459% of the declared worst case for this stage, because the stage dispatched five calls where three were priced and one of the accepted seats was the expensive kimi-k3 rather than the cheap deepseek it replaced. |
| 2026-07-30 | S066 | E-20260730g condition D — 4 dispatches: generation P2 / P5 / P1-reserve, rating P3 | 0.17 worst case from max_tokens 8,000 ×2 and 4,000 ×1 | 0.0903335692 | per-response usage.cost, runs/d-costs.json | P2 Google ($0.0469935, 6,096 completion tokens on a 488-word prompt), P1 as the undeclared reserve at OpenAI ($0.024875), P3 rating at xAI ($0.0071384). One wasted body ($0.0113267). 7 of 12 sites keepable — and the four Futabatei anchors came back at 2, 4, 2, 1, so the threshold that certified them is uncalibrated for the period. 53% of the declared worst case. |
| 2026-07-30 | S066 | the three wasted calls | (inside the stage figures above) | 0.0393540 wasted | per-response usage.cost on the rejected bodies; asserted by analysis/verify.py | Note (b), fifteenth and sixteenth firings: deepseek/deepseek-v4-pro at max_tokens 8,000, finish_reason: length, empty body, at two different providers in one session — GMICloud $0.0145573272 and StreamLake $0.0113266692. Both fell through and both fall-throughs were accepted first call. And a THIRD failure shape neither note (b) nor note (beq) covers: x-ai/grok-4.5 returned stop with content: null and one answer line inside message.reasoning, $0.01347 for nothing; dropping reasoning: {"effort":"low"} and re-issuing to the same slug returned a complete answer. Note (bfb). One of the three failing stages had no declared reserve — note (bfc). |
| 2026-07-30 | S066 | both translation limbs, the two canon texts, the alignment, the permutation nulls and every analysis | 0.00 | 0.000000 | — | No API call. 1,074 Russian words rendered twice — T-svidanie-R06-v1 frozen at 828a8c1, T-svidanie-R09-v1 at b9abe12 — the Russian source established from two witnesses with two emendations and a six-site punctuation variant register, 「あいびき」 fetched and not read until both limbs were committed, and build_items.py, analyse.py, score_a.py, score_d.py and the 62-check analysis/verify.py are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-30 (S067 start) | 31.758610502 | S067 session start, before any project spend. $0.0149421 above S066's closing 31.743668402 — the smallest inter-session drift this ledger has recorded, and consistent with note (bco)'s settling lag on S066's final call rather than with note (abf)'s unattributed opening drift. |
| 2026-07-30 (S067, after two aborted critic sockets) | 31.758610502 | Identical to the opening figure to the digit, taken immediately after the second HTTP 504. Persisted at runs/snap/after-critic-fail.json before being read. This snapshot is what makes the closing excess attributable rather than merely unexplained — see the closing row. |
| 2026-07-30 (S067 end) | 32.02742932 | S067 after eight dispatches. Delta from the opening snapshot 0.268818818 against a per-request sum of 0.2250704184 — an EXCESS of $0.0437483996, the first time in this ledger that the delta has EXCEEDED the sum. Every previous cross-check ran short (note (bco)). The two aborted 504 sockets returned no usage object and are the only dispatches the per-request sum cannot see; a concurrent session on the same key is the standing alternative and cannot be excluded. S061 wrote the check that catches this and recorded that no excess appeared; it appears now. Note (bfe). |
| 2026-07-30 | S067 | E-20260730h-strict-coverage independent pre-run critic — P4 moonshotai/kimi-k3, three dispatches, two aborted | 0.36 worst case from max_tokens 12,000 at list out-price plus the prompt at list in-price, times a 1.5 provider premium (notes (abc), (x), (b)) | 0.066069 (per-request; see the excess row above) | per-response usage.cost; opening snapshot runs/snap/session-open.json (note (bco)) | Attempts 0 and 1 at max_tokens 12,000 returned HTTP 504 "The operation was aborted", no body, no usage object; attempt 2 at max_tokens 8,000 with effort: low returned in 8,443 / out 2,716, stop, provider Moonshot AI. Note (bdl)'s remedy on a new failure shape — note (bfe). 18% of worst case on the per-request figure. P1/P2/P3 are raters or the raters' reserve, so none could critique; P4 is the standing critic seat (S053 fix). Verdict NEEDS-REDESIGN, seven findings — three BLOCKING, two MANDATORY, two ADVISORY, all seven accepted. Note (rr), twenty-fifth consecutive session. Its BLOCKING finding 3 caught this session's own synthetic control repeating the exact defect the control existed to catch, and the repair then found the same defect in two of the 39 published D codes. |
| 2026-07-30 | S067 | E-20260730h condition C2, the three per-option raters — 4 dispatches, P1 / P3 / P5 / P5-reserve | 0.42 worst case from max_tokens 8,000 across 3 calls + 3 reserves | 0.1411029684 | per-response usage.cost, runs/cost.json, recomputed from the stored bodies by analysis/verify.py | P1 OpenAI ($0.03030375), P3 xAI ($0.0576924), P2 Google as the DECLARED reserve in the P5 seat ($0.040146). Note (b), seventeenth firing: deepseek/deepseek-v4-pro at max_tokens 8,000 with effort: low, finish_reason: length, empty body, 29,339 characters of unreturned reasoning, provider GMICloud, $0.0131607684 wasted; rejected body preserved at C2-P5.attempt0.raw (note (bdt)). The reserve was declared in the frozen design for every seat — note (bfc) working as written for the first time since S066 created it. All 41 items answered by all three seats; controls 4 of 4 for every rater, where the S064 instrument scored 1 of 3. 34% of worst case. |
| 2026-07-30 | S067 | E-20260730h the byte-identical repeat control — 1 dispatch, P1 | 0.14 worst case from max_tokens 8,000 | 0.0176985 | per-response usage.cost | P1 OpenAI, in 3,258 / out 2,895, stop. 0.9216 cell agreement against a registered 0.85 floor and against the 0.757 the S064 three-way task returned. 13% of worst case, and the cheapest call of the session bought the control that decides whether anything else may be reported. |
| 2026-07-30 | S067 | the whole translation limb, the recount of 39 published D codes, the item build, the contamination measurement and every analysis | 0.00 | 0.000000 | — | No API call. 955 Russian words of Gogol rendered under R07 with 55 sites dual-coded — Unit A frozen at 0c812d0 before the contamination gate, the rest at 8a11440 — the design and its six predictions frozen at e602222 before the source was read and before the Italian and Bengali logs were opened, and parse_logs.py, build_items.py, recount.py, analyse.py and the 242-check analysis/verify.py with four mutation tests, all four caught are lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
2026-07-30 day total: $3.2273135257 of $5.00 — seven sessions (S061 $0.5664084814, S062 $0.4046048663, S063 $0.5527697100, S064 $0.2024119810, S065 $1.0005113722, S066 $0.2755366964, S067 $0.2250704184). $1.7726864743 headroom. S067: 8 dispatches, 6 billed, 5 accepted, 1 wasted ($0.0131607684) and 2 aborted sockets that appear in the key delta and not in the sum.
S067 came in at 24.5% of its declared worst case, and the composition has one honest wrinkle. Declared $0.92 across three stages; actual $0.2250704184 on the per-request method, $0.268818818 if the key delta is taken instead — 29% of worst case either way, inside the 15–34% band this ledger has always landed in. The wrinkle is that the two figures disagree for the first time in the direction that matters: the delta is larger, so for once the conservative reading is the delta and not the sum. Both are printed above and the day total uses the sum, per config/budget.md's stated method; if the excess is the aborted sockets, the day total understates by $0.0437483996 and the headroom is $1.7289 rather than $1.7727.
(S066's day-total line, superseded by the S067 line above and kept for the audit trail:) 2026-07-30 day total: $3.0022431073 of $5.00 — six sessions (S061 $0.5664084814, S062 $0.4046048663, S063 $0.5527697100, S064 $0.2024119810, S065 $1.0005113722, S066 $0.2755366964). $1.9977568927 headroom. S066: 10 dispatches, 7 accepted, 3 wasted ($0.0393540), two declared reserves fired and one undeclared reserve had to be chosen after a failure.
S066 came in at 39% of its declared worst case and the composition is worth one line. Declared $0.70 across four stages; actual $0.2755366964. The critic stage, which is where S065 overran, landed at 28% of a worst case that had been deliberately inflated 1.5× for a provider premium that did not recur — the same slug routed to Together rather than Fireworks and billed under list. The stage that overran was condition A, at 459%, because three declared calls became five and the accepted seat was the expensive one. Note (abc)'s lesson repeats in a new place: a worst case built per-call is not a worst case when the call COUNT is what varies.
(S065's day-total line, superseded by the S066 line above and kept for the audit trail:) 2026-07-30 day total $2.7267064109 of $5.00 — five sessions, $2.2732935891 headroom.
| 2026-07-30 (S068 start) | 32.02742932 | S068 session start, before any project spend. Identical to S067's closing figure to the digit — a zero inter-session drift, the first this ledger has recorded. |
| 2026-07-30 (S068 end, immediate) | 32.26963482 | Taken seconds after the last dispatch. Delta from the opening snapshot 0.2422055 against a per-request sum of 0.49198902 — a shortfall of $0.2498, the largest this ledger has seen. Not reported as the session figure; see the row below. |
| 2026-07-30 (S068 end, settled) | 32.51941834 | The same key read again after the analysis and verification were written. Delta from the opening snapshot 0.49198902 — EXACT to 1e-8 against the per-request sum. This is the cleanest demonstration note (bco) has: the settling lag is real, it is large in the first minute, and it resolves completely. It also bears on S067's unexplained excess of $0.0437: an immediate closing snapshot is not a valid cross-check in either direction, and S067's excess was read off one. Snapshots at runs/snap/. |
| 2026-07-30 | S068 | E-20260730i-candidate-reach independent pre-run critic — P4 moonshotai/kimi-k3, accepted first call | 0.25 worst case from max_tokens 8,000 at list out-price plus the prompt at list in-price, times a 1.5 provider premium (notes (abc), (x), (b)) | 0.0910722 | per-response usage.cost; opening snapshot runs/snap/session-open.json (note (bco)) | in 10,340 / out 4,015, stop, provider Modal, effort: low at max_tokens 8,000 — S067's working configuration reused rather than rediscovered after its two 504s. 36% of worst case. Verdict NEEDS-REDESIGN, ten findings — three BLOCKING, four MANDATORY, three ADVISORY, all ten accepted. Note (rr), twenty-sixth consecutive session. Its BLOCKING finding 1 is the session's result: the census had produced no n = 2 sites, so nothing in the frozen design could separate the two hypotheses, and the ladder it prescribed is what did. |
| 2026-07-30 | S068 | E-20260730i condition C, three per-option raters + the byte-identical repeat | 0.60 worst case from max_tokens 8,000 across 4 calls (revised from 0.40+0.12 when the payload grew from 49 to 66 items) | 0.1546085 | per-response usage.cost, recomputed from the stored bodies by analysis/verify.py | P1 OpenAI ($0.040569), P3 xAI ($0.0717404), P5 Cloudflare ($0.0120669), repeat P1 OpenAI ($0.0302322). All 66 items answered by all three seats; controls 6/6, 6/6, 3/6, so F2 does not fire. No reserve needed. 26% of worst case, and the cheapest call of the four bought the repeat control. |
| 2026-07-30 | S068 | E-20260730i condition A — 5 dispatches for 3 seats, and one seat was never filled | 0.30 worst case from max_tokens 4,000 across 3 calls | 0.24630832 | per-response usage.cost | P1 OpenAI ($0.026757) and P3 xAI ($0.0678804) accepted. Three wasted bodies, $0.15167092: deepseek-v4-pro at AtlasCloud returned finish_reason: length, empty content, 15,897 characters of unreturned reasoning ($0.02537242, note (b), eighteenth firing); the declared reserve google/gemini-3.6-flash then truncated at 4,000 ($0.04065, 27 of 69 lines) and again at 10,000 ($0.0856485, 67 of 69) — note (bfh). The seat was abandoned rather than bought a fourth time, and condition A is reported on two raters, which is the number RS-20260728i itself used. 82% of the declared worst case for this stage, and 62% of it is waste. |
| 2026-07-30 | S068 | the whole translation limb, the two-witness collation, the option census, the contamination gate, the item build and every analysis | 0.00 | 0.000000 | — | No API call. 1,221 Portuguese words of Machado de Assis rendered under R04 — the project's first PT→EN — with a 45-site option census frozen at 5f10c81 before any candidate was consulted, Unit A frozen at 10b10b6 before the contamination gate, the design and its predictions frozen at 8e07d6d before the source was read, and parse_census.py, build_items.py, prompts.py, analyse.py and the 67-check analysis/verify.py with four mutation tests, all four caught are lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
2026-07-30 day total: $3.7193025457 of $5.00 — eight sessions (S061 $0.5664084814, S062 $0.4046048663, S063 $0.5527697100, S064 $0.2024119810, S065 $1.0005113722, S066 $0.2755366964, S067 $0.2250704184, S068 $0.49198902). $1.2806974543 headroom. S068: 10 dispatches, 7 accepted, 3 wasted ($0.15167092, 31% of the session's spend, all on one unfilled seat).
S068 came in at 43% of its declared worst case, and the composition is the least flattering this ledger has recorded. Declared $1.15 after the amendment; actual $0.49198902. The stage that overran was condition A at 82%, and 62% of that stage was waste — note (abc)'s lesson in the S066 form again: a worst case built per-call is not a worst case when the call COUNT is what varies, and here the count went 3 → 5 for two usable answers. The cross-check is the cleanest in the ledger's history (delta exact to 1e-8 on the settled snapshot) and it retrospectively weakens the reading of S067's excess.
(S067's day-total line, superseded by the S068 line above and kept for the audit trail:) 2026-07-30 day total: $3.2273135257 of $5.00 — seven sessions, $1.7726864743 headroom.
UTC day 2026-07-31
| date | session | what | pre-flight | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-07-31 | S069 | E-20260731-anchor-second-read pre-run critic, first seat, ABANDONED — moonshotai/kimi-k3 (P4) |
0.191 worst case from max_tokens 12,000 (note (abc)) |
~0.1938, inferred, no per-request record | key-usage delta minus the per-request sum; nothing was returned, so there is no usage.cost row |
Note (b), and new note (bfo). The seat held the socket open past twenty minutes and was abandoned with no body, no usage row and nothing written to disk — note (beh) exactly: a fall-through chain protects against a call that returns a failure and does nothing against one that hangs. The abandoned call nevertheless billed at essentially its full cap. This is the first time this project has priced an abandoned hang, and it had been assuming they were free. |
| 2026-07-31 | S069 | E-20260731-anchor-second-read pre-run critic, replacement seat — qwen/qwen3.7-max |
(not separately estimated; the abandonment created it) | 0.042384125 | per-response usage.cost |
in 5,125 / out 7,870, stop, 143.7 s, provider Alibaba. Probed-but-not-selected, not a renderer and not the renderers' reserve, so the S053 role-collision fix holds — thirteenth session running. Verdict NEEDS-REDESIGN, six findings, two BLOCKING, all six accepted, none declined. Note (rr), twenty-second consecutive session, and one of its largest returns: it withdrew the design's own headline question before the run and replaced it with a free measurement and a within-renderer contrast the confound cannot reach. |
| 2026-07-31 | S069 | E-20260731-anchor-second-read the renderer pass — 3 seats + 1 reserve, EN→JA |
0.074 worst case for three seats from max_tokens 4,000, P5 priced at the worst plausible provider (S022 routing caution) |
0.0730877328 | per-response usage.cost, re-summed from the .raw bytes by analysis/verify.py |
P1 OpenAI ($0.012199, stop); P3 xAI ($0.0121604, stop); P5 GMICloud ($0.0062078328, finish_reason: length at 4,000 tokens, 40% of the span — seat FAILED and replaced); reserve P2 Google ($0.0425205, stop, 8,000 cap). 99% of the three-seat worst case, and the overrun is the reserve, which was not in the estimate. Note (bfh)'s lesson held this time: the reserve differed in failure mode and it worked first call. |
| 2026-07-31 | S069 | the translation limb, the census, the contamination gate, both audit arms and every analysis | 0.00 | 0.000000 | — | No API call. Poe ¶4–9 rendered from the English into Japanese (822 → 2,097 characters), the site census frozen and committed before a word was written, the contamination measurement via tools/dependence_check_cjk.py, and stage1.py, counts.py, score.py, render.py, call.py and the 41-check analysis/verify.py with five mutation tests are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-31 | S070 | E-20260731b-figure-audit pre-run critic, ABANDONED at 120 s — qwen/qwen3.7-max | (not separately estimated; the client timeout created it) | 0.000000 | key-usage delta minus the per-request sum, and the residual is zero | Note (bfo), AMENDED. The call was killed by a 120-second client timeout on a seat that had taken 143.7 s the day before. The session's key-usage delta equals the sum of its four accepted calls to 1e-9, so this abandoned call billed nothing — which is the opposite of what S069's twenty-minute hang did. (bfo)'s generalisation is withdrawn; its own case stands. |
| 2026-07-31 | S070 | E-20260731b-figure-audit pre-run critic, replacement — qwen/qwen3.7-max | 0.060 worst case from max_tokens 8,000 (note (abc)) | 0.047567570 | per-response usage.cost | in 5,125→10,724 / out 9,974, stop, provider Alibaba. Probed-but-not-selected seat: not a rater, not the raters' reserve, so the S053 role-collision fix holds — fourteenth session running. Verdict NEEDS-AMENDMENT, five findings, one BLOCKING, all five accepted, none declined. Note (rr), twenty-third consecutive session. Its BLOCKING finding split the lead's twelve composite declarations into endorsed and contested tiers sight-unseen, and the label then fired only in the tier it endorsed. |
| 2026-07-31 | S070 | E-20260731b-figure-audit the positive control — 3 rater seats, condition A | 0.098 worst case for three seats from max_tokens 4,000 | 0.050587150 | per-response usage.cost | P1 OpenAI ($0.017395750, stop); P2 Google ($0.024315000, stop); P3 xAI ($0.008876400, stop). 52% of the three-seat worst case, no reserve fired. Same three seats as E-20260728f, whose five-option block was reused verbatim. |
| 2026-07-31 | S070 | the sweep, the translation limb, the contamination gate and every analysis | 0.00 | 0.000000 | — | No API call. Thirteen published agreement figures recomputed from stored bodies; 蒲松齡〈促織〉 rendered entire from the Chinese (1,828 chars → 2,280 words) with an R06 draft frozen separately first; the contamination measurement via tools/dependence_check.py; and sweep.py, build_items.py, analyse.py, the 195-check verify.py and the six-mutation mutation_test.py are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-31 | S071 | E-20260731c-futabatei-alignment pre-run critic — qwen/qwen3.7-max (probed-but-not-selected) | 0.13 worst case from max_tokens 8,000 (note (abc)) | 0.050161800 | per-response usage.cost | in 5,253 / out 9,585, stop, 174 s, provider Alibaba. Not an alignment seat and not their reserve, so the S053 role-collision fix holds — fifteenth session running. Verdict NEEDS-AMENDMENT, six findings, two BLOCKING, all six accepted, none declined. Note (rr), twenty-fourth consecutive session. Its second BLOCKING finding withdrew the design's own crux (PL4) before the run and forced the search that established no native unforced Japanese comparator exists. |
| 2026-07-31 | S071 | E-20260731c-futabatei-alignment the independent alignment check — 3 seats, 24 items, one call each | 0.18 worst case for three seats from max_tokens 4,000, P3 priced at the worst plausible provider (S022 routing caution) | 0.093463900 | per-response usage.cost | P1 OpenAI ($0.026682, stop); P2 Google AI Studio ($0.0275535, stop); P3 xAI ($0.0392284, stop). 52% of the three-seat worst case; no reserve fired and note (b) did not fire. A reserve (P5) was declared for the stage before dispatch, per note (bfc). |
| 2026-07-31 | S071 | the translation limb, the aligner, the hand alignment, the contamination gate and every analysis | 0.00 | 0.000000 | — | No API call. Turgenev «Свидание» ¶42–¶69 rendered into Japanese twice (669 Russian words, unforced then R09-forced, each frozen in git before the next existed); the dialogue-anchored aligner, the hand alignment, both permutation and length-rate nulls, tools/dependence_check_cjk.py used unmodified, and align.py, hand_alignment.py, build_items.py, analyse.py, followups.py, contamination.py and the 68-check verify.py with six mutation tests are all lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-31 | S072 | E-20260731d-sense-axes pre-run critic — qwen/qwen3.7-max (probed-but-not-selected) | 0.13 worst case from max_tokens 8,000 (note (abc)) | 0.052175175 | per-response usage.cost | in 7,965 / out 9,136, stop, 166 s, provider Alibaba. Not a rater and not the raters' reserve, so the S053 role-collision fix holds — sixteenth session running. Verdict ACCEPT, one finding, MINOR, accepted and applied (a P5 name collision between the reserve seat and prediction P5). Note (rr), twenty-fifth consecutive session, and its lightest return yet: asked to endorse or contest each of the 8 positive-control items sight-unseen, on the S070 precedent, it endorsed all eight, so nothing was excluded and the control set is the lead's set with a second opinion on it. |
| 2026-07-31 | S072 | E-20260731d-sense-axes the CAT block — 3 seats, 40 items, one call each | 0.15 worst case for three seats from max_tokens 6,000 | 0.057972650 | per-response usage.cost, re-summed from the .raw bytes by verify.py | P1 OpenAI ($0.02027225, stop); P2 Google AI Studio ($0.02064, stop); P3 xAI ($0.0170604, stop). 39% of the three-seat worst case; no reserve fired and note (b) did not fire. A reserve (deepseek/deepseek-v4-pro) was declared for the stage before dispatch, per note (bfc). |
| 2026-07-31 | S072 | E-20260731d-sense-axes the GRAD block — the same 3 seats, the same 40 items reshuffled | 0.15 worst case for three seats from max_tokens 6,000 | 0.045052650 | per-response usage.cost, re-summed from the .raw bytes | P1 OpenAI ($0.01718325, stop); P2 Google AI Studio ($0.016017, stop); P3 xAI ($0.0118524, stop). 30% of the three-seat worst case. |
| 2026-07-31 | S072 | the translation limb, the contamination gate, both nulls and every analysis | 0.00 | 0.000000 | — | No API call. Leskov «Левша» chapters 6–8 rendered entire from the Russian (974 words → 1,376 English) with a 54-decision log, frozen in git before the design existed; the contamination gate via tools/dependence_check.py unmodified; and build_items.py, run_critic.py, run.py, analyse.py and the 78-check verify.py with seven mutation tests are all lead work or local computation. Nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4). |
Declared worst case $0.50; actual $0.155200475 — 31%, inside the 15–48% band this ledger has been running. Note (abc); the estimate was built from max_tokens and not from an assumed output length, and no amendment moved it in session.
THE KEY-USAGE CROSS-CHECK RECONCILES THIS SESSION, WHICH IS WORTH SAYING AFTER S071. Opening snapshot, taken immediately before the first call, 35.969639767; settled 36.124840242; delta 0.155200475. Per-request sum over the seven accepted calls 0.155200475. Equal to 1e-9. Note (bfv)'s prescription was followed — snapshots around each call rather than around the session — and no inter-call interval moved: every one of the seven in-session deltas read 0.000000000 at the moment of the call, the settling lag of note (bcx), and the whole amount appeared in the settled reading afterwards. No cross-session delta is computed against S071, whose own cross-check was void.
| 2026-07-31 | S073 | E-20260731e-option-census pre-run critic — qwen/qwen3.7-max (probed-but-not-selected), two dispatches | 0.13 worst case from max_tokens 8,000 (note (abc)) | 0.044181855 | per-response usage.cost; opening key snapshot runs/snap/session-open.json | First dispatch WASTED and BILLED NOTHING: finish_reason: error, content: null, 4,019 reasoning tokens, cost: 0.000000 in the usage block though upstream_inference_cost read 0.025619275 — note (b), nineteenth firing, and the first time this ledger has seen upstream_inference_cost present on a call that was not billed. Second dispatch at max_tokens 12,000: in 5,793 / out 9,180 of which 8,508 reasoning, stop, 167 s, provider Alibaba. The seat needed more than 8,000 tokens to think before emitting anything, which is exactly what note (b) says and exactly what the estimate failed to allow. Not a rater and not the raters' reserve — S053 role-collision fix, seventeenth session running. Verdict NEEDS-AMENDMENT, three findings, one BLOCKING, all three accepted. Note (rr), twenty-seventh consecutive session. |
| 2026-07-31 | S073 | E-20260731e round 1, arm NONE — 3 seats | 0.28 worst case for three seats from max_tokens 6,000 (note (abc)), P3 priced at the worst plausible provider (S022 routing caution) | 0.057058150 | per-response usage.cost, re-summed from the .raw bytes by analysis/verify.py | P1 OpenAI ($0.01251375), P2 Google ($0.023202), P3 xAI ($0.0213424); all stop, all 32 items. 20% of the three-seat worst case; no reserve fired. A reserve (deepseek/deepseek-v4-pro) was declared for every stage before dispatch, per note (bfc). |
| 2026-07-31 | S073 | E-20260731e round 2, arm SHAM — 3 seats, 4 dispatches | 0.28 worst case for three seats | 0.082559050, of which 0.012084400 wasted | per-response usage.cost, re-summed from the .raw bytes | P1 OpenAI ($0.01603625), P2 Google ($0.029256). P3 xAI FAILED with a new shape — finish_reason: stop, content: null, the answer's FIRST ITEM ONLY inside message.reasoning, 1,001 completion tokens, and no reasoning.effort set at all — which is note (bfb)'s shape without the parameter (bfb) blames, and is new note (bga). A plain retry of the identical payload to the identical slug returned all 32 items ($0.0251824), so the reserve was not needed and was not used. |
| 2026-07-31 | S073 | E-20260731e round 3, arm FRAMEWORK — 3 seats | 0.28 worst case for three seats | 0.066622900 | per-response usage.cost | P1 OpenAI ($0.0199825), P2 Google ($0.025548), P3 xAI ($0.0210924); all stop, all 32 items. 24% of the three-seat worst case. |
| 2026-07-31 | S073 | E-20260731e round 4, arm NONE repeated BYTE-IDENTICALLY — the same 3 seats | 0.28 worst case for three seats | 0.056132500 | per-response usage.cost | P1 OpenAI ($0.0100161), P2 Google AI Studio ($0.027996), P3 xAI ($0.0181204). This is the cross-day drift gate's option 1 and the first design to take it by choice, and it is the cheapest thing in this ledger that has ever changed a conclusion: the floor it measured (0.375 on P2, 0.000 on P1) is larger than the effect the design was built to detect, so the headline went from a two-of-three-seat rise to a null. Note the provider change on P2 between two byte-identical requests — Google → Google AI Studio, which note (beo) says is not the same served model. |
| 2026-07-31 | S073 | both contamination gates, the translation limb, the census, the item build and every analysis | 0.00 | 0.000000 | — | No API call. Reymont «Śmierć» ¶0–51 rendered from the Polish (904 words → 1,132 English), the project's first Polish, with an R06 draft frozen in its own commit first and a 40-site option census; plus 308 abandoned words of Sienkiewicz, translated and then dropped when the gate returned 18 tokens. tools/dependence_check.py used unmodified for both gates; prompts.py, call.py, run_critic.py, run.py, analyse.py and the 116-check analysis/verify.py with four mutation tests, all four caught are lead work or local computation. Lead translation is free and is never ledgered (charter §3, A4). |
Declared worst case $1.20; actual $0.306554455 — 26%, inside the 15–48% band this ledger has been running. Note (abc); the estimate was built from max_tokens and no amendment moved it in session.
THE KEY-USAGE CROSS-CHECK RECONCILES EXACTLY, AND IT SETTLES A QUESTION ABOUT WASTED CALLS. Opening snapshot 36.131037442 (taken immediately before the first call, and $0.006197675 above S072's settled close — the settling lag of note (bcx), still present and still small); settled close 36.437591897; delta 0.306554455. Per-request sum over the thirteen accepted-or-billed calls 0.306554455. Equal to 1e-9. Both wasted dispatches therefore billed nothing — the errored critic call and the null-content grok-4.5 call — which is S070's outcome rather than S069's, so note (bfo) stays withdrawn and its own case stands alone. And upstream_inference_cost is not what is billed: the errored critic call carries $0.025619275 in that field and contributed $0.000000000 to the delta.
2026-07-31 day total: $1.012808407 of $5.00 — five sessions (S069 $0.309273057, S070 $0.098154720, S071 $0.143625700, S072 $0.155200475, S073 $0.306554455). $3.987191593 headroom. 20% of the day's cap.
(S072's day-total line, superseded by the S073 line above and kept for the audit trail:) 2026-07-31 day total: $0.706253952 of $5.00 — four sessions (S069 $0.309273057, S070 $0.098154720, S071 $0.143625700, S072 $0.155200475). $4.293746048 headroom. 14% of the day's cap.
(S071's day-total line, superseded by the S072 line above and kept for the audit trail:) 2026-07-31 day total: $0.551053477 of $5.00 — three sessions (S069 $0.309273057, S070 $0.098154720, S071 $0.143625700). $4.448946523 headroom.
The declared worst case was RAISED in session, from $0.25 to $0.30, with the reason written before dispatch — the check's max_tokens went 3,000 → 4,000 against note (b), which has fired sixteen times on empty length returns. Actual $0.1436257, 48% of the amended worst case. Note (abc).
THE KEY-USAGE CROSS-CHECK IS VOID THIS SESSION AND THAT IS ITSELF THE FINDING — note (bfv), new. The key read 34.354965717 at the critic call and 35.505091267 immediately before the alignment check, a rise of $1.149 across an interval in which this session made no call whatever; and the session's opening figure was $1.198 above S070's closing figure of 33.156987017. Two unattributable movements of over a dollar each. The per-request usage.cost sums are primary and are unaffected (CLAUDE.md: per-request costs are primary, the delta is a sanity check) — but the sanity check cannot be run, and reporting the residual as drift would have been wrong. Take snapshots around each CALL, not around the session, and declare the cross-check void when any inter-call interval moves. This is the concurrency hazard NEXT.md has carried since S048, appearing for the first time as a corrupted instrument rather than as a worry about id collisions.
2026-07-31 day total: $0.551053477 of $5.00 — three sessions (S069 $0.309273057, S070 $0.098154720, S071 $0.143625700). $4.448946523 headroom. (This total is the sum of per-request costs across the three sessions; it is NOT reconcilable against the key-usage figure this session, for the reason above.)
(S070's day-total line, superseded by the S071 line above and kept for the audit trail:) 2026-07-31 day total: $0.407427777 of $5.00 — two sessions (S069 $0.309273057, S070 $0.098154720). $4.592572223 headroom.
S070's key-usage delta is exact. Opening 33.058832297, settled 33.156987017, delta 0.098154720; per-request sum 0.098154720. Equal to 1e-9, which is what establishes that the abandoned call cost nothing. (The opening snapshot sat 0.006736 above S069's closing figure — the settling lag of note (bcx), still present.)
The ledgered figure is the KEY DELTA, not the per-request sum, and this is the first day that
distinction has mattered. Opening snapshot 32.74282324, closing 33.052096297, delta
0.309273057; the per-request sum over the five accepted calls is 0.1154717578. The
$0.1938 difference is the abandoned kimi-k3 call, which has no usage row to sum. Note (bet)
says the session-level bound is the valid check and note (bco) says settling runs in the shortfall
direction; this is the excess direction with a known cause, and the honest ledger entry is the
delta. Five billed calls accepted, one dispatch abandoned and billed.
| 2026-07-31 | S074 | E-20260731f-catalogue-reach pre-run critic — moonshotai/kimi-k3, one dispatch, WASTED | 0.20 worst case from max_tokens 12,000 (note (abc)) | 0.195747000 | per-response usage.cost; opening key snapshot runs/snap/critic-open.json | finish_reason: length, 11,997 reasoning tokens, ZERO characters of content. Note (b), TWENTIETH firing, and the first time a cap raised because of this note was still not enough — 12,000 was chosen precisely because S073 lost the same seat at 8,000. Unlike S073's wasted dispatches, this one BILLED. |
| 2026-07-31 | S074 | E-20260731f pre-run critic, re-dispatched to the DECLARED RESERVE — deepseek/deepseek-v4-pro | included above | 0.030371504 | per-response usage.cost | Provider Novita, stop, 135 s, 9,346 reasoning tokens of 10,390. Verdict NEEDS-REDESIGN, six findings, two BLOCKING, all six accepted, amendments A1–A9. The reserve is not a rater in this design, so the S053 role-collision fix is not breached. Note (bfc), working as designed for the first time on a critic fall-through this session. |
| 2026-07-31 | S074 | E-20260731f stage 2b, independent reachability typing — 2 seats, added by amendment A5 to answer the critic's BLOCKING Finding 1 | 0.10 worst case for two seats from max_tokens 4,000 | 0.049995200 | per-response usage.cost | google/gemini-3.6-flash (Google, $0.03036) and qwen/qwen3.7-max (Alibaba, $0.0196352); both stop, both 41 items. Neither is a rater nor the critic. The two calls that made the session's reach figure honest cost 6% of the session. |
| 2026-07-31 | S074 | E-20260731f stage 3 — own-passage second read of the residue, 3 passages × 2 raters, 8 dispatches for 6 cells | 0.32 worst case from max_tokens 6,000 | 0.164076600, of which 0.010324400 on a null retry | per-response usage.cost, re-summed from the .raw bytes by analysis/verify.py | P1 OpenAI throughout, P3 xAI throughout. One cell failed twice in two different shapes: s3-C-P3 first returned a 200 whose body was 2,167 bytes of pure whitespace and no JSON at all (new note (bgc), and it crashed the runner and aborted the two stages behind it), then returned finish_reason: stop with 176 completion tokens all reasoning and zero content (note (bfb)), then on a third dispatch returned all nine items. |
| 2026-07-31 | S074 | E-20260731f stage 4 — the crossover, 4 cells × 2 raters | 0.42 worst case | 0.346179450 | per-response usage.cost | All stop, all items parsed, 0 missing across 8 cells. x-ai/grok-4.5 cost 3.3× openai/gpt-5.6-terra on identical prompts ($0.4364 against $0.1346 over nine comparable cells); note (x)'s list-price assumption held on every call and the difference is output length, not routing. |
| 2026-07-31 | S074 | E-20260731f stage 5 — one cell re-sent BYTE-IDENTICALLY, 2 raters | 0.11 worst case | 0.060393400 | per-response usage.cost | Prompt bytes identical and prompt-token counts identical to the digit, both asserted by the verifier. Flip rate 5 of 34 cells, 14.7%, roughly double RS-20260730d's 7.2% on the same two seats. |
| 2026-07-31 | S074 | the translation limb, both contamination gates, stage 2 and every analysis | 0.00 | 0.000000 | — | No API call. Kielland «Karen» ¶1–15 rendered twice from the Norwegian (799 source words → 862 + 950 English), the project's first Norwegian, under the new regime R10 against two different Tier 1 anchors, with two frozen logs and a frozen opportunity list; tools/dependence_check.py used unmodified for both gates; build_claims.py, checks.py, prompts.py, run.py, run_typing.py, analyse.py and the 130-check analysis/verify.py with five mutation tests are lead work or local computation. Nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4). |
| 2026-07-31 | S075 | E-20260731g-provenance pre-run critic — qwen/qwen3.7-max (probed-but-not-selected), one dispatch after one abandoned | 0.18 worst case from max_tokens 12,000 (note (abc)) | 0.040804105 | per-response usage.cost; re-summed from the .raw bytes by analysis/verify.py; opening key snapshot runs/snap/session-open.json 37.474390751 | in 4,829 / out 8,738 of which 31,547 characters reasoning, stop, 158.9 s, provider Alibaba. 23% of the declared worst case. A reserve (openai/gpt-5.6-terra, a different lab) was declared before dispatch per note (bfc) and was not used; note (b) did not fire — the cap was left at 12,000 rather than raised, which is NEXT.md's reading that the remedy is the fall-through. Verdict NEEDS-AMENDMENT, seven findings, one BLOCKING, all seven accepted, none declined; two answered with a stronger remedy than proposed. Note (rr), twenty-eighth consecutive session. Its BLOCKING finding replaced the lead's self-written contamination floor with a deterministic non-lead null that then killed the session's own leading explanation. This session dispatched no other API call, so the S053 role-collision fix is vacuous rather than satisfied and is recorded as such — nineteenth session running. |
| 2026-07-31 | S075 | one dispatch ABANDONED before the above, 120-second client timeout on a seat that took 158.9 s when re-run | (not separately estimated; the timeout created it) | unknown — this session cannot price it | — | Note (bfo), NEITHER CONFIRMED NOR REFUTED. Nothing was written to disk. The closing key snapshot read 37.474390751, identical to the opening, so the delta is $0.000000 against a per-request sum of $0.040804105 — note (bco)'s settling lag has swallowed the accepted call as well as any abandoned one, and no attribution is possible at hand-off. S069's twenty-minute hang billed nearly its full cap; S070's 120-second timeout billed nothing. This one is a third case and it is recorded as unresolved rather than assigned to either. |
| 2026-07-31 | S075 | the archive sweep, the mutation harness, both translations, the contamination measurement, the anthology null, both length controls and every analysis | 0.00 | 0.000000 | — | No API call. 220 stored response bodies and 39 frozen verifiers swept; 54 verifier runs across four mutation passes; d'Annunzio «La fine di Candia» ¶1–24 rendered twice from the Italian (800 words → 971 and 848 English tokens) with 61 logged decisions and a frozen opportunity list; tools/dependence_check.py used unmodified; and sweep.py, mutate.py, measure.py, call.py, run_critic.py and the 80-check analysis/verify.py with six mutation tests are all lead work or local computation. Nothing in tools/ changed. Lead translation is free and is never ledgered (charter §3, A4). |
Declared worst case $1.30, RAISED IN SESSION to $1.60 with the reason written before the next dispatch — the wasted kimi-k3 dispatch ($0.195747), the reserve critic ($0.030372), and amendment A5's two extra calls. Actual $0.846762554 = 53% of the amended figure, above the 15–48% band this ledger has been running, because eleven of the twenty-one calls were prompt-dominated: each rater cell carries two passages of ~1,500 words. Note (abc); the estimate was built from max_tokens and not from an assumed output length.
THE KEY-USAGE CROSS-CHECK RECONCILES EXACTLY. Opening snapshot 36.437591897 (taken immediately before the first call, and equal to S073's settled close to the digit — no settling lag this session, unlike S073 and S070); settled close 37.284354451; delta 0.846762554. Per-request sum over the twenty-one billed bodies 0.846762554. Equal to 1e-9. Note (bet)'s rule was followed: only the session-level bound is asserted, because the per-stage snapshots are unreadable here — run.py stage3 ran twice after the (bgc) crash and overwrote its own stage snapshots.
| 2026-07-31 | S076 | E-20260731h-carryover pre-run critic — qwen/qwen3.7-max, one dispatch, declared reserve google/gemini-3.6-flash not used | 0.25 worst case from max_tokens 12,000 (note (abc)) | 0.053150150 | per-response usage.cost; opening key snapshot runs/snap/session-open.json 37.523401556 | Provider Alibaba, stop, 195.6 s, 10,775 completion tokens. Verdict NEEDS-REDESIGN — the second this project has taken since S021 — seven findings, three BLOCKING, three MANDATORY, one ADVISORY, all seven accepted, three answered with a stronger remedy than proposed and one literal remedy (raise temperature) declined in writing with a control offered instead. Note (rr), twenty-ninth consecutive session. 21% of the declared worst case. |
| 2026-07-31 | S076 | E-20260731h subjects, seat S1 — openai/gpt-5.6-terra, 9 cells | included in the $0.70 subject line | 0.071335950 | per-response usage.cost, re-summed from the .raw bytes by analysis/verify.py | Provider OpenAI on all nine, all stop, largest completion 1,717 tokens against a 2,200 cap. The cheapest of the three seats by a factor of two. |
| 2026-07-31 | S076 | E-20260731h subjects, seat S2 — x-ai/grok-4.5, 9 cells | included above | 0.166094800 | per-response usage.cost | Provider xAI on all nine, all stop. One cell returned 4,231 completion tokens against the declared 2,200 cap — max_tokens was not enforced for this slug, so note (abc)'s worst case was not in fact an upper bound. New note (bgk). |
| 2026-07-31 | S076 | E-20260731h subjects, seat S3, RUN 2 — deepseek/deepseek-v4-pro at max_tokens 8,000, 9 cells + 1 reserve | 0.35 worst case, declared in amendment A8 before the re-dispatch | 0.135711755 | per-response usage.cost | Providers Together then StreamLake throughout. Eight of nine returned clean at 5,072–7,427 completion tokens; the ninth burned all 8,000 and returned zero characters, falling through to google/gemini-3.6-flash ($0.0329415) — which is a different model, so that cell is VOID as a same-seat repeat. Note (b), twenty-second firing; new note (bgl) on the cap-raise having worked eight times and failed once. |
| 2026-07-31 | S076 | E-20260731h subjects, seat S3, RUN 1 — WASTED, and the fault is the lead's | (not separately estimated; it was the declared run) | 0.150118077 | per-response usage.cost, preserved under runs/void-run1-qwen/ | deepseek-v4-pro returned finish_reason: length with 2,200 tokens and zero characters of content on three cells (note (b), twenty-first firing, and the first ever on a translation SUBJECT rather than a critic or rater), and the reserve that answered them was qwen/qwen3.7-max — this design's own pre-run critic. That is the S053 role collision, written into run.py's reserve table by the lead and invisible to the critic, who was shown design.md and not the runner. All three renderings are VOID and the run is preserved entire. 26% of the session's spend. New note (bgj). |
| 2026-07-31 | S076 | the four lead translations, the opportunity list, both comparator extractions and every measurement | 0.00 | 0.000000 | — | No API call. Mikszáth «Szent Péter esernyője» I.iii, 535 Hungarian words rendered four times (1,366 English words, 110 logged decisions), the project's first Hungarian and its fifteenth source language, with the R10 §6 opportunity list frozen before any of them and the arm order crossed within one work for the first time; tools/dependence_check.py and tools/build_index.py used unmodified; prompts.py, run.py, analyse.py, mutate.py, a multi-turn call.py and the 403-check analysis/verify.py with six mutation tests, six caught and a passing control are lead work or local computation. Nothing in tools/ changed. Lead translation is free and is never ledgered (charter §3, A4). |
S076's declared worst case was $0.45, RAISED IN SESSION to $0.95 with the reason written before the re-dispatch (amendment A8: the voided run 1, and S3's cap raised from 2,200 to 8,000). Actual $0.576410732 = 61% of the amended figure, above the 15–53% band this ledger has been running — and the reason is not estimation error but the $0.150118077 of voided work, which is 26% of the total. Without it the session is $0.426292655, or 45%.
THE KEY-USAGE CROSS-CHECK IS SHORT, WITH THE USUAL CAUSE. Opening snapshot 37.523401556 (taken immediately before the critic dispatch), closing 38.05968395, delta 0.536282394; per-request sum over the twenty-nine billed bodies 0.576410732. Shortfall $0.040128338 — note (bco)'s settling lag in the direction it usually runs, and unlike S075 there is no abandoned dispatch to attribute it to: every call this session returned a body with a usage row. CLAUDE.md makes per-request costs primary and the delta the sanity check, so $0.576410732 is the ledgered figure and note (bet)'s session-level bound is the only one asserted.
2026-07-31 day total: $2.476785798 of $5.00 — eight sessions (S069 $0.309273057, S070 $0.098154720, S071 $0.143625700, S072 $0.155200475, S073 $0.306554455, S074 $0.846762554, S075 $0.040804105, S076 $0.576410732). $2.523214202 headroom. 50% of the day's cap, and the first day in this ledger's history to carry eight sessions.
(S075's day-total line, superseded by the S076 line above and kept for the audit trail:) 2026-07-31 day total: $1.900375066 of $5.00 — seven sessions (S069 $0.309273057, S070 $0.098154720, S071 $0.143625700, S072 $0.155200475, S073 $0.306554455, S074 $0.846762554, S075 $0.040804105). $3.099624934 headroom. 38% of the day's cap. S075 is the cheapest session on this day by a factor of two and its principal unit cost nothing at all: an archive sweep, 54 verifier runs, 1,600 words of translation and every measurement were local computation, and the single call was the pre-run critic.
(S074's day-total line, superseded by the S075 line above and kept for the audit trail:) 2026-07-31 day total: $1.859570961 of $5.00 — six sessions. $3.140429039 headroom. 37%.
| 2026-08-01 | S077 | E-20260801-strata pre-run critic — qwen/qwen3.7-max, one dispatch after one lost to note (bgc); declared reserve moonshotai/kimi-k3 not used | 0.36 worst case from max_tokens 12,000 (note (abc)) | 0.053022120 | per-response usage.cost; opening key snapshot runs/snap/session-open.json 38.114613848 | Provider Alibaba, stop, 10,981 in / 11,121 out of which 36,511 characters reasoning. Verdict NEEDS-AMENDMENT, three findings, one BLOCKING, all three accepted, two answered with a stronger remedy than proposed. Note (rr), thirtieth consecutive session. 15% of the declared worst case. Its BLOCKING finding caught a synthetic control that would have voided the whole run through the design's own F1 — the third mis-built control in this line of work and the third caught only from outside (note (bgo)). |
| 2026-08-01 | S077 | the session's FIRST dispatch, LOST — note (bgc) | (not separately estimated; it was the critic dispatch) | unknown — this session cannot price it | — | HTTP 200 with a body that is not JSON, which raised JSONDecodeError and crashed the runner before any guard existed. The bytes were not preserved; the guard built in response (call.NonJSONBody) preserved the session's second occurrence. No usage object, so the per-request sum is a lower bound, exactly as note (bfe) says. |
| 2026-08-01 | S077 | E-20260801 C1 raters, seat P1 — openai/gpt-5.6-terra, 3 batches (one re-dispatched after a second note-(bgc) body) | 0.18 worst case for the seat from max_tokens 8,000 | 0.131503122 | per-response usage.cost, re-summed from the .raw bytes by analysis/verify.py | Provider OpenAI on all three, all stop, 663 / 664 / 814 characters — the terse seat, and the only one that needed no cap change. $0.111233650 of this is the three accepted primaries; the remaining $0.020269472 is the declared reserve deepseek/deepseek-v4-pro, which answered batch 2 after the note-(bgc) body with finish_reason: length and ZERO characters — note (b), and the primary was then re-dispatched and answered normally. |
| 2026-08-01 | S077 | E-20260801 C1 raters, seat P3 — x-ai/grok-4.5, 3 batches (batch 1 re-dispatched to the primary after the reserve had answered it) | 0.27 worst case from max_tokens 12,000 | 0.114550800 | per-response usage.cost | Provider xAI throughout, all stop. One batch returned 1,331 bytes of pure whitespace — note (bgc), the same slug and the same shape S074 met — and the declared reserve deepseek/deepseek-v4-pro answered it for $0.002637600, which is included in the figure at left ($0.111913200 of accepted primaries plus the reserve). That would make the batch not a seat under the design's F5, so the primary was re-dispatched and answered normally; the reserve's answer is preserved unused under runs/void-p3-b1-reserve/ and no figure is computed from it. |
| 2026-08-01 | S077 | E-20260801 C1 raters, seat P2, RUN 2 — google/gemini-3.6-flash at max_tokens 20,000 with effort: low, 3 batches | 0.24 worst case, declared in amendment A5 before the re-dispatch | 0.053184000 | per-response usage.cost | Providers Google and Google AI Studio, all three stop, all three first call. A third of the cost of the failed run at a cap two and a half times larger — note (bdl)'s shape, change the parameter rather than the seat. |
| 2026-08-01 | S077 | E-20260801 C1 raters, seat P2, RUN 1 — WASTED, all three batches | (not separately estimated; it was the declared run) | 0.186061500 | per-response usage.cost, preserved under runs/void-p2-cap8000/ | Note (b), TWENTY-THIRD firing, and the first time a whole SEAT was lost to it. finish_reason: length on all three batches, 7,996 completion tokens each, returning a truncated prose thinking-summary rather than answers. 32% of the session's spend for nothing. It also exposed a defect in this session's own runner: the note-(bfb) fallback accepted the reasoning field because it had enough lines, and a line count is not a format check — new note (bgp), fixed as amendment A4 before the re-dispatch. |
| 2026-08-01 | S077 | E-20260801 C2 byte-identical repeat (P1, batch 1) and C3 prose-only leak null (P2, 44 items) | 0.10 worst case | 0.042267900 | per-response usage.cost | Both stop. The repeat holds at 0.8947 cell-level (registered P3 ≥ 0.85). The null then beat the panel, 0.5294 to 0.4118 on the blind pool, which is the session's governing result and is not what its failure criterion was written to catch — new note (bgn). |
| 2026-08-01 | S077 | the translation limb, both contamination gates, the item build and every analysis | 0.00 | 0.000000 | — | No API call. Multatuli «Max Havelaar» ch. I ¶1–4, 716 Dutch words into 796 English, the project's first Dutch and sixteenth source language, under R07 v1.0 with 78 logged sites, live options enumerated exhaustively before any code was assigned and both coverage tests applied at the moment of each decision; tools/dependence_check.py and tools/build_index.py used unmodified; parse_logs.py, build_items.py, prompts.py, run.py, analyse.py and the 130-check analysis/verify.py with six mutation tests, six caught, two of them cases the verifier must fail on a corrupted stored body, are lead work or local computation. Nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4). |
S077's declared worst case was $1.10 and it was NOT raised. Actual $0.580589442 = 53%, inside the 15–61% band this ledger has been running — and the notable thing is that it held despite $0.208968572 (36%) of work that produced no accepted cell. The estimate was built from max_tokens per note (abc); amendment A5 raised P2's cap from 8,000 to 20,000 and the seat then cost less than it had at the lower cap, which is note (b)'s remedy paying for itself for the first time in this ledger.
THE KEY-USAGE CROSS-CHECK IS SHORT, WITH THE USUAL CAUSE. Opening snapshot 38.114613848 (taken immediately before the critic dispatch; usage_daily 0.0110333, so $0.011 of the UTC day was already spent by something outside this session), closing 38.634664990, delta 0.520051142; per-request sum over the seventeen billed bodies 0.580589442. Shortfall $0.060538300 — note (bco)'s settling lag in the direction it usually runs. Two dispatches returned no usage object at all (note (bgc)'s non-JSON bodies), so the per-request sum is a lower bound in the sense note (bfe) established; it is nonetheless the ledgered figure, per CLAUDE.md, and note (bet)'s session-level bound is the only one asserted. The opening snapshot was overwritten once when the crashed first stage was re-run — both reads returned the same figure to the digit, so nothing was lost, and the runner defect is recorded as note (bgq).
2026-08-01 day total: $0.580589442 of $5.00 — one session (S077). $4.419410558 headroom. 12% of the day's cap. (The key's own usage_daily read $0.0110333 before this session's first dispatch and $0.531084442 at close; the $0.011 opening figure is not attributable to this project by anything this session can see, and is recorded rather than claimed.)
| 2026-08-01 | S078 | E-20260801b-census-author pre-run critic — qwen/qwen3.7-max; declared reserve moonshotai/kimi-k3 not used | 0.20 worst case from max_tokens 12,000 (note (abc)) | 0.071769075 | per-response usage.cost; opening key snapshot runs/snap/critic-open.json 38.69520329 | Provider Alibaba, stop, 222.3 s, in 11,964 / out 12,231 of which 11,055 reasoning. Probed-but-not-selected: not a rater and not the raters' declared reserve, so the S053 role-collision fix holds — eighteenth session running. Verdict NEEDS-REDESIGN, five findings, two BLOCKING, all five accepted, one prescribed remedy DECLINED with a written reason. Note (rr), thirty-first consecutive session. 36% of the declared worst case. Note (bgo) applied literally — the prompt opened with the control block — and it returned twice over: item-specific control verdicts instead of a blanket endorsement, and an ADVISORY finding that predicted the control failure which then happened at 3 of 3 seats. Its BLOCKING finding 2 rebuilt the payloads as a four-way rotation before any rater call and raised the declared worst case from $1.14 to $1.68, recorded rather than absorbed. |
| 2026-08-01 | S078 | E-20260801b payloads R1–R4, the author-arm rotation — 3 seats each, 12 accepted + 2 wasted | 1.10 worst case for twelve seats from max_tokens 6,000, P3 priced at the worst plausible provider (S022 routing caution) | 0.852516100, of which 0.101253000 wasted | per-response usage.cost, re-summed from the .raw bytes by analysis/verify.py | P1 OpenAI throughout, P3 xAI throughout, all stop, all 30 of 30 lines. R4/P2 cost three dispatches: finish_reason: length twice at 5,996 completion tokens (note (b), twenty-fourth and twenty-fifth firings) before a third returned 30 lines. 78% of the twelve-seat worst case, and the overrun driver is P3, which alone accounts for $0.399 of the $0.853. |
| 2026-08-01 | S078 | E-20260801b payload B, the elicitation-mode arms — 3 seats, 3 accepted + 2 wasted | 0.28 worst case for three seats from max_tokens 6,000; P2 re-dispatched at 20,000 under amendment A6, declared before the call | 0.364954400, of which 0.106525500 wasted | per-response usage.cost | P1 OpenAI and P3 xAI first call, both stop, both 46 of 46. B/P2 is the session's instrument finding: length at 5,996 tokens, then a retry that reported stop while ending mid-item-id at the same 5,996 tokens — new note (bgt). Caught by the registered F4 line count, not by finish_reason; both bodies preserved unread under runs/void-b-p2-cap6000/ and the verifier asserts the analysed body is not one of them. The re-dispatch at the raised cap cost $0.138051 against the failed call's $0.053262 — 2.6× MORE, which is the opposite of S077's result and is why note (b)'s remedy paying for itself is not a rule. |
| 2026-08-01 | S078 | E-20260801b the byte-identical repeat — R1 re-issued to P1 unchanged | 0.10 worst case | 0.034111000 | per-response usage.cost | Provider OpenAI, stop, 30 of 30. The cheapest call of the session and the one that decided how its numbers may be read: it moved k by 0.333–0.667 in every arm, all in one direction, on a request that did not change. Note (bfz), second affirmative application. |
| 2026-08-01 | S078 | the translation limb, all three contamination gates, the item build and every analysis | 0.00 | 0.000000 | — | No API call. Szymański «Stolarz Kowalski» opening, 759 Polish words into 1,045 English, two adjacent passages under two elicitation modes with the order fixed in an instrument frozen before the span was chosen; R06 drafts frozen in their own commits first. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; mode-rule.md, build_items.py, call.py, run_critic.py, run.py, analyse.py and the 458-check analysis/verify.py with six mutation tests, six caught are lead work or local computation. Nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4). |
S078's declared worst case was $1.14, RAISED IN SESSION to $1.68 with the reason written before any rater call (amendment A2: the pre-run critic's BLOCKING finding 2 doubled the author-arm dispatches from six to twelve). Actual $1.323350575 = 79% of the amended figure and 116% of the original — the highest fraction in this ledger's history, and the amended estimate is the only reason it is not an overrun. $0.207778500 (15.7%) produced no accepted cell, all of it one seat truncating.
THE KEY-USAGE CROSS-CHECK IS EXACT. Opening snapshot 38.69520329 (taken immediately before the critic dispatch), closing 40.018553865, delta 1.323350575; per-request sum over the twenty-one billed bodies 1.323350575. Residual −0.000000000 — the first exact-to-1e-9 session-level reconciliation since S073, and every dispatch returned a usage object, so the sum is not a lower bound in note (bfe)'s sense. The key's own usage_daily reads 1.914973317 against a two-session per-request total of 1.903940017; the $0.011033300 difference is exactly the figure S077 recorded as already spent by something outside this project when it opened, and is recorded rather than claimed.
2026-08-01 day total: $1.903940017 of $5.00 — two sessions (S077 $0.580589442, S078 $1.323350575). $3.096059983 headroom. 38% of the day's cap.
| 2026-08-01 | S079 | the CTRL-NEG gate — 3 adjudicating seats, openai/gpt-5.6-terra / google/gemini-3.6-flash / x-ai/grok-4.5, reserve deepseek/deepseek-v4-pro | 0.102 worst case from max_tokens 3,000 (note (abc)); registered in gate/registration.md §7 before dispatch | 0.105042406, of which 0.087967971 wasted | per-response usage.cost; opening key snapshot gate/runs/snap/session-open.json 40.348182665 | P1 stop first call ($0.006756750); P3 stop first call ($0.006690400). Seat 2 cost six dispatches: gemini-3.6-flash length three times at cap 3,000 ($0.0257 each) and deepseek-v4-pro length twice ($0.0055, $0.0054) before a third deepseek dispatch returned ($0.003627285) — note (b), and the reserve answering on its third try to an identical payload at temperature: 0. A 3% OVERRUN, and the cause is not the token cap: the estimate was one dispatch per seat and the caller can dispatch 3 attempts × 2 slugs. Note (abc) amended; attempts capped at 2 for the session's main design. Verdict 2 of 3 for the stated criterion, and all three seats read the design text the same way before seeing the code. |
| 2026-08-01 | S079 | E-20260801c-anchor-instance pre-run critic — qwen/qwen3.7-max (probed-but-not-selected) | 0.16 worst case from max_tokens 12,000 | 0.061464725 | per-response usage.cost | Provider Alibaba, stop, 220.6 s, in 5,242 / out 12,143 of which 11,152 reasoning. Not a rater and not the raters' reserve — S053 role-collision fix, nineteenth session running. Verdict NEEDS-AMENDMENT, four findings, two BLOCKING, all four accepted, one with a different remedy and the reason written. Note (rr), twenty-ninth consecutive session. 38% of its declared ceiling. Its BLOCKING finding 1 moved the primary denominator from 19 propositions to 17 and made the registered bar stricter. |
| 2026-08-01 | S079 | E-20260801c stage A, the per-text attestation — 2 arms × 3 seats, 6 accepted + 4 wasted | 0.76 worst case for six seats from max_tokens 8,000, attempts capped at 2 | 0.313193796, of which 0.060159746 wasted | per-response usage.cost, re-summed from the .raw bytes by analysis/verify.py | All six accepted bodies stop, all 22 of 22 lines plus the END terminator. The waste is the session's own runner defect, not a seat's: the F4 item-count guard was carried forward from S078 with that run's item-id pattern hard-coded, so it counted zero answer lines in every body and rejected four dispatches — AP_P1 $0.028725500, AP_P1_gpt-5.6-terra_try2 $0.011105200, AP_P1_deepseek-v4-pro_try1 $0.014202419 (length, empty body — note (b)), AP_P1_deepseek-v4-pro_try2 $0.006126627. New note (bgw). And the bodies really were truncated at stop / native completed with 303 / 219 / 244 content tokens against a cap of 8,000 — note (bgt) at two labs in one stage. The re-dispatch was a single-seat probe, declared in amendments.md A5 before it ran. |
| 2026-08-01 | S079 | E-20260801c stage F, the forcing probe — 3 seats, Swedish alone | 0.18 worst case for three seats from max_tokens 4,000, attempts capped at 2 | 0.077239900, of which 0.031260000 wasted | per-response usage.cost | P1 stop first call, 8.2 s, $0.006031; P3 stop first call, $0.0141204. P2 length at cap 4,000 (note (b)) then stop on the same slug at the same cap. All three returned 3 of 3 paragraphs, so F4 did not fire. 43% of the declared worst case, and it produced the session's largest finding: three independent seats given only the Swedish return longest runs of 17 / 24 / 22 against Stork 1923, against the lead's 24 on the same material. |
| 2026-08-01 | S079 | an in-flight dispatch killed by the operator — google/gemini-3.6-flash, stage A | (not estimated; it did not exist as a planned call) | ≈0.066959998, INFERRED, no per-request record | key-usage delta minus the per-request sum | The process was killed while diagnosing the F4 defect, with a gemini dispatch open; the runner writes the .raw only after the read returns, so the call billed and wrote nothing. The successful AP_P2 at the same cap billed $0.06726 — the two agree to $0.0003 and no other candidate exists. Third case for notes (bfo)/(bfp): S069 an abandoned hang billed at full cap, S070 an abandoned call billed nothing, S079 an operator-killed call billed and left no record. |
| 2026-08-01 | S079 | the translation limb, the contamination gate, the hand read, stage 0 and every analysis | 0.00 | 0.000000 | — | No API call. Söderberg «Pelsen» entire, 1,273 Swedish words into 1,298 English — the project's thirteenth source language and its first Swedish — under R10 against A-mchugh-presence, with a 32-decision log carrying R10 target codes and an opportunity list frozen before it. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; E-20260731f/analysis/checks.py imported unmodified and run against the anchor's second stored text, which no instrument had read before. fetch.py, build_gate.py, build_items.py, run.py, analyse.py and the 287-check analysis/verify.py with six mutation tests, six caught and each asserted to change bytes (note (bgu)) are lead work or local computation. Nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4). |
S079's declared worst case was $1.262 — $0.102 for the gate, registered separately, plus $1.16 for the design. Actual $0.623900826 = 49%, inside the band this ledger runs. The gate itself overran its own $0.102 by 3% and the reason was structural rather than about tokens: an estimate of one dispatch per seat is not a worst case when the caller can dispatch a seat 3 attempts × 2 slugs. Attempts were capped at 2 for the main design in the same session, and the retry structure was written into its estimate.
$0.179387717 (28.8%) produced no accepted cell, and almost all of it is this session's own defects rather than the seats': $0.0880 in the gate's retry loop, $0.0602 in four dispatches rejected by a guard pointed at the wrong item ids, $0.0313 in one genuine length truncation, plus the killed in-flight call.
THE KEY-USAGE CROSS-CHECK IS SHORT BY EXACTLY ONE KILLED DISPATCH. Opening snapshot 40.348182665 (taken immediately before the gate's first dispatch), closing 40.972083491, delta 0.623900826; per-request sum over the fifteen recorded bodies 0.556940828. Gap $0.066959998, attributed above to the operator-killed gemini dispatch on the arithmetic that it matches the same seat's successful call at the same cap to $0.0003. The larger key-delta figure is what is ledgered, per this page's stated method and the S015 precedent.
2026-08-01 day total: $2.527840843 of $5.00 — three sessions (S077 $0.580589442, S078 $1.323350575, S079 $0.623900826). $2.472159157 headroom. 51% of the day's cap.
| 2026-08-01 | S080 | E-20260801d-name-ground-truth pre-run critic — qwen/qwen3.7-max; declared reserve moonshotai/kimi-k3 not used | 0.16 worst case from max_tokens 12,000 (note (abc)) | 0.050377150 | per-response usage.cost; opening key snapshot runs/snap/session-open.json 40.979552091 | Provider Alibaba, stop, 199.7 s. Probed-but-not-selected: not a rater and not the raters' declared reserve — the S053 role-collision fix, twentieth session running. Verdict NEEDS-AMENDMENT, five findings, one BLOCKING, all five accepted, one prescribed remedy DECLINED with a written reason. Note (rr), thirtieth consecutive session. 31% of its declared ceiling. Its BLOCKING finding was note (bgx) found in this session's own design text: the byte-identical repeat's threshold was an absolute 0.90 standing in for a relative claim. Amended before dispatch — and the relative leg then fired, which the absolute one would not have. |
| 2026-08-01 | S080 | E-20260801d the three rater seats, 138 items each, one call per seat | 0.48 worst case for three seats × 2 attempts from max_tokens 6,000, payload 11,288 input tokens measured, P3 priced at twice list for routing (S022 caution) | 0.264432680, of which 0.145144680 wasted | per-response usage.cost, re-summed from the .raw bytes | P1 OpenAI stop first call, 138 of 138, $0.02758925. P3 xAI stop first call, 138 of 138, $0.0493544. P2 cost four dispatches and an amendment: gemini-3.6-flash returned stop with 22 of 138 lines (5,996 completion tokens, 5,757 reasoning) and then stop with 107 of 138 and no END — note (bgt) twice at the cap S078 met it at — and the reserve deepseek-v4-pro came back 137 of 138 with the terminator present, which was NOT accepted because F3 is registered as exactly one line per item and relaxing a criterion after seeing the body it would reject is note (bgx)'s disease. Its second attempt returned a non-JSON body (note (bgc)) that billed nothing. Amendment A6, declared before the call: primary only, no reserve, cap 20,000. It returned 138 of 138 for $0.106467 — less than the two failed attempts it replaced ($0.118), which is S077's result and not S078's. |
| 2026-08-01 | S080 | E-20260801d the byte-identical repeat — P1's payload re-issued unchanged | 0.10 worst case | 0.018173500 | per-response usage.cost | Provider OpenAI, stop, 138 of 138, 27.3 s. The cheapest call of the session and the one that decided how its agreement figures may be read: 0.9638 against a mean pairwise 0.9734, so the relative leg of amendment A1 fires and κ = 0.9435 is descriptive only. P1 agrees with P3 (0.9928) more than it agrees with itself. Note (bfz), third affirmative application. |
| 2026-08-01 | S080 | the translation limb, the contamination gate, the pre-repair sweep, the partial recomputation and every analysis | 0.00 | 0.000000 | — | No API call. Brennu-Njáls saga chapters 1–2 entire, 1,151 Icelandic words into 1,348 English — the project's fourteenth source language and its first medieval Scandinavian source — under R04 v1.0 with the draft committed separately as an R06 output before revision and a 41-row log frozen before the comparator was opened and tallied by a committed parser. tools/dependence_check.py, tools/ngram_overlap.py, tools/build_index.py and tools/check_balance.py used unmodified; the Gutenberg comparator and the Gutenberg #52319 refetch for the sweep are free. fetch.py, fetch_comparator.py, build_gate.py, build_items.py, prompts.py, run.py, analyse.py, sweep/*.py, parse_log.py and the 93-check analysis/verify.py with six mutation tests, six caught and each asserted to change bytes are lead work or local computation. Nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4). |
S080's declared worst case was $0.76, RAISED IN SESSION to $1.10 by amendment A6 with the reason written before the call. Actual $0.397105980 = 36% of the amended figure and 52% of the original — inside the band this ledger runs, and the amendment is the reason it is not reported against a figure the run had already outgrown.
THE KEY-USAGE CROSS-CHECK IS SHORT BY EXACTLY ONE OVERWRITTEN BODY, AND THAT IS HOW THE DEFECT WAS FOUND. Opening snapshot 40.979552091 (taken immediately before the critic dispatch), closing 41.376658071, delta 0.397105980; per-request sum over the seven parseable stored bodies 0.332877480. Residual $0.064228500 — and the first P2 dispatch's usage.cost was $0.0642285, exact to the cent. That body's .raw was overwritten by the A6 re-dispatch, because call.py labels a primary slug's first attempt with the bare tag: note (bdt)'s preserve-before-parse guard defeated by a label collision. New note (bgz); backlog row opened. The larger key-delta figure is what is ledgered, per this page's stated method. Nothing in the analysis depended on the lost body — it was a rejected seat — and the .meta.json survived, which is the only reason the reconciliation is exact rather than merely close.
$0.145144680 (36.6%) produced no accepted cell, all of it the P2 seat: two gemini bodies failing the registered item guard and one reserve body one line short of it.
2026-08-01 day total: $2.924946823 of $5.00 — four sessions (S077 $0.580589442, S078 $1.323350575, S079 $0.623900826, S080 $0.397105980). $2.075053177 headroom. 58% of the day's cap. (The key's own usage_daily reads 3.273077523 at close against a four-session ledgered total of $2.924946823; the $0.348 difference is not attributable to this project by anything this session can see — S077 recorded $0.011 of the same kind at its open — and is recorded rather than claimed.)
| 2026-08-01 | S081 | E-20260801e-lead-carryover pre-run critic — qwen/qwen3.7-max; declared reserve google/gemini-3.6-flash not used | 0.16 worst case from max_tokens 12,000 (note (abc)) | 0.049782725 | per-response usage.cost; opening key snapshot runs/snap/session-open.json 41.386920671 | Provider Alibaba, stop, 172.7 s. Probed-but-not-selected: not a subject and not the subjects' reserve — the S053 role-collision fix, twenty-first session running. Verdict NEEDS-AMENDMENT, five findings, one BLOCKING, all five accepted in substance, two prescribed remedies DECLINED with written reasons. Note (rr), thirty-first consecutive session. 31% of its declared ceiling. Its BLOCKING finding was the normalisation denominator: min(tokens(a), tokens(b)) moves between the two comparisons the whole design rests on, so Carry could have moved with no change in raw overlap. Amended to the fixed text's own 7-gram count before a word was translated, and the remedy went further than prescribed by giving guard L a declared direction. |
| 2026-08-01 | S081 | E-20260801e the Floor rung — 3 passages × 3 seats, fresh contexts, 8 accepted + 3 wasted | 0.41 worst case for nine seats from max_tokens 3,000 (K, C) and 1,600 (O), attempts capped at 2, P3 at twice list and P5 at four times list (S022 routing caution) | 0.138043428, of which 0.013038552 wasted | per-response usage.cost, re-summed from the .raw bytes by analysis/verify.py | P1 OpenAI and P3 xAI stop on every first call, 6 of 6. P5 deepseek-v4-pro cost all three failures: length twice on pair K (providers StreamLake, Baidu) exhausting that seat's two attempts, and once on pair O before Ionstream returned. Note (b), three firings in one run. No amendment was made and no cap was raised: design §9.3 already said a passage with two live seats reports its Floor on two seats, labelled — so the failure was absorbed by a rule written before the data instead of by a repair after it. 34% of the declared worst case. |
| 2026-08-01 | S081 | the three translations, all extraction, the ladder, both controls, the leak sensitivity, the contamination measurement and every analysis | 0.00 | 0.000000 | — | No API call. Kielland «Karen» ¶1–15 (799 Norwegian words), d'Annunzio «La fine di Candia» ¶1–24 (800 Italian words) and Tarchetti «Un osso di morto» opening (357 Italian words) — 1,956 source words rendered blind against archived pairs, each frozen in its own commit before the next was begun, two of them under R10 with their own opportunity lists written before translating and without consulting the earlier session's. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; material/extract.py, prompts.py, run.py, run_critic.py, analyse.py and the 503-check analysis/verify.py with six mutation tests, six caught and each asserted to change bytes on disk are lead work or local computation. Nothing in tools/ changed this session, and one live defect in tools/ngram_overlap.extract() was found, reported and deliberately not repaired. Lead translation is free and is never ledgered (charter §3, A4). |
2026-08-01 day total: $3.112772976 of $5.00 — five sessions (S077 $0.580589442, S078 $1.323350575, S079 $0.623900826, S080 $0.397105980, S081 $0.187826153). $1.887227024 headroom. 62% of the day's cap. (Key-usage cross-check: opening snapshot 41.386920671, closing 41.544405927, delta $0.157485256 against a per-request sum of $0.187826153. The delta is SMALLER than the sum — the ordinary settling lag note (bcx) names, and the fourth session in which it has run in that direction. Per-request sums are ledgered, per this page's stated method.)
| 2026-08-01 | S083 | E-20260801f-tierD-run pre-run critic — qwen/qwen3.7-max; declared reserve moonshotai/kimi-k3 not used | 0.16 worst case from max_tokens 12,000 (note (abc)) | 0.068078625 | per-response usage.cost; opening key snapshot runs/snap/session-open.json 41.583192123 | Provider Alibaba, stop, 252.1 s, 11,748 in / 11,469 out of which 10,202 reasoning. Probed-but-not-selected: not a juror (P1/P2/P5) and not a juror's reserve — the S053 role-collision fix, twenty-second session running. Verdict NEEDS-AMENDMENT, seven findings, two BLOCKING, all seven accepted, two prescribed remedies declined with written reasons. Note (rr), thirty-second consecutive session. 43% of its declared ceiling. Its blocking pair killed two sentences the design would otherwise have published: that the repaired sham lets the run catch a length response, and a pre-registered reading calling Garnett and Hapgood "same-quality". |
| 2026-08-01 | S083 | the control-arm-spec ratification vote — P3 x-ai/grok-4.5, the standing gate on that page, routed before dispatch | 0.05 worst case from max_tokens 6,000 | 0.024095600 | per-response usage.cost | Provider xAI, stop, 73.1 s. A panel member (so a non-Anthropic panel vote, charter §8) and deliberately not one of this design's jurors. Verdict RATIFY-WITH-AMENDMENT, A1 applied verbatim; the S051 backlog obligation is discharged. And it returned the citing design's own reading of R1–R4 INCORRECT, which added §6.8 before anything was dispatched. 48% of its ceiling, and the cheapest correction of the session. |
| 2026-08-01 | S083 | E-20260801f stage 1, the prior positive control — 1 item × 2 orderings × 3 jurors | 0.2896 full-stage reservation (§10), reserved against $1.7950 headroom before the first call | 0.046363, none wasted | per-response usage.cost, re-summed by analysis/verify.py | Caps P1=3500, P2=6000, P5=10000, sized from S034's measured per-call maxima (886 / 2,839 / 5,893). One transport failure (Connection reset by peer) retried and billed nothing; 6 of 6 accepted, 0 failures. 16% of the reservation. The gate fired (2 of 3 units at +1, none −1), so stage 2 was entered. |
| 2026-08-01 | S083 | E-20260801f stage 2, the rebuilt sham — 5 items × 2 orderings × 3 jurors | 1.3470 full-stage reservation, reserved against $1.7487 headroom | 0.252122, none wasted | per-response usage.cost | 30 of 30 accepted, 0 failures, 493 s. Providers routed across eight (OpenAI ×12, Google ×8, Google AI Studio ×4, DeepInfra ×5, GMICloud ×2, Ionstream ×2, Novita ×2, Cloudflare ×1) — note (x). 19% of the reservation, and the per-call mean $0.0084 against S034's $0.0121 on the same task, most of it P1's halved price (S061). Result: IN BAND. |
| 2026-08-01 | S083 | the translation limb, the contamination gate, the extract() audit, the materials build, the Metric A pre-check and all analysis | 0.00 | 0.000000 | — | No API call. Korolenko «Лес шумит» opening, 276 Russian words into 396 English under R04 v1.0, draft frozen separately as an R06 output in its own commit first, with a decision log frozen before any evaluation was designed. tools/dependence_check.py, tools/ngram_overlap.py, tools/metric_a.py, tools/run_tierD.py, tools/build_index.py and tools/check_balance.py used unmodified; materials/build.py, analysis/rules.py, analysis/score.py, analysis/metric_a_precheck.py and the 723-check analysis/verify.py with six mutation tests, six caught and each asserted to change bytes on disk (note (bgu)) are lead work or local computation. The experiment's local call.py carries a one-line fix for note (bgz)'s label collision; nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4). |
S083's declared worst case for what it actually dispatched was $1.847 ($0.16 critic + $0.05 vote + $0.2896 + $1.3470). Actual $0.390659225 = 21%. The two later stages of the design were not entered: stages 3–5 reserve $2.795 and the day had $1.497 left, so the run stopped at the stage boundary its dispatch order exists to make survivable, with the deferral written into NEXT.md. Nothing was wasted: 36 of 36 dispatched calls were accepted, and no body failed a guard.
THE KEY-USAGE CROSS-CHECK IS EXACT TO 1.55 × 10⁻⁷. Opening snapshot 41.583192123 (taken immediately before the critic dispatch), closing 41.973851503, delta 0.390659380; per-request sum over the 36 juror bodies plus the critic and the vote 0.390659225. The residual is float noise, not a lost body — the first session since S077 whose two accounting paths agree to better than a cent, and it is worth naming why: every attempt in this session was uniquely labelled, which is note (bgz)'s repair applied at the point of use.
2026-08-01 day total: $3.503432201 of $5.00 — seven sessions (S077 $0.580589442, S078 $1.323350575, S079 $0.623900826, S080 $0.397105980, S081 $0.187826153, S082 $0.00, S083 $0.390659225). $1.496567799 headroom. 70% of the day's cap.
| 2026-08-01 | S084 | E-20260801g-purpose-row pre-run critic — qwen/qwen3.7-max; declared reserve moonshotai/kimi-k3 not used | 0.16 worst case from max_tokens 12,000 (note (abc)) | 0.035414750 | per-response usage.cost; opening key snapshot runs/snap/session-open.json 41.986007903 | Provider Alibaba, stop, 147.8 s, 5,638 in / 6,124 out of which 4,554 reasoning. Probed-but-not-selected, and the S053 role-collision fix is satisfied trivially rather than by design: this experiment dispatches exactly one call and has no raters at all, which is stated rather than claimed as a control — twenty-third session the fix has held. Verdict NEEDS-REDESIGN (not NEEDS-AMENDMENT), seven findings across six headings, four BLOCKING, all accepted, one prescribed remedy declined with a written reason, two accepted in a stronger form than prescribed. Note (rr), thirty-third consecutive session. 22% of its declared ceiling — the cheapest critic of the eight. What it bought, before a word was translated: the primary's confirming licence cut to nothing (an elastic typology makes residue = 0 nearly free), a co-primary added, all three coding controls replaced because they tested meta-decisions where the primary asks about rendering decisions, a frozen purpose clause rewritten (it was naturalness restated as a licence and would have manufactured a trivial trade at every site), the controls' "blind" withdrawn rather than weakened, and one interpretive claim struck. |
| 2026-08-01 | S084 | the three translations, both gates, all coding and every analysis | 0.00 | 0.000000 | — | No API call. Lagerlöf «En julgäst» ¶1–3 (216 Swedish words), «De fågelfrie» ¶42 (218) and the matched R11 pair on «Gudsfreden» ¶5–8 (276 words rendered twice, 393 + 301 English tokens) — 926 Swedish words, the project's second woman author, each frozen in its own commit before the next was begun and both pair arms frozen before any coding. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; material/fetch.py, material/gate.py, material/spec_check.py, analysis/controls.py, parse_logs.py, divergence.py, pair_sites.py, score_controls.py, score_pair.py and the 86-check analysis/verify.py with three mutation tests, three caught and each asserted to change bytes on disk (note (bgu)) are lead work or local computation. call.py was copied unchanged from E-20260801f. Nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4). |
S084's declared worst case was $0.16, for the one call it made. Actual $0.035414750 = 22%.
THE KEY-USAGE CROSS-CHECK IS EXACT. Opening snapshot 41.986007903 (taken immediately before the critic dispatch), closing 42.021422653, delta 0.035414750; the single recorded body's usage.cost is 0.035414750. Zero residual — the second consecutive session whose two accounting paths agree, and for the same reason as S083's: every attempt uniquely labelled, note (bgz)'s repair applied at the point of use. Nothing was wasted: one dispatch, one acceptance.
Two frozen discard rules fired and cost nothing but work. The design's V3 contamination gate fired on two successive candidate works (17 and 16 contiguous tokens against a 1899 comparator, against six null cells at 3–4), and both were discarded before the studied pair was built — the order note (bcd) prescribes, working as a selection gate rather than as a flag.
2026-08-01 day total: $3.538846951 of $5.00 — eight sessions (S077 $0.580589442, S078 $1.323350575, S079 $0.623900826, S080 $0.397105980, S081 $0.187826153, S082 $0.00, S083 $0.390659225, S084 $0.035414750). $1.461153049 headroom. 71% of the day's cap. ARM-tierD-run stages 3–5 reserve $2.795 and still do not fit; they were deferred a second time, on arithmetic rather than on judgment.
| 2026-08-01 | S085 | D-20260801-10 ratification — independent adversarial review (openai/gpt-5.6-terra, P1) + ratifying vote (google/gemini-3.6-flash, P2) | 0.148 (worst case from max_tokens) | 0.0354295 | per-response usage.cost | review 0.0147865 (in 7,625 / out 876, provider OpenAI) + vote 0.020643 (in 8,832 / out 986, provider Google). Both finish_reason: stop; max_tokens 8,000 with reasoning: {effort: low}, i.e. notes (b) and (abc) applied together rather than one of them. Verdict B-AMENDED, unanimous, against the opening session's default of no change. Landed at 24% of the worst case. |
| 2026-08-01 | S085 | D-20260801-11 ratification — independent adversarial review (deepseek/deepseek-v4-pro, P5) + ratifying vote (x-ai/grok-4.5, P3) | 0.089 (worst case from max_tokens) | 0.0173426234 | per-response usage.cost | review 0.0052962234 (in 3,163 / out 2,542, provider Baidu — list price honoured, no Venice-style excursion on the seat that produced one at S022) + vote 0.0120464 (in 3,975 / out 719, provider xAI). Both stop. Verdict C, unanimous, retiring a sense the list has carried since run one. Landed at 19% of the worst case. |
| 2026-08-01 | S085 | the whole principal unit — the second long work chosen, span 1 translated, the study limb | 0.00 | 0.000000 | — | No API call. The three-route contamination search, the 2,699-word Finnish→English span with its 23-decision frozen log, the ten-rule binding register, the 573-paragraph source preparation, and the 23-locus comparison against Hertzberg's 1886 Swedish are all lead work. Lead translation is free and is never ledgered (charter §3, A4). |
2026-08-01 day total: $3.591619074 of $5.00 — nine sessions (S077 $0.580589442, S078 $1.323350575, S079 $0.623900826, S080 $0.397105980, S081 $0.187826153, S082 $0.00, S083 $0.390659225, S084 $0.035414750, S085 $0.0527721234). $1.408380926 headroom. 72% of the day's cap, and the first day in this ledger's history to carry nine sessions. (The key's own usage_daily reads 4.158489928 at close against the nine-session ledgered total; the $0.567 difference is not attributable to this project by anything this session can see — S077 and S080 both recorded the same shape on this day — and is recorded rather than claimed.) ARM-tierD-run stages 3–5 reserve $2.795 and still do not fit; deferred a third time, on arithmetic rather than on judgment.
| 2026-08-02 | S086 | E-20260801f-tierD-run stage 3, the held-out arm — 2 items × 2 orderings × 3 jurors | 0.579 full-stage reservation (§10), reserved against $5.0000 headroom on a fresh UTC day | 0.161661344 | per-response usage.cost, re-summed from the stored bodies by analysis/verify.py; opening key snapshot runs/snap/session-open-S086.json 42.262070476, usage_daily 0 | Caps P1=3500, P2=6000, P5=10000, unchanged from S083. 12 of 12 accepted, 0 failures, 513 s. 28% of the reservation. The arm did not fire (3 to Garnett, 0 to Hapgood, 3 split) — the pre-registered outcome, and the one that leaves a detection claim alive. One call breached its declared per-call worst case: H-HA__o0__P5 on SiliconFlow, 12,237 completion tokens against a max_tokens of 10,000, finish_reason: stop, $0.04140828 against $0.038775 — note (bgk), second firing, first on this vendor. No budget consequence; the stage reservation absorbed it eight times over. |
| 2026-08-02 | S086 | E-20260801f stage 4, targeted heavy — 4 items × 2 orderings × 3 jurors | 1.108 full-stage reservation, reserved against $4.8383 headroom | 0.189338748 | per-response usage.cost | 24 of 24 accepted, 0 failures, 370 s. 17% of the reservation. Detection fires at ceiling: 12 of 12 units at +1, none −1, all three jurors 4 of 4 taken separately. Specificity does not fire — drop(naturalness) +1.12 against §6.6's ≤ 0.75 — and that one number is the whole verdict. |
| 2026-08-02 | S086 | E-20260801f stage 5, targeted light (drawn 3 sites) — 4 items × 2 orderings × 3 jurors | 1.108 full-stage reservation, reserved against $4.6490 headroom | 0.162968897 | per-response usage.cost | 24 of 24 accepted, 0 failures, 291 s. 15% of the reservation, and the cheapest of the three stages. Detection fires at ceiling and all three specificity conditions fire (drop(naturalness) +0.50). Not claimed as a pass: the light cell does not gate, and making it the primary after watching it pass is the move the discipline exists to stop (RS-20260802-tierD-verdict §3). |
| 2026-08-02 | S086 | the verdict page, the four sensitivity analyses, the verifier extension and all scoring | 0.00 | 0.000000 | — | No API call. analysis/stages345.py (new), the stage-3–5 half of analysis/verify.py (1,664 checks, 0 failures, nine mutation tests, nine caught and each asserted to change bytes on disk, note (bgu)), the P5-excluded sensitivity, prediction 8, the F3 within-stratum check and the five-sense recomputation after literary-quality's retirement are all lead work or local computation. tools/run_tierD.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. |
S086's declared worst case for what it dispatched was $2.795 (the three full-stage reservations). Actual $0.513968989 = 18%. The whole 96-call experiment, across S083 and S086, cost $0.812454145 against a design worst case of $4.432 — 18%, and its central estimate of $1.16 — 70%. Nothing was wasted: 48 of 48 dispatched calls were accepted, no retry was needed, no body failed a guard, and no .attemptN.json or .err file exists.
THE KEY-USAGE CROSS-CHECK LEAVES $0.025883017 UNACCOUNTED, AND NO PROJECT BODY IS MISSING. Opening snapshot 42.262070476 (taken immediately before the first dispatch, usage_daily 0 — a fresh UTC day with the full cap intact), closing 42.801922482, delta 0.539852006; the key's own usage_daily reads 0.539852006 at close, identical to the delta, so the whole of today's key spend is that figure. The per-request sum over the 48 stored bodies is 0.513968989, leaving $0.025883017. Unlike S080's residual this one has no in-session explanation: 48 dispatches, 48 stored bodies, zero discarded attempts, zero error files — checked by listing rather than assumed. Concurrent non-project key use is the standing attribution at this order of magnitude (S042, S045, S049) and settling lag from S085's 2026-08-01 spend is the other; both are recorded rather than one being asserted. The per-request sum is what is ledgered, per this page's stated method, which is also the more conservative direction for the project's own accounting.
2026-08-02 day total: $0.513968989 of $5.00 — one session (S086 $0.513968989). $4.486031011 headroom. 10% of the day's cap. The stages deferred three times on arithmetic at S083, S084 and S085 fitted on the first fresh day with $2.20 of the reservation to spare, exactly as RS-20260801f §8 predicted.
| 2026-08-02 | S087 | E-20260802-voice-warrant pre-run critic — deepseek/deepseek-v4-pro (P5), a non-rater seat | 0.079 worst case from max_tokens 8,000 (note (abc)) → raised to 0.158 in session, amendment A1 | 0.030478153 | per-response usage.cost; opening key snapshot runs/snap/session-open-S087.json 43.319309166, usage_daily 1.05723869 | Two dispatches, and the first returned nothing. At max_tokens 8,000 the seat spent 7,999 of 8,000 completion tokens on reasoning and returned zero characters, finish_reason: length, billing $0.014185606302 — note (abc)'s sharpened form, the cap bounds reasoning as well as visible output, firing on the session's first call. Cap raised to 24,000; second dispatch stop, 5,947 chars, $0.01629254655, provider StreamLake. Verdict NEEDS-AMENDMENT, eight findings, two BLOCKING, all eight accepted, one prescribed remedy accepted with a written modification. Note (rr), thirty-fourth consecutive session. Both BLOCKING findings were about the CONTROL, not the argument, and both were right: the critic prompt contained the design in full, whose §3.2 tabulates the item strata, so the home assignment taken alongside it was discarded and re-taken blind. |
| 2026-08-02 | S087 | E-20260802-voice-warrant blind home assignment — deepseek/deepseek-v4-pro (P5), stateless, items + definitions only | 0.079 | 0.023229 | per-response usage.cost | Provider Parasail, stop, 79.8 s, 468 chars, 24 of 24 lines. Amendment A3. Matched the lead's strata on 9 of 24, against the same seat's 15 of 24 when it had the design in hand — the framing effect the critic predicted, measured. |
| 2026-08-02 | S087 | E-20260802-voice-warrant second home assignment — moonshotai/kimi-k3 (P4) | 0.180 → 0.600 after amendment A4 | 0.5030868 | per-response usage.cost | THREE COMPLETED DISPATCHES, ZERO CHARACTERS OF CONTENT, AND THE MOST EXPENSIVE LINE IN THIS RUN BY A FACTOR OF TWO. 5,997 of 6,000 reasoning tokens ($0.101382), then 5,997 of 6,000 again ($0.0958524), then — after the cap was raised to 20,000 — 19,997 of 20,000 ($0.3058524, 682 s), finish_reason: length every time. A fourth attempt was stopped in flight and billed nothing. The design's amendment A4 had pre-declared the fallback and A5 took it: the agreed subset is two-way. The verdict does not depend on the seat — F3's floor is 16 and P5 supplies only 9. New note (bhf): measure a seat's reasoning appetite before a registered control depends on it; max_tokens is the wrong instrument for measuring it. |
| 2026-08-02 | S087 | E-20260802-voice-warrant stage 1 — WARRANT and EXTENT, 24 items × 2 axes × 3 seats, order counterbalanced | 0.672 (3 seats × 2 calls × 2 attempts, max_tokens 8,000) | 0.156291550 | per-response usage.cost | 6 of 6 accepted on the first attempt, 24 of 24 answer lines each, all stop, 190 s total. P1 openai/gpt-5.6-terra 0.010361 + 0.01884925 (OpenAI); P2 google/gemini-3.6-flash 0.036036 (Google AI Studio) + 0.0434085 (Google); P3 x-ai/grok-4.5 0.0111344 + 0.0365024 (xAI). 23% of the declared worst case. No cap was breached — note (bgk) checked after the fact by the verifier, 10 of 10 bodies within cap. |
| 2026-08-02 | S087 | the translation, both gates, the item set, all scoring and verification | 0.00 | 0.000000 | — | No API call. Heine «Die Harzreise» ¶124 (421 German words → 466 English, 30-decision frozen log) and the ¶118 gate span (203 → 226), the two-comparator seven-cell contamination gate, the 24-item set, analysis/score.py and the 131-check analysis/verify.py with two mutation tests (both asserting byte change, note (bgu); every written file snapshotted and restored, note (bhd)) are all lead work or local computation. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. call.py copied unchanged from E-20260801g. Lead translation is free and is never ledgered (charter §3, A4). |
S087's declared worst case, after amendments A1 and A4, was $1.68. Actual $0.713085503 = 42% — and the fraction is high for the worst reason: 71% of the spend went to a seat that returned no data. Against the original $0.70 declaration it is 102%, i.e. the run overran its first estimate, and note (bhf) is what that bought.
THE KEY-USAGE CROSS-CHECK IS EXACT — AND ONLY BECAUSE TWO DESTROYED BODIES WERE WRITTEN DOWN BEFORE THEY WERE DESTROYED. Opening snapshot 43.319309166, closing 44.032394668, delta 0.713085502. The per-request sum over the 10 surviving stored bodies is 0.597517897; two billed bodies were overwritten when amendments A1 and A4 re-invoked their stages and call.py restarted its attempt counter at 1 ($0.014185606302 and $0.101382, together $0.115567606). 0.597517897 + 0.115567606 = 0.713085503, agreeing with the key delta to 1e-9. Note (bgz) is amended for this: its S080 repair labels every attempt uniquely and does not cover a collision between invocations; workshop/experiments/E-20260802-voice-warrant/runs/OVERWRITTEN.md carries the record.
The $0.543 that is not this project's, and it reconciles too. The key's usage_daily reads 1.770324192 at close against a ledgered day total of 1.227054492. The difference is $0.543269701, and it decomposes exactly: $0.025883017 is S086's own unattributed residual (recorded there) and $0.517386684 was billed to the key between S086's closing snapshot and S087's opening one, by something no session can see. Concurrent non-project key use is the standing attribution at this order of magnitude (S042, S045, S049, S086) and it is recorded rather than asserted. The per-request sum is what is ledgered, per this page's stated method.
2026-08-02 day total: $1.227054492 of $5.00 — two sessions (S086 $0.513968989, S087 $0.713085503). $3.772945508 headroom. 25% of the day's cap.
| 2026-08-02 | S088 | E-20260802b-pole-uniformity pre-run critic — x-ai/grok-4.5 (P3), a seat taking no part in the coding stage | 0.20 worst case from max_tokens 24,000 (notes (abc), (bhf)) | 0.0673904 | per-response usage.cost; opening key snapshot runs/snap/session-open-S088.json 44.231851868 | One dispatch, stop, 7,927 chars, 175.2 s, 6,136 of 8,003 completion tokens on reasoning. Verdict NEEDS-AMENDMENT, nine findings, five BLOCKING, all nine accepted — note (rr), thirty-fifth consecutive session. Every BLOCKING remedy was taken as written; two of them (omission-codes-D, the asymmetric licensing rule) move against the result the lead expected, which is the only defence available to amendments taken after the renderings had been read. |
| 2026-08-02 | S088 | E-20260802b independent coding, seat 1 — openai/gpt-5.6-terra (P1), stateless, site book + six renderings only | 0.15 | 0.023322 | per-response usage.cost | Provider OpenAI, stop, 40.2 s, 349 chars, 21 of 21 answer lines, first attempt. Never shown the design, the hypothesis or the lead's codes. |
| 2026-08-02 | S088 | E-20260802b independent coding, seat 2 — google/gemini-3.6-flash (P2), same payload | 0.15 | 0.0619755 | per-response usage.cost | Provider Google, stop, 35.1 s, 349 chars, 21 of 21 lines, first attempt; 7,433 of 7,649 completion tokens on reasoning for a 349-character answer — the reasoning tax on a mechanical coding task, and the reason this seat cost 2.7× the other for identical output. |
| 2026-08-02 | S088 | D-20260802-12 ratification (the gate, not the principal unit) — independent adversarial review deepseek/deepseek-v4-pro (P5) + routed review vote x-ai/grok-4.5 (P3) | 0.15 | 0.017744733 | per-response usage.cost | review 0.005760183 (provider Ionstream, 45.7 s) + vote 0.0119844 (xAI, 16.6 s). Verdict B, unanimous, both against S087's provisional default of A — the third consecutive ratification to go against a no-change default. Applied to wiki/goodness-senses.md the same session. |
| 2026-08-02 | S088 | the two translations, the contamination gate, the site book, the lead's coding, all scoring and all verification | 0.00 | 0.000000 | — | No API call. Iliad XII.310–327 rendered twice from the Greek alone (R07 184 words, R08 165, 20-decision logs each), the XII.331–341 gate span and its two-comparator dependence_check (0 shared 7-grams, longest runs 6 and 5 against the published pair's own 4), the 21-site book, analysis/score.py and the 181-check analysis/verify.py with three mutation tests (all caught, each asserting the bytes on disk changed, note (bgu)) are all lead work or local computation. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. call.py copied unchanged from E-20260802-voice-warrant. Lead translation is free and is never ledgered (charter §3, A4). |
S088's declared worst case was $0.50 (design §10) and the ratification gate was budgeted separately at $0.15. Actual $0.170432483 — 26% of the $0.65 reserved. Five dispatches, five accepted on the first attempt, zero retries, zero discarded attempts, zero .err files, and no call breached its declared cap (note (bgk) checked after the fact).
THE KEY-USAGE CROSS-CHECK IS EXACT. Opening snapshot 44.231851868, closing 44.402284351, delta 0.170432483, against a per-request sum over the five stored bodies of 0.170432483 — agreement to 1e-9, on five calls across four providers, and the second exact cross-check in three sessions. Both snapshots written to disk before being read (note (bco)); no body was overwritten, so the S087 collision (note (bgz) as amended) did not recur.
2026-08-02 day total: $1.397486975 of $5.00 — three sessions (S086 $0.513968989, S087 $0.713085503, S088 $0.170432483). $3.602513025 headroom. 28% of the day's cap.
| 2026-08-02 | S089 | E-20260802c-regime-scoring pre-run critic (P3 x-ai/grok-4.5, 1 call) | 0.13 | 0.062304400 | per-response usage.cost | raw: workshop/experiments/E-20260802c-regime-scoring/runs/. Overran the design's own stated $0.08 line and the overrun is recorded rather than the line re-based. Verdict NEEDS-AMENDMENT, 10 findings, all accepted before dispatch |
| 2026-08-02 | S089 | E-20260802c-regime-scoring scoring block: 10 items × 2 orderings × 3 jurors | 2.696 (full-stage reservation) | 0.515743035 | per-response usage.cost, summed | raw: .../runs/scores/. 60 of 60 accepted first attempt, 0 parse failures, 0 transport errors. Realised at 19% of the reservation, the same proportion as the S086 Tier D stage |
2026-08-02 day total: $1.975534410 of $5.00 — four sessions (S086 $0.513968989, S087 $0.713085503, S088 $0.170432483, S089 $0.578047435). $3.024465590 headroom. 40% of the day's cap, and the largest single-session spend of the four.
The cross-check is NOT exact and the gap is stated rather than absorbed. Opening 45.164290590, closing 45.896845469, delta $0.732554879 against a per-request sum of $0.578047435 — $0.154507444 unattributed. Per this page's stated method the per-request sum is primary; the residual is the same order and shape as the $0.543269701 of non-project drift this same UTC day already carries on S087's row, and it is recorded as unattributed rather than assigned.
| 2026-08-02 | S090 | E-20260802d-deference-carry pre-run critic (P5 deepseek/deepseek-v4-pro, 1 accepted body) | 0.15 worst case from max_tokens 24,000 (notes (abc), (bhf)) | 0.014593728 | per-response usage.cost | raw: workshop/experiments/E-20260802d-deference-carry/runs/. Provider DigitalOcean, stop, 10,820 chars, 38.6 s, 0 reasoning tokens. Verdict NEEDS-AMENDMENT, five findings, three BLOCKING, all five accepted — note (rr), thirty-sixth consecutive session. Two amendments (A2 kin-term code, A3 unframed arm) made the design able to refute the lead's own register rule, and one of them did |
| 2026-08-02 | S090 | E-20260802d stage 1 — contrast-subject translations, 2 sites × 3 seats + 1 unframed arm | 0.52 (9 calls × 2 attempts, max_tokens 4,000) | 0.094170700 | per-response usage.cost, re-summed from the stored bodies by analysis/verify.py | 9 of 9 accepted on the first attempt, all stop, 0 failures, 0 retries, 0 .err files, no body above its cap (note (bgk), checked on every body by the verifier). Routing (note (x)): P1 OpenAI ×3, P2 Google ×2 / Google AI Studio ×1, P3 xAI ×3. P2 spent 2,053 / 2,498 / 2,882 completion tokens on 767–1,300-character translations — 4–6× P1's cost for the same passage, the third instance of the note (bhf) shape |
| 2026-08-02 | S090 | span 2 of «Köyhää kansaa», the 13-locus Hertzberg comparison, all coding and all verification | 0.00 | 0.000000 | — | No API call. The 1,782 → 2,758-word span with its 23-decision frozen log, the two errata, the alignment of Finnish ¶77–134 to Swedish ¶74–131, the 13-locus classification and the 61-check analysis/verify.py with three mutation tests (all caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)) are lead work or local computation. tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. call.py copied unchanged from E-20260802b-pole-uniformity. Lead translation is free and is never ledgered (charter §3, A4) |
S090's declared worst case was $0.67 after the critic's amendments raised the subject count 6 → 9. Actual $0.108764428 — 16% of the reservation. Ten dispatches, ten accepted on the first attempt.
AN EXACT CROSS-CHECK, AND THE FIRST THIS LEDGER HAS ON A KILLED DISPATCH. The critic prompt was first sent in the foreground and killed in flight at a 2-minute tool timeout before the run was backgrounded; whether it had billed was open. Session-open snapshot 46.007974247, snapshot after the accepted critic body 46.022567975, delta 0.014593728 — the accepted body's per-request cost to 1e-9. The killed dispatch billed exactly zero, established rather than assumed, which is the question note (bcx) usually has to leave open.
Stage 1's own delta is 0.019216000 (46.022567975 → 46.041783975) against a per-request sum of 0.094170700: $0.074954700 still unsettled at close, the ordinary lag direction and not the (bet) excess direction. Per-request sums are ledgered, per this page's stated method. Between-session drift from S089's close (45.896845469) to this session's open was $0.111128778, consistent with S089's own $0.154507444 residual settling after its closing snapshot.
2026-08-02 day total: $2.084298838 of $5.00 — five sessions (S086 $0.513968989, S087 $0.713085503, S088 $0.170432483, S089 $0.578047435, S090 $0.108764428). $2.915701162 headroom. 42% of the day's cap.
| 2026-08-02 | S091 | E-20260802e-displaced-marking pre-run critic (P5 deepseek/deepseek-v4-pro, 1 body) | 0.20 worst case from max_tokens 24,000 (note (abc)) | 0.009172451 | per-response usage.cost | raw: workshop/experiments/E-20260802e-displaced-marking/runs/. Provider StreamLake, stop. Verdict NEEDS-AMENDMENT, four findings, two BLOCKING, all four accepted — note (rr), thirty-seventh consecutive session. A1 (an independent naturalness audit of the decoy) and A2 (a manual audit of all 126 census windows, and the population claim downgraded) are the two that could hurt the lead |
| 2026-08-02 | S091 | E-20260802e stage 1 — 3 seats × 2 rotations grading 4 renderings × 6 sites | 0.55 | 0.220316800 | per-response usage.cost, re-summed from stored bodies by analysis/verify.py | 6 accepted, 4 rejected. P1 and P3 clean on first attempt. P2 gemini-3.6-flash returned finish_reason: length on all four attempts at max_tokens 4,000 (4,177–5,691 reasoning chars) and recovered at 16,000 under declared amendment A5. The four failed bodies billed $0.133473 |
| 2026-08-02 | S091 | E-20260802e stage 1b — the decoy naturalness audit (P4 moonshotai/kimi-k3), critic amendment A1 | 0.25 | 0.298161000 | per-response usage.cost | 1 accepted, 1 rejected. The rejected body returned length with 16,803 reasoning characters and ZERO content characters for eighteen one-line ratings; it billed $0.062820. The A5 retry at max_tokens 20,000 routed to Fireworks rather than Moonshot AI and cost 3.7× the failed attempt — the S022 routing caution, live. This stage overran its own $0.25 line and the overrun is recorded rather than the line re-based |
| 2026-08-02 | S091 | E-20260802e stage 2 — followability, 3 seats × 20 frozen log excerpts | 0.25 | 0.269987400 | per-response usage.cost | 3 accepted, 2 rejected (both P2, both length, $0.073197, recovered under A5). Result: pairwise κ 0.630 / 0.631 / 0.774, all above the pre-registered 0.60 bar |
| 2026-08-02 | S091 | the census, the six forced re-renderings, the decoys, the positive controls, all classification and all verification | 0.00 | 0.000000 | — | No API call. 126 → 61 → 13 census, the 8 Class A / 8 Class B classification, the four renderings at each of six sites in four source languages, and the 62-check analysis/verify.py with three mutation tests (all caught, each asserting the bytes changed, note (bgu); every mutated file snapshotted and restored, note (bhd)) are lead work or local computation. tools/ unmodified. call.py copied unchanged from E-20260802d. Lead translation is free and is never ledgered (charter §3, A4) |
S091's declared worst case was $1.20, raised to $1.40 by amendment A5. Actual $0.797637651 — 57% of the reservation, the highest fraction of a declared worst case this project has spent, and the reason is on the record rather than in the noise.
$0.269490000 — 34% of the spend — bought nothing. Seven of eighteen bodies returned finish_reason: length and were rejected under note (b); all seven were hidden-reasoning overruns on two seats, and all seven recovered at a raised max_tokens with byte-identical prompts. Note (bhf) has fired three times on price; this is its first firing on capacity, and note (bhq) now records that max_tokens must be sized from a seat's measured reasoning appetite plus the answer, never from the answer alone.
2026-08-02 day total: $2.881936489 of $5.00 — six sessions (S086 $0.513968989, S087 $0.713085503, S088 $0.170432483, S089 $0.578047435, S090 $0.108764428, S091 $0.797637651). $2.118063511 headroom. 58% of the day's cap, and the largest single-session spend of the six.
| 2026-08-02 | S092 | E-20260802f-licensed-strangeness pre-run critic (P2 google/gemini-3.6-flash, a seat taking no part in the rating stage) | 0.19 worst case from max_tokens 24,000 (notes (abc), (bhq)) | 0.079264500 | per-response usage.cost; opening key snapshot runs/snap/session-open-S092.json 48.049987491 | Provider Google, stop, 5,452 chars, 39.3 s, 7,119 of 8,574 completion tokens on reasoning. Verdict NEEDS-AMENDMENT, five findings, three BLOCKING, all five accepted — note (rr), thirty-eighth consecutive session. Two could have voided the run: A2 struck a LICENSED edit that was a forced etymological gloss rather than a feature of the Greek, and A1 computed the design's own matching criterion and found it failing at 38.3% against a 30% ceiling before any datum existed — note (bhr) |
| 2026-08-02 | S092 | E-20260802f condition N (no source) — 3 seats × 13 attribution items + 4 naturalness ratings | 0.38 | 0.060470664 | per-response usage.cost, re-summed from stored bodies by analysis/verify.py | 3 of 3 accepted on the first attempt, all stop, 17 of 17 answer lines each. P1 openai/gpt-5.6-terra (OpenAI) $0.014977; P3 x-ai/grok-4.5 (xAI) $0.0394764; P5 deepseek/deepseek-v4-pro (StreamLake) $0.006017263824 |
| 2026-08-02 | S092 | E-20260802f condition S (the Greek supplied) — same 3 seats, same items, dispatched after every condition-N body had returned | 0.38 | 0.064910571 | per-response usage.cost | 3 of 3 accepted on the first attempt. P1 $0.01575975; P3 $0.0415484; P5 $0.007602421398. Six of the seven bodies in this session spent over 95% of their completion tokens on reasoning for a 143-character answer; no body came within 7,000 tokens of its cap, note (bgk) checked by the verifier on every body |
| 2026-08-02 | S092 | the four arms, the edit sets, the contamination gate, all scoring and all verification | 0.00 | 0.000000 | — | No API call. Three derived renderings of Iliad XII.310–327 with a frozen edit log, the nine-cell dependence_check gate, analysis/score.py and the 81-check analysis/verify.py with two mutation tests (both caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)) are lead work or local computation. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. call.py copied unchanged from E-20260802e. Lead translation is free and is never ledgered (charter §3, A4) |
S092's declared worst case was $1.20. Actual $0.204645735 — 17% of the reservation, and 7 of 7 dispatches were accepted on the first attempt: no retries, no .err files, no finish_reason: length, nothing rejected. That is the direct answer to S091, where 34% of the spend bought nothing; the caps were sized from note (bhq) rather than from the answer length, and the reasoning tokens (2,056–7,119 for a 143-character answer) show the note was measuring the binding quantity.
One exact cross-check and one named residual. Session-open 48.049987491 → the snapshot taken when the rating stage opened, 48.129251991: delta 0.079264500, the critic body's per-request cost to 1e-9. Session-open → final close (48.247030804) is 0.197043313 against a per-request sum of 0.204645735; the residual $0.007602422 equals the last dispatched body's cost to 1e-9, i.e. that call had not settled when the closing snapshot was taken — the ordinary lag direction, note (bcx), and for once fully attributable. Per-request sums are ledgered, per this page's stated method. All snapshots written to disk before being read (note (bco)).
2026-08-02 day total: $3.086582224 of $5.00 — seven sessions (S086 $0.513968989, S087 $0.713085503, S088 $0.170432483, S089 $0.578047435, S090 $0.108764428, S091 $0.797637651, S092 $0.204645735). $1.913417776 headroom. 62% of the day's cap.
| 2026-08-02 | S093 | D-20260802-13 ratification (the gate, not the principal unit) — independent adversarial review, moonshotai/kimi-k3 (P4), first dispatch, REJECTED | 0.25 worst case from max_tokens 14,000 (notes (abc), (bhq)) | 0.221532 | per-response usage.cost; opening key snapshot session-open-S093.json 48.390045095 | finish_reason: length. 14,000 completion tokens, ZERO content characters. Provider Together. This bought nothing and it is 59% of the session's entire spend. Note (b) was written at S044 about this same model doing this same thing, and its prescribed fix — cap the reasoning effort, a generous max_tokens alone is not enough — was not applied on the first dispatch. The note is amended for it: on a seat with a recorded history, the cap is not advisory |
| 2026-08-02 | S093 | same review, retry under note (b)'s own fix (--reasoning-effort low, max_tokens 16,000) | included above | 0.047242200 | per-response usage.cost | Provider Together, stop, 50.9 s, 2,392 completion tokens. Verdict C, eight findings, one BLOCKING — and the BLOCKING one is real: RS-20260802f's naturalness item quotes the sense minus the clause, so the run never varied the clause and "the clause is inert" is withdrawn. The retry cost 21% of the failure it replaced |
| 2026-08-02 | S093 | D-20260802-13 routed vote — openai/gpt-5.6-terra (P1), a blind rating seat in E-20260802f that never saw the design, hypothesis or results | 0.10 | 0.019125500 | per-response usage.cost | Provider OpenAI, stop, 23.3 s. Verdict C, accepts all eight findings, id perceived-source-carriage, six conditions. Applied to wiki/goodness-senses.md the same session |
| 2026-08-02 | S093 | E-20260802g-programme-divergence — 22-site rule-application coding, 2 blind seats | 0.20 worst case from max_tokens 14,000 (note (bhq)) | 0.086262800 | per-response usage.cost, re-summed from the stored bodies by analysis/verify.py | 2 of 2 accepted on the first attempt, both stop, 22 of 22 answer lines each, 0 retries, 0 .err files. P5 deepseek/deepseek-v4-pro (provider Alibaba) $0.0313644, 115.1 s; P3 x-ai/grok-4.5 (provider xAI) $0.0548984, 159.4 s. 43% of the reservation. Both registered controls fired clean |
| 2026-08-02 | S093 | the complete reading, both regimes, both translations, the contamination gate, the site book, the lead's coding, all analysis and all verification | 0.00 | 0.000000 | — | No API call. The 438,198-character Arnold–Newman volume read entire; R12 and R13 written; Iliad XXIV.486–512 rendered twice from the Greek alone with 22- and 20-decision frozen logs; the XXIV.552–570 gate span and its two-comparator dependence_check; the 22-site book; the lead's own 22 codes; and the 32-check analysis/verify.py with two mutation tests (both caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)) are all lead work or local computation. tools/dependence_check.py, tools/ratify_vote.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4) |
S093's declared worst case was $0.55 across the two stages. Actual $0.374162500 — 68% of the reservation, and the highest fraction any session has spent (the previous high was S091's 57%). The reason is one rejected body: $0.221532, 59% of the spend, for zero content characters. Excluding it, the session spent $0.15263 on five accepted bodies against a $0.55 reservation — 28%, the ordinary range. The lesson is not that the cap was too small; it is that note (b) existed, named the model, prescribed the fix, and was not applied.
THE KEY-USAGE CROSS-CHECK IS EXACT, AND FOR THE FIRST TIME THE RESIDUAL IS ZERO. Opening snapshot 48.390045095, closing 48.764207595, delta 0.374162500, against a per-request sum over the five stored bodies of 0.374162500 — agreement to 1e-9 with a residual of exactly 0.000000000, across five calls on four providers. No lag, no drift, nothing unattributed. Both snapshots written to disk before being read (note (bco)); no body was overwritten. Between-session drift from S092's close (48.247030804) to this session's open was $0.143014291, unattributed and consistent with the concurrent non-project key use this ledger has recorded at this order of magnitude since S042.
2026-08-02 day total: $3.460744724 of $5.00 — eight sessions (S086 $0.513968989, S087 $0.713085503, S088 $0.170432483, S089 $0.578047435, S090 $0.108764428, S091 $0.797637651, S092 $0.204645735, S093 $0.374162500). $1.539255276 headroom. 69% of the day's cap.
| 2026-08-03 | S094 | E-20260803-a4-set pre-run critic (P4 moonshotai/kimi-k3, a seat taking no part in any other stage) | 0.19 worst case from max_tokens 16,000 (notes (abc), (bhq)) | 0.070260 | per-response usage.cost; opening key snapshot runs/snap/session-open-S094.json 49.914103845 | Provider Moonshot AI, stop, 133.9 s, 6,136 chars. Note (b)'s prescribed fix — reasoning: {"effort": "low"} — was applied on the FIRST dispatch to the seat the note is about, against S093, where it was not and the first dispatch billed $0.221532 for zero content characters. Verdict NEEDS-AMENDMENT, eleven findings, three BLOCKING, all eleven accepted — note (rr), streak restored after S093's breach. One of the three accepted BLOCKING amendments then introduced note (bhs)'s defect |
| 2026-08-03 | S094 | E-20260803-a4-set edit-neutrality check (P3 x-ai/grok-4.5), critic amendment A8 | 0.04 | 0.010834 | per-response usage.cost | Provider xAI, stop, 19.9 s, 11 of 11 answer lines. Verdict "quality-neutral" with a predicted direction (naturalness up, voice/style-correspondence down) that the run's measured floor agrees with on sign. A check, not a gate: the edit list was frozen and was not revised on it |
| 2026-08-03 | S094 | E-20260803-a4-set log-reading stage — P3, six frozen translator's logs, dispatched in full before the first rating call | 0.22 | 0.095542 | per-response usage.cost | 6 of 6 accepted on the first attempt, all stop, 19.6–38.9 s. The seat saw no source and no translation |
| 2026-08-03 | S094 | E-20260803-a4-set rating stage — 3 jurors × 8 items × 2 passes, one text per call, strictly sequential | 1.80 across both passes | 0.472781 | per-response usage.cost, re-summed from the stored raw bodies by analysis/verify.py | 48 accepted, 1 rejected. rate_PAN_P5_p1 returned finish_reason: length and billed $0.016288625 for nothing — 2.5% of the session, against S093's 59%; it recovered on a byte-identical retry to the same slug. P5 routed across ten providers in 32 calls with no price excursion |
| 2026-08-03 | S094 | the contamination gates, the fresh translation, the null's edit list, all scoring and all verification | 0.00 | 0.000000 | — | No API call. Three gate runs of tools/dependence_check.py (two firing, one re-checking a published figure and reproducing it), the 220-word Daudet and 233-word Sand gate translations, the 503-word T-clos-des-ames-R04-v1 and its ten-edit paraphrase, analysis/score.py and the 640-check analysis/verify.py with three mutation tests (all caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)) are lead work or local computation. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4) |
S094's declared worst case was $2.21 across four stages. Actual $0.649417427 — 29% of the reservation, and 54 of 55 dispatches were accepted on the first attempt.
THE KEY-USAGE CROSS-CHECK IS THE CLOSEST THIS LEDGER HAS RECORDED. Opening snapshot 49.914103845, closing 50.563521268, delta 0.649417423, against a re-sum over all 55 stored raw bodies (including the rejected one) of 0.649417427 — a residual of −4 × 10⁻⁹, which is floating-point and not lag. S093 was the first exact zero on five calls; this is fifty-five calls across five slugs and thirteen providers. Both snapshots written to disk before being read (note (bco)). Between-session drift from S093's close (48.764207595) to this session's open was $1.149896250, unattributed and larger than the usual order of magnitude for the concurrent non-project key use this ledger has recorded since S042 — noted rather than explained, since a whole UTC day elapsed between the two sessions.
2026-08-03 day total after S094: $0.649417427 of $5.00. $4.350582573 headroom.
| date | session | what | pre-flight | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-03 | S095 | E-20260803b-honorific-carry pre-run critic, pass 1 (P3 x-ai/grok-4.5, takes no part in scoring) |
0.16 worst case from max_tokens 24,000 (notes (abc), (bhf)) |
0.0569544 | per-response usage.cost; opening key snapshot runs/snap/session-open-S095.json 50.567528368 |
Provider xAI, stop, 111.8 s, 4,635 of 6,153 completion tokens on reasoning. Verdict NEEDS-REDESIGN, nine findings, five BLOCKING, all nine accepted — note (rr), thirty-eighth consecutive session. Five of the nine were one defect from different angles: no outcome of the design could have changed the translation. The amendment that fixed it (A1) is what produced Erratum 4 |
| 2026-08-03 | S095 | E-20260803b pre-run critic, pass 2 on the amended design (P3) |
0.16 | 0.0759304 | per-response usage.cost |
Provider xAI, stop, 155.7 s, 7,191 of 8,886 completion tokens. Verdict NEEDS-AMENDMENT, nine findings, three BLOCKING, all nine accepted. B6 halved the run: call.py posts temperature: 0, so the design's two "repeats" were the same draw and would have double-counted in the robustness check and the permutation test — 36 calls cut to 18. ⚠ This dispatch overwrote pass 1's stored body (same tag; note (bgz)'s shape, not its case) — recorded at runs/critic-pass1-body.NOTE.md, not silently repaired |
| 2026-08-03 | S095 | E-20260803b scoring stage — 3 arms × 3 seats × 2 orderings |
0.40, plus a 0.40 retry reserve | 0.130995926 | per-response usage.cost, re-summed from the stored raw bodies by analysis/verify.py |
18 of 18 accepted on the first attempt, all stop, 0 failures, 0 retries, 0 .err files. 33% of the reservation. Routing (note (x)): P1 OpenAI ×6, P2 Google / Google AI Studio ×6, P5 ×6. P2 cost 14× P5 for identical 40–55-character answers — the note (bhf) shape again |
| 2026-08-03 | S095 | span 3 of «Köyhää kansaa», the Hertzberg comparison, the 21-site census, all scoring and all verification | 0.00 | 0.000000 | — | No API call. The 1,795 → 2,778-word span with its 18-decision frozen log, the alignment against Hertzberg 1886, the second-person census, analysis/score.py (frozen before dispatch, critic B3) and the 31-check analysis/verify.py with two mutation tests (both caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)) are lead work or local computation. tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4) |
S095's worst case was $1.94 as first frozen and $0.94 after the second critic pass halved the run. Actual $0.263880726 — 28% of the revised reservation, and 20 of 20 dispatches accepted on the first attempt.
The key-usage cross-check closes at a residual of −2 × 10⁻⁹, the closest this ledger has recorded. Opening snapshot 50.567528368, closing 50.831409092, delta 0.263880724, against a re-sum over all 20 stored raw bodies of 0.263880726. Between-session drift from S094's close (50.563521268) to this session's open was $0.004007100 — two orders of magnitude smaller than the gap at S094's own open, and these two sessions ran on the same UTC day.
2026-08-03 day total: $0.913298153 of $5.00 — two sessions (S094 $0.649417427, S095 $0.263880726). $4.086701847 headroom. 18% of the day's cap.
S096 — 2026-08-03 (UTC), E-20260803c-occupied-slot
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-03 | S096 | E-20260803c pre-run critic, pass 1 (P3 x-ai/grok-4.5, grades nothing in this experiment) |
0.20 worst case from max_tokens 24,000 (notes (abc), (bhf)) |
0.0623784 | per-response usage.cost; opening key snapshot runs/snap/session-open-S096.json 51.112336122 |
Provider xAI, stop, 131.4 s, 3,995 of 6,798 completion tokens on reasoning. Verdict NEEDS-REDESIGN, ten findings, seven BLOCKING, all ten accepted — note (rr), thirty-ninth consecutive session. Four of the seven BLOCKING findings were one defect: the design could not distinguish the slot is spent from the translator patched harder where patching was easier. The hand-written repair arm was withdrawn from scoring and replaced by a mechanical one |
| 2026-08-03 | S096 | E-20260803c pre-run critic, pass 2 on the redesigned arm (P3) |
0.20 | 0.0653484 | per-response usage.cost |
Provider xAI, stop, 176.1 s, 4,854 of 6,890 completion tokens. Verdict NEEDS-AMENDMENT, eleven findings, four BLOCKING, all eleven accepted (two in part, with the unaccepted part written out). Caught two registered relations no wording of the excerpt could have conveyed and a token fingerprint in the repair lexicon, both fixed before dispatch. The two passes are the session's most consequential act: without them S096 would have reported a false positive |
| 2026-08-03 | S096 | E-20260803c grading stage — 3 seats x 2 rotations x 2 halves, plus one attention-check call per seat |
1.50 across both stages, incl. a 0.20 retry reserve | 0.551652808 | per-response usage.cost, re-summed from the stored raw bodies by analysis/verify.py |
15 scored bodies of 21 billed. P2 google/gemini-3.6-flash billed $0.458625 — 68% of the session — against P1's $0.046731 and P5's $0.046297 for the same five payloads, and $0.162 of it bought four rejected bodies (finish_reason: length or no terminator at max_tokens 4,000). All four answered on a byte-identical re-dispatch at 12,000, recorded as amendment A14 rather than done silently. Note (bhq)/(bhf) a fourth time, now with a rejection rate attached. P5 routed across five providers in five calls with no price excursion |
| 2026-08-03 | S096 | the contamination gates, the translation, the census, all scoring and all verification | 0.00 | 0.000000 | — | No API call. Two runs of tools/dependence_check.py (the pre-selection gate, fired at 13 tokens, and the post-freeze span check, 15 tokens), the 229-word gate rendering, the 1,981 -> 2,208-word T-kammacher-R04-v1 with its 19-decision frozen log, the twelve-site census, analysis/score.py and the 117-check analysis/verify.py with two mutation tests (both caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)) are lead work or local computation. tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter S3, A4) |
S096's declared worst case was $2.10 across three stages. Actual $0.679379608 - 32% of the reservation, in the ordinary range.
The key-usage cross-check closes at a residual of exactly zero. Opening snapshot 51.112336122, final 51.791715730, delta 0.679379608, against a per-request re-sum over all 25 stored raw bodies of 0.679379608 - agreement to 1e-9, across 25 calls on ten providers, with the four rejected bodies included in both figures. Between-session drift from S095's close (50.831409092) to this session's open was $0.280927030 - an order of magnitude above the S094->S095 gap and consistent with the concurrent non-project key use this ledger has recorded since S042; it is not attributed to this project.
2026-08-03 day total: $1.592677761 of $5.00 - three sessions (S094 $0.649417427, S095 $0.263880726, S096 $0.679379608). $3.407322239 headroom. 32% of the day's cap.
S097 — 2026-08-03 (UTC), E-20260803d-class-uniformity
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-03 | S097 | E-20260803d pre-run critic, dispatch 1 (P4 moonshotai/kimi-k3) |
0.26 worst case from max_tokens 16,000 (note (abc)) |
0.260301 | per-response usage.cost; opening key snapshot runs/snap/session-open-S097.json 51.925936255 |
BOUGHT NOTHING. Provider Fireworks, finish_reason: length, 16,547 reasoning tokens, ZERO content characters, 223.2 s. Guard (b) rejected it correctly and the raw body is preserved. Note (b)'s twenty-fourth firing and the fourth on this seat: its standing amendment — set effort: low on the FIRST dispatch to a seat with this history — was not applied, because call.py was inherited from S096 and had no reasoning parameter at all. Half this session's API spend, for nothing |
| 2026-08-03 | S097 | E-20260803d pre-run critic, pass 1 re-dispatched (P4, reasoning: {"effort":"low"}, new tag, reserve declared) |
0.26 | 0.051360 | per-response usage.cost |
Provider Fireworks, stop, 76.0 s, 5,580 chars — a fifth of the failed dispatch, for a complete verdict. NEEDS-AMENDMENT, eight findings, four BLOCKING, all eight accepted — note (rr), fortieth consecutive session. Finding 1 was fatal to stage 2: the item glosses paraphrased the Russian morphology, so each gloss WAS the calque whose availability was the question — note (bhy) |
| 2026-08-03 | S097 | E-20260803d pre-run critic, pass 2 on the amended design (P4) |
0.26 | 0.117210 | per-response usage.cost |
Provider Fireworks, stop, 106.0 s, 8,613 chars. NEEDS-AMENDMENT, nine findings, three BLOCKING, all nine accepted. Finding 1 saved the run: the registered primary measured departure from the translator's modal strategy, which can only detect cause when the modal strategy IS the calque — and Garnett's modal is substitute, so the run would have reported a wrong-signed false negative about the translator the census is actually about. The primary was rewritten before dispatch |
| 2026-08-03 | S097 | E-20260803d stage 1, strategy coding — 3 seats × 14 items |
0.07, plus a 0.20 retry reserve | 0.037390 | per-response usage.cost, re-summed from the stored raw bodies by analysis/verify.py |
P1 25.4 s and P3 12.7 s accepted first call. P5 deepseek/deepseek-v4-pro, provider StreamLake, failed twice at max_tokens 4,000 WITH effort: low — length, empty body, $0.0064953504 each. Note (b)'s twenty-fifth firing; note (bdl)'s remedy (raise the cap, do not fall through) worked at 12,000, accepted first call in 89.2 s |
| 2026-08-03 | S097 | E-20260803d stage 2, property coding — 3 seats × 22 items, Russian only |
0.05 | 0.054687 | per-response usage.cost, re-summed by analysis/verify.py |
P1 and P3 accepted first call. P5 failed twice again at 4,000, note (b)'s twenty-sixth firing, and again answered first call at 12,000 in 104.1 s. All four registered controls passed and 42 of 42 stage-1 cells took a seat majority |
| 2026-08-03 | S097 | the contamination gate, the 3,825-word translation, the class table, all scoring and all verification | 0.00 | 0.000000 | — | No API call. One run of tools/dependence_check.py (the published pair, DEPENDENT? at 14 tokens, registered before the run rather than discovered after it), the whole night talk of «Бежин луг» with its fourteen-item class record written at translation time, analysis/score.py frozen before dispatch and amended only on critic findings, and the 55-check analysis/verify.py with two mutation tests (both caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)). tools/dependence_check.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4) |
S097's declared worst case was $0.84, revised to $1.10 by amendment A1 after the first critic dispatch failed. Actual $0.520948932 — 47% of the revised reservation — and $0.260301 of it, exactly half, bought nothing. Without note (b)'s firing the session would have cost $0.26.
Five rejected bodies are preserved on disk and are billed in both the per-request re-sum and the totals above, per note (bdt).
The key-usage cross-check does NOT close this session, and the residual is recorded rather than reconciled. Opening snapshot 51.925936255, closing 52.624133459, delta 0.698197204, against a per-request re-sum over all 13 stored raw bodies of 0.520948932 — residual $0.177248272. Between-session drift from S096's close (51.791715730) to this session's open was a further $0.134220525. Both fall in the concurrent non-project key use this ledger has recorded since S042; per-request costs are primary and the delta is the sanity check (CLAUDE.md), so the figures above stand on the stored bodies.
2026-08-03 day total: $2.113626693 of $5.00 — four sessions (S094 $0.649417427, S095 $0.263880726, S096 $0.679379608, S097 $0.520948932). $2.886373307 headroom. 42% of the day's cap.
S098 — 2026-08-03 (UTC), D-20260803-14 ratification gate + E-20260803e-purpose-index
| date | session | item | pre-flight estimate (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-03 | S098 | ratification gate, D-20260803-14 independent adversarial review (P2 google/gemini-3.6-flash, the only panel seat with no role of any kind in E-20260803d) |
0.099 worst case from max_tokens 12,000 (note (abc)) |
0.0230685 | per-response usage.cost; opening key snapshot runs/snap/session-open-S098.json 52.811520842 |
Provider Google, stop, 12.6 s, 5,724 in / 1,931 out. Note (b)'s prescribed effort: low applied on the FIRST dispatch to a seat with four rejected bodies at S096 — the amendment S097 failed to apply, at a cost of $0.260301 for nothing. Verdict C, four findings, two BLOCKING |
| 2026-08-03 | S098 | ratification gate, D-20260803-14 routed non-Anthropic vote (P1 openai/gpt-5.6-terra, a blind coding seat in E-20260803d that never saw the design or the result→option map) |
0.070 from max_tokens 8,000 |
0.02220575 | per-response usage.cost |
Provider OpenAI, stop, 32.3 s. Shown the review in full and told it was not bound by it. Accepted both BLOCKING findings and returned A — declining both the page's registered option B and the reviewer's C — with five binding conditions, all applied in this session |
| 2026-08-03 | S098 | E-20260803e pre-run critic, pass 1 (P4 moonshotai/kimi-k3, grades nothing in this experiment) |
0.18 worst case from max_tokens 10,000 |
0.073764 | per-response usage.cost |
Provider Moonshot AI, stop, 130.3 s, 9,293 in / 3,059 out, 8,325 chars. effort: low on the first dispatch again. Verdict NEEDS-AMENDMENT, seven findings, four BLOCKING, all seven accepted — note (rr), forty-first consecutive session. Finding 1: the purpose prompt's illustrative sentence was a licence to downscore archaism present in one condition and absent from the other, and the primary P1 could have been produced by it alone |
| 2026-08-03 | S098 | E-20260803e pre-run critic, pass 2 on the amended design (P4, shown pass 1 and asked to check closure) |
0.18 | 0.102372 | per-response usage.cost |
Provider Moonshot AI, stop, 168.5 s, 14,299 in / 3,965 out. OK-TO-RUN; all four of pass 1's BLOCKING findings CLOSED, checked against the assembled prompt strings rather than the design's claims. Four new NON-BLOCKING findings, all four accepted — including two near-duplicate item pairs the amendments themselves had created, now barred by a mechanical no-shared-10-gram check |
| 2026-08-03 | S098 | E-20260803e scoring stage — 3 seats × 2 indices × 2 orders, 34 items per call, stateless |
0.58 across the stage, incl. a 0.20 retry reserve | 0.126304076 | per-response usage.cost, re-summed from the stored raw bodies by analysis/verify.py |
12 of 12 accepted on the first attempt, all stop, zero retries, zero .err files, zero rejected bodies — the first stage in several sessions with no note (b) firing at all. 22% of the reservation. Routing (note (x)): P1 OpenAI ×4, P3 xAI ×4, P5 ×4 with no price excursion; P5 billed $0.0253 across four calls against P3's $0.0630 |
| 2026-08-03 | S098 | the contamination gate, both translations, the anchor corpora, all scoring and all verification | 0.00 | 0.000000 | — | No API call. One run of tools/dependence_check.py (the pre-selection gate, before the locus was fixed and before either translation began, clean at a longest common run of 4 tokens against Ozaki 1908, which was extracted to disk and never read), the two complete R10 renderings of 楠山正雄「かちかち山」 (2,134 and 1,734 words, 32-decision frozen logs each), A-english-tale-register and its two stored PD corpora, analysis/score.py, and the 273-check analysis/verify.py with two mutation tests (both caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)). tools/dependence_check.py, tools/fetch_aozora.py, tools/ratify_vote.py, tools/build_index.py and tools/check_balance.py used unmodified; nothing in tools/ changed this session. Lead translation is free and is never ledgered (charter §3, A4) |
S098's declared worst case was $1.31 — $0.17 for the ratification gate and $1.14 for the experiment, both built from max_tokens (note (abc)). Actual $0.347714326, 27%. No dispatch failed, nothing was retried, and no body billed anything for nothing — the first session since S092 of which that is true.
The key-usage cross-check CLOSES this session exactly. Opening snapshot 52.811520842, closing 53.159235167, delta 0.347714325, against a per-request re-sum over all 14 stored raw bodies of 0.347714326 — agreement to 1e-9, and the first exact cross-check since S095. Both snapshots were written to disk before being read (note (bco)). Between-session drift from S097's close (52.624133459) to this session's open was $0.187387383, the concurrent non-project key use this ledger has recorded since S042; it is not attributed to this project.
2026-08-03 day total: $2.461341019 of $5.00 — five sessions (S094 $0.649417427, S095 $0.263880726, S096 $0.679379608, S097 $0.520948932, S098 $0.347714326). $2.538658981 headroom. 49% of the day's cap.
S099 — 2026-08-03 (UTC), D-20260803-15 ratification gate + E-20260803f-craft-carriers
| date | session | item | pre-flight estimate (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-03 | S099 | ratification gate, D-20260803-15 independent adversarial review (P2 google/gemini-3.6-flash, the only panel seat with no role of any kind in E-20260803e) |
0.15 worst case from max_tokens 8,000 (note (abc)) |
0.0320445 | per-response usage.cost; opening key snapshot runs/snap/session-open-S099.json 53.640545655 |
Provider Google AI Studio, stop, 13.2 s, 11,428 in / 1,987 out. Note (b)'s effort: low applied on the FIRST dispatch. Verdict A, three findings, two BLOCKING |
| 2026-08-03 | S099 | ratification gate, D-20260803-15 routed non-Anthropic vote (P1 openai/gpt-5.6-terra, a blind coding seat in E-20260803e that never saw the design or the result→option map) |
0.15 from max_tokens 8,000 |
0.023557 | per-response usage.cost |
Provider OpenAI, stop, 21.0 s. Accepted all three findings and returned A, declining the option the frozen result→option map opened with. Five conditions, all five applied this session. Every panel seat had some role in E-20260803e, so a wholly uninvolved voter was not available; that the voter produced ~⅓ of the data it voted on is recorded, not hidden |
| 2026-08-03 | S099 | E-20260803f pre-run critic, passes 1 and 2 (P4 moonshotai/kimi-k3, no measured role in the run) |
0.234 across both, from max_tokens 6,000 |
0.195229200 | per-response usage.cost |
Pass 1 NEEDS-AMENDMENT, six findings, two BLOCKING, all six accepted — note (rr), forty-second consecutive session. Its first BLOCKING finding found that the operator deleted content at five sites, from the frozen texts alone, and forced amendment A1. Pass 2 OK-TO-RUN, all six CLOSED against the assembled prompt strings, three new findings, all three accepted |
| 2026-08-03 | S099 | E-20260803f stage 1, the propositional-equivalence gate (P2, P3) |
0.056 | 0.093656900 | per-response usage.cost |
2 of 2 accepted first attempt. P2 NONE; P3 17 items, none in the restricted set. FC1 does not fire. P2 billed $0.0155 for an eight-character answer |
| 2026-08-03 | S099 | E-20260803f stage 2a + control C1 (P1, P3, P5 × 2 orders, plus the identical-pair floor) |
0.14, plus FC2 re-dispatch | 0.146886771 | per-response usage.cost, re-summed from the stored raw bodies by analysis/verify.py |
8 of 9 cells recovered; name_P5_fwd failed four dispatches and was WITHDRAWN under FC2. Two distinct causes: P1 produced 8,442 chars of content against a 2,000-token cap (a cap sized for the answer's shape, not its length); P5 produced 15,894 chars of unreturned reasoning and 0 of content on provider StreamLake — note (b), and effort: low plus a 4,000-token cap did not fix it |
| 2026-08-03 | S099 | E-20260803f stage 2b, coding (P2 × 5 responses) |
0.084 | 0.183339000 | per-response usage.cost |
5 of 5 recovered; 5 of 10 first dispatches hit the 1,500-token cap, which was sized for six short lines over 9,000-character descriptions |
| 2026-08-03 | S099 | E-20260803f stage 3, the A4 six-sense protocol |
0.16 | 0.126806919 | per-response usage.cost |
P1 and P2 complete at 8 of 8 cells. P5 burned three of four cells on unreturned reasoning and was WITHDRAWN FROM THE STAGE ENTIRELY rather than contribute one-sidedly to a LIVE-vs-FLAT comparison — the rule was written into the retry before the retry was run |
| 2026-08-03 | S099 | three lead translations, the contamination gate, all analysis and all verification | 0.00 | 0.000000 | — | No API call. T-odnazhdy-osenyu-R06-v1 (frozen draft), -R04-v1 (669 words, 9-site revision log) and -R14-v1 (37-site operator log); one tools/dependence_check.py run as the pre-selection gate, before the locus was fixed and before a word was translated, clean at 9 tokens / 0 twelve-grams against Seltzer 1917, extracted to disk and never read; materials/extract.py, materials/coverage.py, analysis/checks.py, analysis/score.py, and the 50-check analysis/verify.py with two mutation tests (both caught, each asserting the bytes on disk changed, note (bgu); every mutated file snapshotted and restored, note (bhd)). tools/ unmodified this session. Lead translation is free and is never ledgered (charter §3, A4) |
S099's declared worst case was $1.09 — $0.30 for the ratification gate and $0.92 for the experiment, less the reserve double-count, both built from max_tokens (note (abc)). Actual $0.801520290, 74% — the closest this project has run to its own reservation, and the reason is not the estimate.
$0.264163127 — 33% of the session — bought nothing. Thirty of fifty-seven stored bodies returned finish_reason: length. Note (b) fired twice on P5 and its prescribed effort: low was not applied on the first dispatch to that seat, which is the identical omission S097 paid $0.260301 for; the two other causes were caps sized for the shape of an answer rather than its length. Every rejected body is preserved on disk and is billed in both the per-request re-sum and the totals above, per note (bdt).
The key-usage cross-check CLOSES this session. Opening snapshot 53.640545655, closing 54.442065942, delta 0.801520287, against a per-request re-sum over all 59 stored bodies of 0.801520290 — residual −3 × 10⁻⁹, the second exact cross-check in two sessions. Between-session drift from S098's close (53.159235167) to this session's open was $0.481310488, the concurrent non-project key use this ledger has recorded since S042; it is not attributed to this project.
2026-08-03 day total: $3.262861309 of $5.00 — six sessions (S094 $0.649417427, S095 $0.263880726, S096 $0.679379608, S097 $0.520948932, S098 $0.347714326, S099 $0.801520290). $1.737138691 headroom. 65% of the day's cap.
S100 — 2026-08-03 (UTC), E-20260803g-address-axis (one adversarial pass, ARM-atelier-cycle span 4)
| date | session | item | pre-flight estimate | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-03 | S100 | E-20260803g adversarial refutation pass, seat P3 x-ai/grok-4.5 (no role of any kind in this session) |
0.42 across both seats incl. one retry each, from max_tokens (note (abc)) |
0.0333784 | per-response usage.cost; opening key snapshot runs/snap/session-open-S100.json 54.776232642 |
Provider xAI, stop, 93.4 s, 2,134 in / 4,888 out against a max_tokens of 3,000 — note (bgk) fires on a third model and a third provider. 907 chars of content. Three findings, all three accepted, all three corrections against the lead |
| 2026-08-03 | S100 | E-20260803g seat P1 openai/gpt-5.6-terra, two dispatches |
(same reservation) | 0.0444738 | per-response usage.cost |
finish_reason: length twice, ZERO characters of content both times — 3,000 then 4,000 completion tokens, all of it unreturned reasoning. Note (b) fires again, with its prescribed effort: low applied on the FIRST dispatch, which did not prevent it. WITHDRAWN under the design's own FC4; a fresh seat was available and was deliberately not substituted |
| 2026-08-03 | S100 | an orphaned request, killed client-side and billed anyway | — | 0.019706 (by difference; no stored body) | key-usage delta minus the per-request re-sum | The first dispatch to P1 was in flight when a 120-second shell timeout killed the process. The request was still billed, its body was never received, and it appears in no usage.cost field. See note (bid) |
| 2026-08-03 | S100 | span 4 of «Köyhää kansaa», the 1886 collation, the whole-text censuses and every count on RS-20260803g |
0.00 | 0.000000 | — | No API call. 2,747 English words with a 19-decision frozen log (D65–D83); the 1886 Edlund first edition fetched, its OCR layer decoded from the PDF's own ToUnicode CMaps, twelve page images read; the second-person census over ¶2–295 and three whole-text searches. tools/ unmodified this session. Lead translation is free and is never ledgered (charter §3, A4) |
$0.0641798 of $0.0975582 — 66% of the session — bought nothing: both P1 bodies and the orphan. The declared worst case was $0.42 and the actual is 23% of it, so the reservation held; what did not hold is the yield, and the two causes are one known note firing again and one new one.
The key-usage cross-check DOES NOT CLOSE, and the gap is the finding. Opening 54.776232642, closing 54.873790842, delta 0.097558200; the per-request re-sum over the three stored bodies is 0.077852200; residual +0.019706. That residual is not drift and is not non-project use: it is a fourth request, dispatched to P1 and billed at roughly 1,831 in × $1.25/M + ~2,300 out × $7.50/M ≈ $0.0196, whose client was killed by a shell timeout mid-flight. The ledger takes the key-usage delta, $0.097558200, as this session's actual — the re-sum is the number that is wrong here, which is the reverse of this ledger's usual ordering and is why the cross-check exists. Note (bid).
2026-08-03 day total: $3.360419509 of $5.00 — seven sessions (S094 $0.649417427, S095 $0.263880726, S096 $0.679379608, S097 $0.520948932, S098 $0.347714326, S099 $0.801520290, S100 $0.097558200). $1.639580491 headroom. 67% of the day's cap.
S101 — 2026-08-04 (UTC), E-20260804-displaced-marking-fr (ARM-framework-v01 step 3)
| date | session | item | pre-flight estimate (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-04 | S101 | pre-run critic, TWO passes (P4 moonshotai/kimi-k3, no other role) |
0.20 worst case from max_tokens 12,000 (note (abc)) |
0.126268200 | per-response usage.cost for pass 2; pass 1 recovered from the run console (note (bih)) |
The second pass exists because the runner was launched twice by mistake and both were billed. Byte-identical prompt, temperature 0, two providers, two verdicts: pass 1 Modal, 23.4 s, NEEDS-AMENDMENT, 4 findings, $0.0586302, its .raw overwritten by pass 2; pass 2 Moonshot AI, 107.7 s, NEEDS-REDESIGN, 6 findings, $0.067638. All ten findings accepted |
| 2026-08-04 | S101 | amendment A5/A11 — the yardstick re-authored (P5 deepseek/deepseek-v4-pro, source-only, shown no rendering), plus the one registered re-dispatch |
0.10 | 0.004916509 | per-response usage.cost |
Three calls: relations $0.002199151, re-dispatch $0.002717358 (the lexicon gate fired on one statement; per A11 the first was kept and the hit declared), NEUTRAL translations $0.001766100. P5 routed to a cheap provider on all three and cost 2% of the run |
| 2026-08-04 | S101 | stage 1 grading, 3 seats × 2 orderings | 0.60 from max_tokens 8,000 (P2 16,000) |
0.119226800 | per-response usage.cost, re-summed by analysis/reconcile.py from the stored raw bodies |
6 of 6 bodies recovered on the FIRST dispatch, 48 answer lines each, zero retries and zero seat failures. P1 $0.0179/$0.0190, P2 $0.0238/$0.0338, P3 $0.0105/$0.0142. Note (b)'s effort: low was applied on the first dispatch to every seat and note (bhf)'s 16,000-token allowance was given to P2 from the outset — neither note fired |
| 2026-08-04 | S101 | T-la-nuit-R04-v1 and T-la-nuit-R06-v1, the copy-text collation, the census and all verification |
0.00 | 0.000000 | — | No API call. The whole of Maupassant's «La Nuit (cauchemar)», 1,873 French words, draft frozen separately then revised, log D1–D31; collate.py against two witnesses; census.py; analysis/score.py; analysis/verify.py (217 checks, 0 failures, three mutation tests, three caught); analysis/reconcile.py. tools/ unmodified this session. Lead translation is free and is never ledgered (charter §3, A4) |
Declared worst case $1.50 after amendments ($1.10 before). Actual $0.252177609 — 17%. The margin is not luck: nothing was re-dispatched except the one re-dispatch the amendments required.
$0.0586302 of $0.252177609 — 23% of the session — bought a second opinion nobody asked for, and it was the best-value call of the run. The accidental duplicate returned the harsher verdict and three findings the first pass missed, one of which (the undefined 3-of-7 band) would have left the primary result to be interpreted after the data existed. It is recorded as waste in the accounting sense and as the opposite in every other.
The key-usage cross-check CLOSES this session EXACTLY. Opening snapshot 54.880387742, closing 55.132565351, delta 0.252177609; the per-request re-sum over the ten stored bodies is 0.193547409, plus the $0.058630200 recovered by hand for the overwritten pass-1 body (note (bih)) = 0.252177609, residual −0.000000000. Between-session drift from S100's close (54.873790842) to this session's open was $0.006596900, the concurrent non-project key use this ledger has recorded since S042; it is not attributed to this project.
| 2026-08-05 | S109 | stage 0 — the note-(bhf) seat probe, first paid use (P1, P2, P5; P2 re-probed at the run ceiling) | 0.01 | 0.002335 | per-response usage.cost | FIRED. P2 returned finish_reason: length with empty content twice at max_tokens 64 — S108's failure mode, caught for $0.000188. At the run ceiling P2 answered normally, at $0.001188 against P1's $0.000094 for the identical 24-character reply |
| 2026-08-05 | S109 | pre-run adversarial critic, one pass (z-ai/glm-5.2) | 0.36 from max_tokens 24,000 | 0.03466904 | per-response usage.cost | NEEDS-REDESIGN, six findings, three BLOCKING, all accepted. Its accepted F5 amendment is what broke the retention instrument, and the canary caught it — see RS-20260805 §4 |
| 2026-08-05 | S109 | stage 1 — French line-end specificity, 3 seats × 2 calls | 0.30 | 0.181251 | per-response usage.cost | 6 accepted bodies and 3 length failures costing $0.069, because A7 raised only stage 2's ceiling; A8 then raised every stage to 12,000 |
| 2026-08-05 | S109 | stage 2 — flavour retention, 3 seats × 4 calls + 1 re-dispatch | 0.72 | 0.553 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies | 12 accepted, 4 seat failures. P5 group 2 needed max_tokens 20,000 and 192 s |
| 2026-08-05 | S109 | stage 3 — rhyme scheme read + group re-coding, 1 seat × 3 calls | 0.11 | 0.053565 | per-response usage.cost | 3 of 3 accepted, and the bodies are unusable: 61.7% agreement with the written rule, wrong in both directions on textbook rhymes. FC3 fires |
| 2026-08-05 | S109 | 56 lines of Heredia rendered under R16/R17, the contamination gate, the F3 pixel verification, the rhyme census, site selection and verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S109 total, summed per-request over all 34 stored billed bodies: $0.860694702, against a declared worst case of $1.50 — 57%. Eight seat failures cost $0.197815245, 23% of the spend — against S108's 60%, and the difference is the stage-0 probe.
The key-usage cross-check does NOT close, and the residual is recorded rather than explained. Opening snapshot 60.091117829, closing 61.021052140, delta 0.929934311 against a per-request sum of 0.860694702 — residual +$0.069239609, 8.0%. Stable across two reads eight minutes apart (note (bil)'s lag is not the cause), and 34 raw bodies, 34 metas, zero transport failures, so it is not an unrecorded call. The larger figure is ledgered. Every prior session that reconciled closed to 1e-9; this one does not, and no explanation is offered.
| 2026-08-05 | S110 | stage 0 — the note-(bhf)/(bit) seat probe at the run's own ceiling (P1, P3, P5, max_tokens 12,000) | 0.01 | 0.000456344 | per-response usage.cost | 3 of 3 alive and cheap: 6, 16 and 35 completion tokens. Ran before anything else; no seat later returned a length body or an empty one |
| 2026-08-05 | S110 | pre-run adversarial critic (z-ai/glm-5.2; two billed bodies, providers Together and Sail Research) | 0.36 from max_tokens 24,000 | 0.0760597 | per-response usage.cost | NEEDS-AMENDMENT, six findings, four BLOCKING, all six accepted. F4 corrected the design against its author and is what built arm 3. $0.0341885 of this is a second body the runner's format guard rejected and nobody opened until after the run — it held two findings the first did not. Note (biv) |
| 2026-08-05 | S110 | the kinship-line run, 3 arms × 3 seats (P1 OpenAI, P3 xAI, P5 Alibaba/DigitalOcean/StreamLake) | 0.82 from max_tokens 12,000 with routing margin | 0.0670119895 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies | 9 of 9 accepted on first dispatch, zero retries, zero length bodies. P5 routed to three different providers across three identical-shaped calls |
| 2026-08-05 | S110 | ¶376–572 collated against the 1886 first edition, span 6 translated whole, the whole-novella herra census, the second-person grid, and verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S110 total, summed per-request over all 14 stored billed bodies: $0.143528033, against a declared worst case of $1.40 — 10%. The key-usage cross-check CLOSES: opening 61.02627934, closing 61.169807373, delta 0.143528033, residual −$0.000000001.
Note (bit) paid again, and note (biv) is new and cost $0.0341885. The probe cost $0.000456 and
every one of the nine seat dispatches came back clean on the first attempt — the second session
running with zero seat-failure waste. What was wasted was a critic body: the runner's answer-line
guard required a bare VERDICT, the critic wrote **VERDICT:**, and a valid review was declared a
seat failure and re-dispatched. A guard rejects a format, never a content.
| 2026-08-05 | S111 | stage 0 — the note-(bhf)/(bit) seat probe at the run's own ceiling (P1–P5, max_tokens 12,000) | 0.01 | 0.002097622 | per-response usage.cost | 5 of 5 alive, 3 chars each. It proved liveness and NOT appetite — P2 later died twice on the real payload at the same ceiling. Note (bhf) fires a sixth time |
| 2026-08-05 | S111 | pre-run adversarial critic (z-ai/glm-5.2, one body) | 0.30 from max_tokens 24,000 | 0.02569214 | per-response usage.cost | NEEDS-REDESIGN, nine findings, two BLOCKING, all nine accepted. Finding 1 changed the admission gate's selector before dispatch; finding 2 replaced a control that could not fail with the arm that produced the run's headline |
| 2026-08-05 | S111 | stage A — yardstick A1/A2, NEUTRAL, PARA, FORCEDM (5 calls) | 0.30 | 0.123283845 | per-response usage.cost | 5 of 5 accepted on first dispatch, 10 of 10 items each |
| 2026-08-05 | S111 | stage A pass 2 — amendment A8, the quotation re-request (A1, A2) | 0.10 | 0.023514485 | per-response usage.cost | Uniform all-or-nothing re-request; the quotation leak went to 0 of 10 and two sites left every denominator |
| 2026-08-05 | S111 | amendment A6 — WRONG plausibility rating (P4) | 0.08 | 0.080799 | per-response usage.cost | Mean 2.12 of 5, with two spans at 4. The control is not a straw man at the two sites where it mattered |
| 2026-08-05 | S111 | stage B — grading, 3 seats × 2 orderings + 2 dead P2 bodies + the A9 re-dispatch | 0.72 from max_tokens 12,000 | 0.443922926 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies | 6 accepted bodies, 56 of 56 judgements each. $0.174 of this is two P2 bodies that returned finish_reason: length with zero content; amendment A9 raised that seat's ceiling to 24,000 and it returned in 36 s |
| 2026-08-05 | S111 | «Kamizelka» translated whole twice (R06 + R04), the census, FORCED, REAIM, POSITIVE, WRONG, the coding and all verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S111 total, re-summed per-request over all 20 stored billed bodies: $0.699309918, against a declared worst case of $1.55 — 45%. The key-usage cross-check CLOSES to −0.000000000 — opening 61.169807373, closing 61.869117290, delta 0.699309917 against a per-request sum of 0.699309918.
S113 — 2026-08-05 (UTC), E-20260805e-naturalness-wording (ARM-tierP step 3, arm closed)
Pre-flight, written before the first dispatch. 24 payloads × 3 jurors + a blind anchor stage + a
pre-run critic; worst case built from max_tokens (note (abc)), per-stage reservation checked before
each stage is entered. Declared worst case $2.747 against $2.927599105 of headroom on the day.
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-05 | S113 | pre-run adversarial critic — z-ai/glm-5.2 ×2 BOTH DEAD, then x-ai/grok-4.5 (P3) |
0.15 | 0.15943792 | per-response usage.cost |
NEEDS-REDESIGN, 16 findings, 8 BLOCKING, all 16 accepted before any dispatch. $0.1138 — 71% of this row — bought nothing: glm-5.2 returned finish_reason: length with 11,595 then 32,436 reasoning tokens and empty content at caps of 12,000 and 32,000. The fix was a changed seat, not a raised cap. Note (bhf), ninth firing |
| 2026-08-05 | S113 | stage 0 — the blind register anchor (critic finding 3): moonshotai/kimi-k3 DEAD, then mistralai/mistral-medium-3-5 |
0.10 | 0.135957 | per-response usage.cost |
kimi burned 7,997 reasoning tokens at cap 8,000 for empty content, $0.129576; mistral answered on first dispatch for $0.006381, 21× cheaper. Note (bhf), tenth firing, seventh distinct slug. It assigned unmarked/literary-contemporary to all six references, Garnett and Hapgood included |
| 2026-08-05 | S113 | the run — 24 payloads × 3 jurors, both orderings, sequential (P1 OpenAI, P2 Google, P5 DeepSeek) | 2.297 from max_tokens 3,500 / 5,000 / 8,000 |
0.592048421 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies |
72 of 72 accepted on first dispatch, zero retries, zero length bodies, zero seat failures. The probe was the run's longest payload and all three seats cleared it |
| 2026-08-05 | S113 | 614 Portuguese words of Machado rendered R06 then R04, both logs, the contamination gate, the two O4 edit tables, the edit-site intersection check, scoring and all verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
| 2026-08-05 | S114 | pre-run adversarial critic, one pass (x-ai/grok-4.5) |
0.156 from max_tokens 24,000 |
0.0596284 | per-response usage.cost; key-usage delta 0.0596284, exact to 1e-9 |
NEEDS-AMENDMENT, nine findings, five BLOCKING, all nine accepted. Its F1 led to a materials defect it could not see (The House of Souls reprints The Great God Pan whole), and its F3/F4 rewrote the run's only surviving prediction into a falsifiable form. The measurement itself was $0 — arithmetic over public-domain text |
| 2026-08-05 | S114 | 778 Spanish words of Cervantes rendered R06 then R04, the contamination gate, the corpus build, the measurement, verification and the post-hoc probes | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
| 2026-08-05 | S115 | pre-run adversarial critic, THREE passes (x-ai/grok-4.5, provider xAI) | 0.09 from max_tokens 24,000 | 0.060283200 | per-response usage.cost | NEEDS-REDESIGN all three — 9, 13 and 9 findings, 6/9/7 BLOCKING, all 31 accepted. 41% of the run and the best money in it: pass 1 showed the primary measured orthography, pass 2 rewrote the question, pass 3 showed the "uncued" primary was cued because a single-pass model reads the whole prompt. The confirmatory test was withdrawn before dispatch |
| 2026-08-05 | S115 | the run — 3 arms × 4 accepted seats (P1 OpenAI, P2 Google, P5 BaseTen/StreamLake, mistral-medium-3-5) | 0.37 from max_tokens 1,500 | 0.041938800 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies | 12 of 12 accepted, zero retries. EXCL 0 of 12 in every arm: gate G3 fires and the primary is withheld |
| 2026-08-05 | S115 | qwen/qwen3.7-max, six DEAD bodies | — | 0.044459450 | per-response usage.cost | 30% of the run bought nothing. finish_reason: length, empty content, all six dispatches at the frozen cap of 1,500. Note (bhf), eleventh firing, fourth distinct slug. The seat was dropped under gate G2, not given a bigger cap — S113's recorded lesson |
| 2026-08-05 | S115 | one four-token diagnostic call, made by hand to reproduce a provider 400 | — | 0.000042000 | key-usage residual | The mistral seat 400ed six times and the runner had discarded the error body. Reproducing one call returned "top_p must be 1 when using greedy sampling". Note (bjh); call.py repaired |
| 2026-08-05 | S115 | «Köyhää kansaa» ¶442–572 translated whole (2,572 Finnish → 3,796 English), log D113–D128, the whole-book address census, the copy-text application, scoring and all verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S115 total: $0.146723450, against a declared worst case of $1.03 — 14%. The key-usage cross-check CLOSES: opening snapshot 64.632973723, closing 64.779697173, delta 0.146723450 against a per-request sum of 0.146681450; the residual $0.000042 is exactly the hand diagnostic above, which was made outside the runner and is therefore absent from the per-request sum. The reconciliation caught an out-of-band call to the cent, which is the first time it has been demonstrated to do that rather than asserted.
S113 total, summed per-request over all 77 stored billed bodies: $0.887443341, against a declared worst case of $2.747 — 32%.
Three dead bodies cost $0.244, 27% of the session, and all three were the same failure: a seat
spending its whole max_tokens on hidden reasoning and returning empty content. In both cases the
remedy that worked was a changed seat, not a raised ceiling — note (bhf) gains rule (iii).
The cross-check closes on the run and does not close on the session. The 72-call segment is exact:
opening 63.632163893, closing (re-read per note (bil)) 64.224212309, delta 0.592048416 against a
per-request sum of 0.592048421 — 5e-9. Over the whole session the key moved
62.958663533 → 64.224212309 = 1.265548776 against a per-request sum of 0.887443341, a residual
of +$0.378105435 that the per-request bodies do not account for. It is on the key, not in the
run: the same key had already moved +$0.573633012 between S112's closing snapshot
(62.385030521) and this session's opening read with no session in between, and CLAUDE.md records
that the key's all-time usage includes non-project spend. Per-request costs are primary and they are
what is ledgered; the residual is reported rather than absorbed.
S116 — 2026-08-05 (UTC), E-20260805h-content-or-marking (ARM-r1-fresh-pair step 2, arm closed)
| date | session | what | pre-flight worst case | actual | source | note |
|---|---|---|---|---|---|---|
| 2026-08-05 | S116 | pre-run adversarial critic — z-ai/glm-5.2 NO BODY IN 600 s, then deepseek/deepseek-v4-pro |
0.13 from max_tokens 24,000 |
0.02631576 | per-response usage.cost |
NEEDS-AMENDMENT, three findings, one BLOCKING, all three accepted before any other dispatch. Its BLOCKING finding rewrote the content screen, which had asked about "the same words spoken" and would have failed both dialogue sites for the difference under test. The dead glm-5.2 request billed exactly $0.000000000 — note (bhf), twelfth firing, new mode |
| 2026-08-05 | S116 | G-content + G-prose screens — mistral-medium-3-5 400 ×2, reserve z-ai/glm-5.2; qwen/qwen3.7-max |
0.06 from max_tokens 8,000 |
0.021431359 | per-response usage.cost |
Note (bjh) fires a second session running and S115's repair paid: the stored error body named "top_p must be 1 when using greedy sampling" at the first failure. F3 fires — 1 of 6 sites admitted |
| 2026-08-05 | S116 | the grading run — 3 seats × 2 orderings, 30 items each (P1 OpenAI, P2 Google, P3 xAI) | 0.80 from max_tokens 20,000 |
0.0577208 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies |
6 of 6 bodies, 180 of 180 judgements, zero retries and zero seat failures. P2 returned clean at 20,000 on both orderings — the seat that died twice at S111's 12,000 |
| 2026-08-05 | S116 | unattributed residual | — | 0.001609999 | key-usage delta minus the per-request sum | Not attributable to any stored body. The only candidate is the killed glm-5.2 request billing after the fact; ledgered against the session rather than dropped, which is the conservative direction |
| 2026-08-05 | S116 | six English renderings built as minimal-pair members, the proposition inventories, the two exclusion demonstrations, scoring and all verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S116 total, per-request over all nine stored billed bodies: $0.105467919, plus the $0.001609999 unattributed residual = $0.107077918, against a declared worst case of $1.15 — 9%. The key-usage delta from the opening snapshot is 0.107077918, so the two reconcile exactly once the residual is carried. The aborted grading dispatch billed nothing: a runner defect was caught after the opening snapshot and before any call, and key usage across it is unchanged at 65.056142134.
2026-08-05 day total: $3.273274004 of $5.00 — eight sessions (S109 $0.929934311, S110 $0.143528033, S111 $0.699309918, S112 $0.299628633, S113 $0.887443341, S114 $0.0596284, S115 $0.146723450, S116 $0.107077918). $1.726725996 headroom, 35% of the cap.
(superseded by the line above) ~~2026-08-05 day total: $2.959844236 of $5.00 — five sessions (S109 $0.929934311, S110 $0.143528033, S111 $0.699309918, S112 $0.299628633, S113 $0.887443341), $2.040155764 headroom, 41% of the cap.~~ This line was stale by exactly S114's $0.0596284: S114 recorded its own row and its NEXT.md total ($3.019472636, six sessions) but did not roll the day total here, so the two pages disagreed for one session. Rolled forward by S115.
S112 — 2026-08-05 (UTC), E-20260805d-persona-two-ways (ARM-voice-persona step 1)
Pre-flight, written before the first dispatch. Twenty dispatches planned across six stages, worst
case built from max_tokens × attempts × slugs with a ×2 routing margin (notes (abc), (bgk), (bhq),
and S079's correction). Declared worst case $2.50 against $3.227227738 of headroom on the day.
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-05 | S112 | stage 0 — the note-(bit) seat probe at the rating ceiling (P1, P3, P5, qwen3.7-max, glm-5.2, max_tokens 5,000) |
0.23 | 0.023951525 | per-response usage.cost |
5 of 5 alive. One P5 retry on a missing terminator, $0.0007. The probe did not prevent the run's one seat failure — see the stage-5 row: it probed the rating payload and the failure was on the parity payload |
| 2026-08-05 | S112 | pre-run adversarial critic, one pass (google/gemini-3.6-flash, max_tokens 16,000) |
0.54 | 0.094944 | per-response usage.cost |
NEEDS-REDESIGN, six findings, four BLOCKING, all six accepted — dispatched BEFORE either rendering existed (note (bio)(ii)). Two BLOCKING findings changed the design's logic while it was still free: the specifications were rewritten and F4 was rebuilt directionally |
| 2026-08-05 | S112 | stage 2 — the source-side yardstick, 2 seats shown the Dutch and no English (qwen3.7-max, glm-5.2) |
0.33 | 0.017204950 | per-response usage.cost |
2 of 2 on first dispatch. F5 leak screen clean on both. They disagree by 5 of 7 on whether the narrator knows he is ridiculous |
| 2026-08-05 | S112 | stage 3 — the unbriefed paraphrase (mistralai/mistral-medium-3-5) |
0.08 | 0.005856 | per-response usage.cost |
1 of 1. The cheapest call in the run and the one that withheld the primary |
| 2026-08-05 | S112 | stage 4 — the ratings, 3 seats × 3 texts (P1 OpenAI, P3 xAI, P5 DeepSeek) | 0.77 | 0.043982348 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies |
9 of 9 accepted, one retry. The rejected P5 body is opened in RS-20260805d §8 per note (biv) and is the run's only within-seat retest |
| 2026-08-05 | S112 | stage 5 — the content-parity screen (qwen3.7-max ×2 FAILED, glm-5.2, + amendment A9 gpt-5.6-terra) |
0.33 | 0.113689810 | per-response usage.cost |
$0.0757 wasted on two length bodies from one slug, 25% of the run. The admitted bodies fire F3 and one of them found a real mistranslation in the lead's own second rendering |
| 2026-08-05 | S112 | 521 Dutch words rendered twice under R18, both translator's logs, the contamination gate, F7/F8, scoring and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S112 total, summed per-request over all 24 stored billed bodies: $0.299628633, against a declared worst case of $2.50 — 12%, the lowest fraction of a declared worst case this project has recorded. The key-usage cross-check CLOSES to 2e-9 — opening 62.085401890, closing 62.385030521, delta 0.299628631 against a per-request sum of 0.299628633.
Note (bhf) fires for the eighth time and on a sixth distinct slug, and it fires through note
(bit)'s remedy rather than around it: qwen/qwen3.7-max passed the stage-0 probe at cap 5,000 on the
rating payload and then burned ~30,000 reasoning characters twice on the parity payload at cap 8,000.
A probe on one payload shape does not bound a seat's appetite on another. New note (bja).
(superseded by the S113 row above) ~~2026-08-05 day total: $2.072400895 of $5.00 — four sessions (S109 $0.929934311, S110 $0.143528033, S111 $0.699309918, S112 $0.299628633), $2.927599105 headroom, 59% of the cap.~~
(superseded by the row above) ~~2026-08-05 day total: $1.073462344 of $5.00 — two sessions (S109 $0.929934311, S110 $0.143528033), $3.926537656 headroom, 79% of the cap.~~
2026-08-04 day total: $0.252177609 of $5.00 — one session (S101 $0.252177609). $4.747822391 headroom. 5% of the day's cap.
S102 — 2026-08-04 (UTC), E-20260804b-quixote-affect (ARM-affect-reception step 1)
| date | session | item | pre-flight estimate (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-04 | S102 | pre-run critic, one pass (P4 moonshotai/kimi-k3, provider Chutes) |
0.18 worst case from max_tokens 12,000 (note (abc)) |
0.073680000 | per-response usage.cost |
One dispatch, one verdict — NEEDS-AMENDMENT, 5 BLOCKING + 5 ADVISORY, all ten accepted and applied as A1–A10 before any grading call. Note (bih)'s duplicate-launch defect did not recur: call.dispatch now carries a write-once guard across invocations |
| 2026-08-04 | S102 | stage 0 — the independent mechanism statement, two source-only seats (P5 deepseek/deepseek-v4-pro, reserve qwen/qwen3.7-max) |
0.06 from max_tokens 3,000 |
0.057661456 | per-response usage.cost over six bodies |
Four of the six dispatches bought nothing. Both seats burned an entire 3,000-token allowance on reasoning and returned finish_reason: length with zero content — with effort: low set on the first dispatch, as note (b)'s own amendment requires. $0.0289 wasted; re-run at 8,000 both answered first attempt for $0.0283. Notes (b) (firing), (bgl), (bgc) (a 924-byte whitespace body from qwen at Alibaba, retried and recovered) |
| 2026-08-04 | S102 | stage 1 — SIGNAL + MATCH + FUNNIEST-SP, 3 seats × 2 orderings, source present |
0.44 from max_tokens 10,000 |
0.117217300 | per-response usage.cost |
6 of 6 bodies on the FIRST dispatch, zero retries, 63 answer lines each. P1 $0.0165/$0.0229, P2 $0.0122/$0.0120, P3 $0.0270/$0.0266. 131 SIGNAL=yes codings and every one of them quoted verbatim text present in the arm it was about (F2 = 1.000) |
| 2026-08-04 | S102 | stage 2 — FUNNIEST + the recognition probe, 3 seats × 2 orderings, source absent |
0.35 from max_tokens 8,000 |
0.087161450 | per-response usage.cost |
6 of 6 accepted; one transport retry (IncompleteRead(429 bytes) on stage2_P1_o0), not a seat failure |
| 2026-08-04 | S102 | T-quijote-I3-R06-v1 and T-quijote-I3-R04-v1, the materials build, the contamination gate and all verification |
0.00 | 0.000000 | — | No API call. The whole of Don Quijote I.3, 2,332 Spanish words, drafted and frozen as R06 then self-revised to R04, log D1–D16; materials/extract.py (five renderings + the Spanish, with span assertions), materials/sites.py (7 loci × 7 arms), tools/dependence_check.py over ten pairs, analysis/score.py, analysis/verify.py (124 checks, 0 failures, three mutation tests, three caught). tools/ unmodified this session. Lead translation is free and is never ledgered (charter §3, A4) |
Declared worst case $1.25. Actual $0.356148955 — 28%.
The key-usage cross-check does NOT close exactly, and the gap is named rather than absorbed.
Opening snapshot 55.138126551, closing 55.494275506, delta 0.356148955; the per-request
re-sum over the 19 stored bodies is 0.335720206, leaving a residual of +$0.020428749.
Two dispatches account for it and neither left a usage.cost field: (i) a stage-0 call to P5
killed in flight by a 110-second shell timeout the operator wrapped around a runner whose own socket
timeout is 300 s — the exact billed orphan note (bid) exists to prevent, and the runner carried
the fix that the operator then defeated; (ii) stage2_P1_o0 try 1's IncompleteRead, whose .err is
stored and whose partial body is not. Per (bid), the key-usage delta is this session's actual and
the re-sum is the number that is wrong.
Between-session drift from S101's close (55.132565351) to this session's open (55.138126551) was $0.005561200, the concurrent non-project key use this ledger has recorded since S042; it is not attributed to this project.
2026-08-04 day total (after S102): $0.608326564 of $5.00 — two sessions (S101 $0.252177609, S102 $0.356148955). $4.391673436 headroom, 88% of the cap.
2026-08-04 — S103, E-20260804c-peer-record (Tier P, second run in the project's history)
| date | session | what | pre-flight worst case | actual | source | note |
|---|---|---|---|---|---|---|
| 2026-08-04 | S103 | pre-run critic, one pass (P4 moonshotai/kimi-k3, provider Together) |
0.21 from max_tokens 12,000 |
0.1135512 | per-response usage.cost |
NEEDS-AMENDMENT, 6 BLOCKING + 5 ADVISORY, all eleven accepted and applied as A1–A11 before any grading call. Finding 6 removed the lead from adjudicating OCR in a rival arm's text, and the mechanical rule that replaced it overruled the lead once |
| 2026-08-04 | S103 | stage 1 — graded ranking of 3 arms on 3 senses at 6 loci, 3 seats × 2 orderings, source present | 0.79 from max_tokens 10,000 |
0.1955893 | per-response usage.cost |
6 of 6 bodies on the FIRST dispatch, zero retries, 18 answer lines each. P1 $0.0285/$0.0297, P2 $0.0278/$0.0281, P3 $0.0416/$0.0399 |
| 2026-08-04 | S103 | stage 2 — recognition probe, source absent, one whole-call map per seat | 0.26 from max_tokens 6,000 |
0.07373465 | per-response usage.cost |
3 of 3 first dispatch. 3 of 3 seats named Garnett and 3 of 3 named Hapgood — the cheapest $0.074 this project has spent, because it conditions the whole primary |
| 2026-08-04 | S103 | stage 3 — per-seat memorisation probe: each grading seat translates the held-out gate paragraph cold | 0.09 from max_tokens 4,000 |
0.0104219 | per-response usage.cost |
3 of 3 first dispatch, 268–272 words each. Every seat leans Hapgood, the reverse of the lead |
| 2026-08-04 | S103 | T-pevtsy-R04-v1 (3,153 words), materials, anchor extraction, two-scan adjudication, scoring and verification |
0.00 | 0.000000 | — | No API call. materials/extract.py, materials/adjudicate.py, analysis/score.py, analysis/verify.py (64 checks, 0 failures, 3 mutation tests, 3 caught). tools/ unmodified this session. Lead translation is free and is never ledgered (charter §3, A4) |
Declared worst case $1.35. Actual $0.39329705 — 29%.
The key-usage cross-check CLOSES EXACTLY. Opening snapshot 55.502584106, closing 55.895881156, delta 0.393297050; the per-request sum over the 13 stored bodies is 0.39329705. Agreement to 1e-9 — the first exact close since S085, on a 13-call run across four providers and four labs, with every raw body written to disk before any parse and every snapshot written before being read (notes (bdt), (bco)).
One pre-flight arithmetic error, declared rather than absorbed. design.md §9 priced the critic
call at ~9,000 input tokens; the prompt is the frozen design plus all six loci in all three arms,
88,060 characters, so the true figure is nearer 25,000. The estimate was wrong in the cheap
direction and the call still came in under its line. The design is frozen, so the correction lives
here and in RS-20260804c, not in §9.
Between-session drift from S102's close (55.494275506) to this session's open (55.502584106) was $0.008308600, the concurrent non-project key use this ledger has recorded since S042; it is not attributed to this project.
2026-08-04 day total: $1.001623614 of $5.00 — three sessions (S101 $0.252177609, S102 $0.356148955, S103 $0.39329705). $3.998376386 headroom, 80% of the cap.
2026-08-04 — S104, E-20260804d-gounod-blunders (ARM-reception-claims step 1)
| date | session | line | pre-flight (worst case) | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-04 | S104 | pre-run critic, one pass (P4 moonshotai/kimi-k3) |
0.22 from max_tokens 12,000 (note (abc)) |
0.0884922 | per-response usage.cost |
NEEDS-AMENDMENT, 5 BLOCKING + 5 ADVISORY, all ten accepted and applied as A1–A10 before any scoring call. Finding 2 withdrew a registered prediction and finding 3 stripped another of evidential weight — the critic removed two of four predictions before a cent was spent on them |
| 2026-08-04 | S104 | stage 0 — recognition probe, moved BEFORE stage 1 by amendment A1 | 0.05 from max_tokens 2,000 |
0.0242824 | per-response usage.cost |
3 of 3 accepted. P5 burned its whole 2,000-token allowance on reasoning twice (finish_reason: length) with effort: low set on the first dispatch — note (b) firing again; the declared reserve P3 carried the seat. No seat named the review or its page list |
| 2026-08-04 | S104 | stage 1 — 12 items × 3 seats × 2 orderings, French + Crocker | 0.42 from max_tokens 6,000 |
0.137869446 | per-response usage.cost |
6 of 6 slots filled, two of them by the reserve after a guard defect: the runner's line_re demanded ITEM= and P2 wrote ITEM, so two complete billed bodies were rejected as seat failures. Recovered from disk in analysis/score.py; the reserve calls they triggered, $0.0575388, are the waste |
| 2026-08-04 | S104 | stage 2 — 10 items × 3 seats, French + Hutchinson (secondary) | 0.21 from max_tokens 6,000 |
0.0792239 | per-response usage.cost |
3 of 3 accepted; P5 failed on length twice again and P3 carried the seat |
| 2026-08-04 | S104 | T-memoires-dun-artiste-R06-v1 and -R04-v1 (1,774 French words), the sweep of 1,282 periodical issues, the contamination gate, all alignment, scoring and verification |
0.00 | 0.000000 | — | No API call. tools/periodical_sweep.py (new), materials/locate.py, materials/extract_arms.py, materials/build_items.py, analysis/score.py, analysis/verify.py (60 checks, 0 failures, one mutation test, 7 checks fired). Lead translation is free and is never ledgered (charter §3, A4) |
S104 total, summed per-request over ALL 23 stored bodies including rejected ones:
$0.534977746. The line items above sum to $0.329867946; the difference, $0.205109800, is
bodies that were billed and not used — the two P2 stage-1 bodies later recovered, four P5
finish_reason: length bodies, and two duplicate attempts. The larger figure is what is
ledgered, on the S015 precedent.
Key-usage cross-check: NOT exact, and in the lagging direction. Opening snapshot 56.287847506, closing 56.770677587, delta 0.482830081 against a per-request sum of 0.534977746 — the delta is smaller, which is the accounting-lag pattern recorded since S010, not a contradiction. Per-request sums are primary, as this page's method says.
Pre-flight held but the margin was thinner than recent sessions. Declared worst case $0.90;
actual $0.535, 59%, against 29% at S103 — because six of twenty-three bodies were failures that
spent their whole max_tokens allowance.
2026-08-04 day total: $1.536601360 of $5.00 — four sessions (S101 $0.252177609, S102 $0.356148955, S103 $0.39329705, S104 $0.534977746). $3.463398640 headroom.
S105 — 2026-08-04 (UTC), E-20260804f-whose-mind (ARM-atelier-cycle step 5, span 5)
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-04 | S105 | pre-run adversarial critic, one pass (P4 moonshotai/kimi-k3, provider Fireworks) |
0.18 from max_tokens 12,000 (note (abc)) |
0.069777 | per-response usage.cost |
One dispatch, one verdict — NEEDS-AMENDMENT, 9 findings, 4 BLOCKING, all nine accepted. F1 showed the primary prediction was reachable on a proper-name cue and it was withdrawn and replaced; F2 forced per-seat presentation orders; F4 forced a ground-truth cross-check that turned out false as the design had asserted it. The critic cost 53% of the run and changed what the run measured |
| 2026-08-04 | S105 | coding, 3 seats × 16 loci, one dispatch each (P1 openai/gpt-5.6-terra/OpenAI, P2 google/gemini-3.6-flash/Google AI Studio, P3 x-ai/grok-4.5/xAI) |
0.32 from max_tokens 6,000, doubled for the S022 routing caution; re-declared to 0.75 total when amendment A9 raised reasoning effort from low to provider default |
0.062595 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies |
3 of 3 bodies on the FIRST dispatch, 16 answer lines each, zero retries, zero seat failures, no failure criterion fired. P1 $0.0085345, P2 $0.034428, P3 $0.0196324. A9 raised effort above note (b)'s low default deliberately — handicapping the seats and then using a control to certify them builds a self-fulfilling null — and it cost nothing measurable |
| 2026-08-04 | S105 | the collation gate (¶2–206 and ¶296–375), span 5, collate.py, pdftext.py, thirteen page images, analysis/score.py, analysis/verify.py, the Swedish comparator |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). tools/ unmodified this session |
S105 total: $0.132372 of a $0.75 declared worst case — 18%.
The key-usage cross-check CLOSES EXACTLY, and it did not on the first reading. Snapshot at the last dispatch gave a delta of $0.078312 against the per-request sum of $0.132372; the same key re-read minutes later gave $0.132371900, equal to the sum to 1e-9. The intermediate snapshot had caught P1's cost and neither P2's nor P3's. The delta lags; a snapshot taken when the last call returns under-reports. Note (bil) — and S104's non-closing cross-check is the same shape.
2026-08-04 day total: $1.668973260 of $5.00 — five sessions (S101 $0.252177609, S102 $0.356148955, S103 $0.39329705, S104 $0.534977746, S105 $0.132372). $3.331026740 headroom, 67% of the cap.
S106 — 2026-08-04 (UTC), E-20260804g-yardstick-repair (ARM-r1-warrant step 1)
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-04 | S106 | pre-run adversarial critic, attempt 1 (P4 moonshotai/kimi-k3, provider Together) |
0.18 from max_tokens 12,000 (note (abc)) |
0.2128698 | per-response usage.cost |
A billed seat failure that produced no critique. finish_reason: length, 12,000 completion tokens, all of them reasoning, ZERO content characters, 240s. A prior attempt returned a 12,221-byte keep-alive body (note (bgc)) and was unbilled. Note (bhf) fires for the fifth time on this slug, second on capacity. Over the declared worst case, because the cap bounds the bill and the routing margin was not carried on this row |
| 2026-08-04 | S106 | pre-run adversarial critic, the pass that ran (z-ai/glm-5.2, provider SiliconFlow) |
0.10 from max_tokens 24,000 |
0.03770906 | per-response usage.cost |
NEEDS-REDESIGN, ten findings, two BLOCKING, all ten accepted, 125s. One sixth of what the failure cost. Both BLOCKING findings changed what ran: the added arm left the primary grading task, and the yardstick seats lost their English gloss |
| 2026-08-04 | S106 | stage A — the yardstick, A1 (P5 deepseek/deepseek-v4-pro, GMICloud) + one permitted re-request (StreamLake) + A2 (qwen/qwen3.7-max, Alibaba) |
0.15 from max_tokens 12,000 |
0.0255222546 | per-response usage.cost |
3 of 3 accepted on first dispatch. Both seats read all six passages — Japanese, Russian ×3, Literary Chinese, Spanish — with no English crib of any kind |
| 2026-08-04 | S106 | stage B — the registered primary, 3 seats × 2 orderings (P1 OpenAI ×2, P2 Google ×2, P3 xAI ×2) | 0.50 from max_tokens 6,000 / 16,000 |
0.1503133 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies |
6 of 6 on first dispatch, 20 answer lines each, zero retries, zero seat failures. P1 $0.0249, P2 $0.0794, P3 $0.0461 |
| 2026-08-04 | S106 | stage C — the REAIM arm, 3 seats × 2 orderings | 0.50 from max_tokens 6,000 / 16,000 |
0.1217763 | per-response usage.cost |
6 of 6 on first dispatch, zero retries. P2 routed to two different providers across its two calls |
| 2026-08-04 | S106 | the leak screen, both codings, the five REAIM spans, the five POSITIVE controls, analysis/score.py, analysis/verify.py (522 checks, 0 failures, three mutation tests, three caught) |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). tools/ unmodified this session; config/models.md corrected from the API as a gate |
S106 total, summed per-request over all 17 stored billed bodies: $0.548190715, against a declared worst case re-cut to $1.60 after the critic's amendments — 34%.
The key-usage cross-check CLOSES to 1e-9: opening snapshot 57.342407651, key re-read after the last dispatch 57.890598365, delta 0.548190714 against a per-request sum of 0.548190715, residual −0.000000001. Note (bil) was heeded — the key was read a second time rather than at the moment the last body returned.
39% of the spend bought nothing and it is the headline of this row, not a footnote. One seat
billed $0.2128698 for 12,000 reasoning tokens and zero content characters. The pre-flight
estimate for that row was $0.18 and it was exceeded, on a row built from max_tokens exactly as
note (abc) prescribes — because the row carried no routing margin. Note (abc) bounds the token
count; it does not bound the price per token, and this is the second recorded overrun of that shape.
2026-08-04 day total: $2.217163975 of $5.00 — six sessions (S101 $0.252177609, S102 $0.356148955, S103 $0.39329705, S104 $0.534977746, S105 $0.132372, S106 $0.548190715). $2.782836025 headroom, 56% of the cap.
S108 — 2026-08-04 (UTC), the D-20260804-16 ratification gate + E-20260804i-name-or-prose (ARM-tierP step 2)
Pre-flight, written before the first dispatch. Gate: 2 calls, declared worst case $0.42. Run:
9 dispatches, declared $0.86, revised to $1.05 over 11 after the pre-run critic's F4 abolished
the four-arm stage. Both built from max_tokens per note (abc) with a ×2 routing margin per (bhq).
Combined declared worst case $1.47 against $2.281349592 of headroom. Everything fit; nothing
was deferred.
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-04 | S108 | D-20260804-16 ratification — adversarial review (nvidia/nemotron-3-ultra-550b-a55b, Together) + routed panel vote (moonshotai/kimi-k3, Modal) |
0.42 | 0.327597600 | per-response usage.cost |
Review RATIFY-A, vote B; the vote governs. One wasted body: the voter's first attempt hit finish_reason: length at cap 8,000 and cost $0.156402 — more than the entire adversarial review |
| 2026-08-04 | S108 | pre-run adversarial critic (z-ai/glm-5.2, Baidu) |
0.20 | 0.0619817352 | per-response usage.cost |
NEEDS-REDESIGN, 6 findings, 3 BLOCKING, all six accepted before any grading call. F4 alone added three dispatches and is why every PR1 figure comes from a block the constructed pair is absent from |
| 2026-08-04 | S108 | stage 0 — fork classification, A1 (P5, Ionstream) accepted; A2 failed on 4 bodies across qwen/qwen3.7-max and moonshotai/kimi-k3, then re-dispatched under amendment A7 to nvidia/nemotron-3-ultra-550b-a55b at cap 20,000 |
0.07 | 0.3211999 | per-response usage.cost |
$0.3016 of this bought nothing — four bodies, finish_reason: length, zero content characters, on two slugs. The substitute seat returned all 12 verdicts for $0.0157. Note (bhf), sixth firing |
| 2026-08-04 | S108 | stage 1a — G v H, P1/P2/P3 |
0.37 | 0.0854556 | per-response usage.cost |
3 of 3 first dispatch |
| 2026-08-04 | S108 | stage 1b — C v E, P1/P2/P3 |
0.37 | 0.2536298 | per-response usage.cost |
P2's good body was rejected by this run's own guard — one needless retry of the same seat and two reserve calls after it, $0.126113 wasted; the rejected body was recovered from disk and is what stage 1b reports. Note (bgw), second firing |
| 2026-08-04 | S108 | stage 2 — misattributed, P1/P3 accepted; P2 failed on 4 bodies | 0.24 | 0.2038645 | per-response usage.cost |
$0.1727 wasted; PR3 runs on two seats and the result says so |
| 2026-08-04 | S108 | «Певцы»'s six loci rendered twice (4,900 English words), the 12-site extraction, the FC3 contamination gate, the lead fork map, scoring and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S108 total, summed per-request over all 24 stored billed bodies: $1.253729249 — gate $0.327597600 plus run $0.926131649 — against a combined declared worst case of $1.47, i.e. 85%. Inside the estimate, and the estimate is the only thing that went right.
The key-usage cross-check closes twice, exactly. Gate: opening 58.83017868, settled close 59.15777628, delta 0.3275976 against a per-request sum of 0.3275976. Run: opening 59.15777628, settled close 60.083907929, delta 0.926131649 against 0.9261316492 — agreement to 2e-10. Note (bil) fired a third time, and at the size of the last call: the read taken immediately after the last gate body gave 59.01054108, short by exactly $0.1472352, and thirty seconds later was exact.
The line that matters in this row, and it is not the total. $0.756841 of $1.253729 — 60% of the
session, and $0.600439 of $0.926132, 65% of the run — bought nothing. Thirteen billed bodies returned no usable answer: ten hit
finish_reason: length with zero or near-zero content at their cap (note (bhf), now fired on
three different slugs, so it is not one seat's property), and three were spent re-trying a body
that was already correct because the runner's answer-line guard required a format the seat did not
emit (note (bgw)). The verification discipline caught none of this, because none of it is a
verification question: the raw bytes were on disk (note (bdt)) and the recovered body is in the
result. What would have caught it is one cheap probe per seat per task shape before the run —
which is rule (i) of note (bhf), written at S087 and still not paid for.
A between-session key-usage gap is recorded again and is not project spend. The pre-session read
was 58.83017868 against S107's settled close of 58.407119498, a gap of $0.423; the same shape
appeared S103→S104 (+0.392) and S104→S105 (+0.439). Per CLAUDE.md, per-request costs are primary
and the all-time key figure includes non-project spend; the day total below is the per-request sum.
2026-08-04 day total: $3.972379657 of $5.00 — eight sessions (S101 $0.252177609, S102 $0.356148955, S103 $0.39329705, S104 $0.534977746, S105 $0.132372, S106 $0.548190715, S107 $0.501486433, S108 $1.253729249). $1.027620343 headroom, 21% of the cap. S108 is the most expensive session in the project's history, and it is the wastage and not the evidence that made it so.
S107 — 2026-08-04 (UTC), E-20260804h-affect-halves (ARM-affect-reception step 2)
Pre-flight, written before the first dispatch. Fifteen dispatches planned: 1 critic
(max_tokens 24,000), 2 yardstick seats (12,000), 12 grading calls (8,000). Built from
max_tokens per note (abc) and carried with a ×2 routing margin, which is the correction note
(bhq) and S106's overrun row demand — note (abc) bounds the token count, not the price per token.
Declared worst case $1.60 against $2.782836025 of headroom on the day. The run fits; no
stage is deferred.
| date | session | what | pre-flight worst case | actual | method | notes |
|---|---|---|---|---|---|---|
| 2026-08-04 | S107 | pre-run adversarial critic, one pass (z-ai/glm-5.2, provider Alibaba) |
0.15 from max_tokens 24,000 with routing margin |
0.029674002 | per-response usage.cost |
NEEDS-AMENDMENT, 3 findings, 2 BLOCKING, all three accepted, 118s, 6,681 reasoning tokens. Both BLOCKING findings changed what the run may conclude. F1 found both frozen purpose specs paraphrasing a half of the definition under test — the mechanical check had passed them — and F2 showed the localization control could not fail, which added four sites |
| 2026-08-04 | S107 | stage 0 — the yardstick, A1 (P5 deepseek/deepseek-v4-pro, SiliconFlow) + A2 (qwen/qwen3.7-max, Alibaba) |
0.15 from max_tokens 12,000 with routing margin |
0.05207233144 | per-response usage.cost |
2 of 2 on first dispatch. Both seats shown the twelve Russian passages and no English of any kind, and they returned compatible accounts at 12 of 12 sites — the same criterion fired at 6 of 7 in S102 |
| 2026-08-04 | S107 | stage 1 — H1, 3 seats × 2 orderings (P1 OpenAI ×2, P2 Google ×2, P3 xAI ×2) | 0.65 from max_tokens 8,000 with routing margin |
0.2020218 | per-response usage.cost, re-summed by analysis/verify.py from the stored raw bodies |
6 of 6 on first dispatch, 12 answer lines each, zero retries. P1 $0.0173, P2 $0.0951, P3 $0.0896 |
| 2026-08-04 | S107 | stage 2 — H2, 3 seats × 2 orderings | 0.65 from max_tokens 8,000 with routing margin |
0.2177183 | per-response usage.cost |
6 of 6 on first dispatch, zero retries. P2 routed to Google AI Studio on both calls where stage 1 routed to Google |
| 2026-08-04 | S107 | «Хамелеон» rendered whole twice, the flattening arm, the three редакции diff, the contamination gate, site selection, scoring and verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S107 total, summed per-request over all 15 stored billed bodies: $0.501486433, against a declared
worst case of $1.60 — 31%. Fifteen dispatches, fifteen accepted bodies, zero retries, zero
seat failures, zero finish_reason: length. Nothing was wasted, which is the first such row since
S105.
The key-usage cross-check CLOSES to 0.000000000 — opening snapshot 57.905633065, closing 58.407119498, delta 0.501486433 against a per-request sum of 0.50148643344.
Note (bil) fired, and this row is the reason it exists. The key read at the moment the last body returned gave 58.328509598, a delta of 0.422876533 and a residual of −$0.0786 — a shortfall big enough to look like a billing discrepancy. Twenty seconds later the same endpoint returned the exact figure and the residual went to zero. A key-usage cross-check taken at the moment of the last response is not a cross-check, and a session that had reconciled immediately would have filed a false anomaly against its own ledger.
The pre-run critic cost $0.029674002 — 6% of the run — and it is the most valuable line here. Its two BLOCKING findings added the failure criterion that tested the design's deepest threat and re-registered what the run was allowed to conclude, both before a grading call was made.
2026-08-04 day total: $2.718650408 of $5.00 — seven sessions (S101 $0.252177609, S102 $0.356148955, S103 $0.39329705, S104 $0.534977746, S105 $0.132372, S106 $0.548190715, S107 $0.501486433). $2.281349592 headroom, 46% of the cap.
| 2026-08-06 | S120 | pre-run adversarial critic, one pass (x-ai/grok-4.5, provider xAI) | 0.25 from max_tokens 24,000 with routing margin | 0.036896400 | per-response usage.cost | NEEDS-AMENDMENT, 8 findings, 6 BLOCKING, all eight accepted, run against the design AND the completed coding before a single seat call. 34% of the session's spend and it changed the result. F1 established that P1 was close to unfalsifiable — the predictor was the coder and free variation was the residual class — and forbids the sentence the arm wanted to write. F2 found the coding protective; applying it moved the attributable count 10 → 13 and turned P1 from a pass into a FAIL. F5 refused the lead's "held on substance" scoring of two named predictions |
| 2026-08-06 | S120 | the blind coding run, 3 seats × 2 blocks of 14 items (P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash, P5 deepseek/deepseek-v4-pro) | 0.58 from max_tokens 4,000 × 6 with routing margin and the two-attempt cap | 0.072132180 | per-response usage.cost | 6 of 6 first dispatch, 84 of 84 cells, zero retries, zero seat failures, zero finish_reason: length. P1 $0.008974, P2 $0.057731, P5 $0.005428 — P2 is 80% of the seat spend for 33% of the cells. P5 routed to two different providers in two calls (DigitalOcean, CoreWeave) |
| 2026-08-06 | S120 | «Köyhää kansaa» span 1 re-rendered whole, the copy-text gate, the diff, the causal coding, the strata build and all verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S120 total, summed per-request over all 7 stored billed bodies: $0.109028580, against a declared worst case of $0.90 — 12%.
The key-usage cross-check CLOSES to 0.000000000 — opening snapshot 67.644145744, delta 0.109028580 against a per-request sum of 0.109028580. Note (bil) fired a third time and the rule worked: the snapshot taken at the moment the last body returned gave 67.722136894, a residual of −$0.031, and a re-read after the analysis was written gave the exact figure. A session that had reconciled immediately would have filed a false anomaly for the third time in sixteen sessions.
2026-08-06 day total: $1.431275507 of $5.00 — four sessions (S117 $0.614681164, S118 $0.665695363, S119 $0.041870400, S120 $0.109028580). $3.568724493 headroom, 71% of the cap.
| 2026-08-06 | S121 | pre-run adversarial critic, one pass (moonshotai/kimi-k3, provider Moonshot AI) | 0.36 from max_tokens 24,000 | 0.349863000 | per-response usage.cost | NEEDS-AMENDMENT, 3 BLOCKING + 9 notes, all three BLOCKING accepted in full, run against the frozen design, the sites and both lead renderings before any other call. B1: the judgement→site aggregation rule was unregistered and every primary sat on it. B2: no mismatched-statement control, so every recovery rate would have been uninterpretable — the control was added with its own failure criterion. B3: three sentences of the frozen design overclaimed, including a licence sentence saying seats recover a relation they are in fact handed. 31% of the session's spend and the best-value line in it. The declared line was $0.30 against a true max_tokens worst case of $0.36 — note (abc): the estimate was built low, and it is the line that overran |
| 2026-08-06 | S121 | yardstick, one seat shown the German and no English of any kind, plus one leak-screen regeneration (x-ai/grok-4.5, xAI) | 0.10 | 0.020698800 | per-response usage.cost | 4 of 12 statements flagged first pass, all regenerated once; 1 still flags and is adjudicated (du inside due; formal in a NULL statement). F3 does not fire. Role-direction check 12 of 12 correct |
| 2026-08-06 | S121 | parity screen, two seats, neither a grader (x-ai/grok-4.5; moonshotai/kimi-k3) | 0.10 | 0.063405400 | per-response usage.cost | 9 of 9 CLOSE/FORCE pairs judged to state the same facts; F4 does not fire, and per amendment A6 its passing shows little |
| 2026-08-06 | S121 | the blind grading run, 33 of 36 calls (S1 openai/gpt-5.6-terra, S2 google/gemini-3.6-flash, S3 deepseek/deepseek-v4-pro) | 1.40 from max_tokens 12,000 × 36 with routing margin | 0.530502577 | per-response usage.cost | INCOMPLETE. 3 of 36 calls outstanding (o2/S3 b3–b5) after repeated provider hangs past the 600s timeout. No figure is computed from a partial grading. Every returned body is stored and the runner reuses it without re-billing, so finishing costs three calls. A first dispatch of this stage was discarded whole when a body returned finish_reason: length; the cap was raised and a hard guard added |
| 2026-08-06 | S121 | «Michael Kohlhaas», both long dialogue scenes rendered twice (2,725 German words each way), the contamination gate, site selection, the textual census and all verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
| 2026-08-06 | S122 | the blind grading run finished, the outstanding 3 of 36 (S3 deepseek/deepseek-v4-pro, provider pinned) | 0.10 from max_tokens 12,000 × 3 with routing margin | 0.008616480 | per-response usage.cost, and the key-usage delta agrees exactly | All 33 stored bodies re-used without re-billing. All three returned from DigitalOcean in seconds once PROVIDER_ONLY was written into run.py — S121 pinned routing in a session that then ended, and the pin was never committed. Timeout halved to 300s |
| 2026-08-06 | S121 | CORRECTION ROW, late billing — one hung deepseek/deepseek-v4-pro call from S121's grading stage, billed by the provider and never delivered to this process | — | 0.012835977 | key-usage delta between S121's closing snapshot (69.451202818) and S122's opening one (69.464038795), with no call in between | This identifies the unknown component of S121's $0.170699695 residual. S121's true spend is $1.148005249, not $1.135169272; S121's own pages are left as S121 wrote them and the correction is carried here |
| 2026-08-06 | S122 | the whole Luther interview rendered a third time (1,421 German words) under the device the frozen design had excluded, plus the dependence measurement and all verification | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S121 total, per-request over the 38 stored billed bodies: $0.964469577, against a declared worst case of $1.90 (the frozen $1.20, raised by amendment A4 when the critic's block rule doubled the grading calls) — 51%.
The key-usage cross-check DOES NOT CLOSE, and the residual is explained rather than smoothed. Opening 68.316033546, closing 69.451202818, delta $1.135169272 against the stored-body sum of $0.964469577 — a residual of $0.170699695. Two known components and one unknown: the discarded first grading dispatch, nine bodies that returned and were deleted when the stage was re-dispatched (eight logged at $0.1021425; the ninth is the truncated body, whose cost was never printed), and any of the hung DeepSeek calls billed by the provider but never delivered to this process. Here the delta is the authority and the per-request sum is not — the reverse of this project's usual order, and stated so the row is not read the wrong way round. The session's true spend is $1.135169272.
deepseek/deepseek-v4-pro routed to EIGHT distinct providers in nine calls — Alibaba,
Cloudflare, DigitalOcean, GMICloud, Ionstream, SiliconFlow, StreamLake, Together — at $0.0016 to
$0.0071 for comparable blocks. config/models.md's standing routing caution is about cost; this
run adds latency: same slug, same block size, thirty seconds to past a ten-minute timeout
depending on the draw. Routing for that slug was pinned to the three providers that returned cleanly;
the slug is unchanged, so the seat is still the one config/models.md specifies. S122 addendum:
that pin lived only in the session that made it and was never committed — it is now PROVIDER_ONLY
in the experiment's run.py, which is where a resume can actually pick it up.
~~2026-08-06 day total: $2.566444779 — five sessions.~~ SUPERSEDED by the S122 line below, which adds S122's spend and the late-billed S121 correction row.
S122 total: $0.008616480, against a declared worst case of $0.10 — 9%. The cross-check closes exactly: opening 69.464038795, closing 69.472655275, delta $0.008616480, identical to the three stored bodies' per-request sum. It is exact because only three calls were made and none of them hung.
~~2026-08-06 day total: $2.587897236 of $5.00 — six rows across six sessions (S117 $0.614681164, S118 $0.665695363, S119 $0.041870400, S120 $0.109028580, S121 $1.148005249 incl. the late-billed correction, S122 $0.008616480). $2.412102764 headroom, 48% of the cap.~~ SUPERSEDED by the S123 block below.
S123 — 2026-08-06 (UTC), E-20260806f-persona-crossing (ARM-voice-crossing step 1, arm constituted)
| date | session | what | pre-flight worst case | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-06 | S123 | stage 0 — independent pre-run critic, FIRST ATTEMPTS, moonshotai/kimi-k3 |
inside the frozen $1.50 | 0.242587600 | key-usage delta 69.472655275 → 69.715242875 | Nothing was bought. One stored body: finish_reason: length, zero content characters, 16,000 of 16,000 completion tokens spent on hidden reasoning, $0.2414856. Two further attempts were killed by the wrapper before their responses returned and were billed anyway, $0.0011020 — which contradicts note (bhf)'s S116 record that a killed request bills zero. The slug had already failed this exact role at S106; note (bhf) rule (iii) says change the seat, and the design had not read it |
| 2026-08-06 | S123 | stage 0 — the critic that worked, nvidia/nemotron-3-ultra-550b-a55b (non-panel) |
— | 0.016529400 | per-response usage.cost |
NEEDS-AMENDMENT, 15 findings, 6 BLOCKING. One fifteenth of the failure it replaced. Five BLOCKING and three advisory accepted; BLOCKING 1 and 2 rejected with two lines of algebra and the figure they asked for reported anyway |
| 2026-08-06 | S123 | stage 1 — SOURCE profiles, first dispatch, 4 works × 4 seats |
0.15 from max_tokens 1,200 with routing margin |
0.082456250 | per-response usage.cost |
12 of 16 clean. google/gemini-3.6-flash returned length with the JSON truncated mid-object on 4 of 4 — 1,104–1,152 of 1,196 completion tokens were reasoning. Those four bodies, $0.0416715, are in runs/discarded/ and bought nothing |
| 2026-08-06 | S123 | stage 1 re-dispatch under amendment A9 — S2 alone raised to max_tokens 3,000, sized from the seat's measured appetite (note (bhq)), byte-identical payload |
0.09 from the raised cap | included below | per-response usage.cost |
4 of 4 returned stop |
| 2026-08-06 | S123 | stages 2–6 — CLOSE 16, PANEL translations 2, FLAT 8, PANEL profiles 8, recognition probe 8 |
0.86 from the raised caps with routing margin | included below | per-response usage.cost |
42 of 42 on first dispatch, zero retries, zero seat failures. 0 of 48 profile cells missing |
| 2026-08-06 | S123 | ten renderings — four works in four languages, R06 draft + R04 close each, plus two R14 flattenings; the four contamination measurements; the design, analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4) |
S123 total: $0.854134980, against a declared worst case of $1.70 (the frozen $1.50, revised
by amendment A9 when the S2 ceiling was raised) — 50%.
The cross-check does not close to the cent, and the residual is named rather than smoothed.
Opening 69.472655275, closing 70.326790255, delta $0.854134980. Stored bodies sum to
$0.836503580 (60 in runs/ plus 4 in runs/discarded/). Residual $0.017631400, all of it
the two killed pre-run critic attempts whose responses never reached this process. Here the
delta is the authority, and the residual is the session's most transferable instrument fact:
a killed request is not a free request.
28% of this session's spend bought nothing — $0.2415 for a zero-content critic body and $0.0417 for four truncated profiles — and both failures were foreseeable from note (bhf), which the design did not consult before choosing its seats.
2026-08-06 day total: $3.442032216 of $5.00 — seven rows across seven sessions (S117 $0.614681164, S118 $0.665695363, S119 $0.041870400, S120 $0.109028580, S121 $1.148005249, S122 $0.008616480, S123 $0.854134980). $1.557967784 headroom, 31% of the cap.
S124 — 2026-08-06 (UTC), E-20260806g-berman-negative-pole (ARM-berman-occurrence step 2, arm closed resolved)
| date | session | what | pre-flight worst case | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-06 | S124 | stage 0 — independent pre-run critic, one pass (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider Together) |
0.10 from max_tokens 16,000 |
0.032205000 | per-response usage.cost, and the key-usage delta agrees exactly |
NEEDS-AMENDMENT, 10 findings, 5 BLOCKING. Eight accepted, one taken as a declared limit, one rejected with two lines of algebra — the second consecutive session in which a critic's arithmetic BLOCKING finding was wrong while its design findings were worth many times the fee. Its finding 1 is the session's best money: F4, the gate separating read as low from read as bad, could not fire as written, because the prompt asks each seat what drove its most extreme of three codes. Note (bjx) |
| 2026-08-06 | S124 | stage 2 — the coding run, 12 sites × 3 seats (P1 gpt-5.6-terra, P3 grok-4.5, P5 deepseek-v4-pro) |
0.645 from max_tokens 2,500 × 36 with the 1.5× routing margin |
0.365426787 | per-response usage.cost |
36 of 36 seats returned, zero seat failures, zero retries, 459 codes, 153 cells. 57% of the declared worst case. $0.0128132 of it bought nothing: two deepseek-v4-pro bodies at finish_reason: length with zero content characters and ~9,000 tokens of hidden reasoning, from GMICloud then SiliconFlow, while the same slug at DigitalOcean had returned clean for $0.00094. Note (bhf) with the provider as the variable. P2 (gemini-3.6-flash) was dropped from the panel before dispatch on this project's own record (6,000-token cap at S118; 80% of seat spend for 33% of cells at S120) and P5 added — the whole coding stage cost less than S118's, on the same site count |
| 2026-08-06 | S124 | declared mid-run amendment A10 — routing pinned to DigitalOcean for that slug, run stopped and restarted |
— | included above | — | The 5 completed bodies were re-used without re-billing, which is what the .done.json cache is for. The pin was written into the committed call.py, not into the result page alone — note (bjt), fired and honoured. Post-pin the seat returned in 4–20s at ~$0.001 a call |
| 2026-08-06 | S124 | stage 3 — F4's second independent classifier (nvidia/nemotron-3-ultra-550b-a55b, non-panel, took no coding seat) |
0.05 from max_tokens 4,000 |
0.010059600 | per-response usage.cost |
22 of 22 classified. Disagrees with the lead on 1 line of 22 (4.5%), and the one disagreement is the boundary case the frozen rubric was written to adjudicate. Bought by critic finding 3, which refused to let an unreproducible lead judgment gate the primary |
| 2026-08-06 | S124 | CORRECTION ROW, killed call — the request in flight when run_coding.py was killed to apply A10 |
— | 0.016075170 | key-usage delta, the whole of the session's residual | Note (bjv), fired a second session running. The per-request sum is $0.407691387; the key-usage delta over the session is $0.423766557. The residual is the kill, not a mystery — and it is exactly the cost note (bjv) says a kill is not free |
| 2026-08-06 | S124 | R21 minted; Verga «Cavalleria rusticana» ¶50–81 rendered a fourth time under it, with the rule set frozen before translating and the log frozen before the design; the contamination measurement, the site cutter and its three assertions, the machine limb, scoring and all verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). The arm the entire primary rests on cost nothing |
Session total $0.423766557 (key-usage delta 70.627810778 → 71.051577335), against a declared
worst case of $0.95. 45% of the declared figure, and the two lines that overran nothing were the
ones built from max_tokens per note (abc).
2026-08-06 day total: $3.865798773 of $5.00 — eight rows across eight sessions (S117 $0.614681164, S118 $0.665695363, S119 $0.041870400, S120 $0.109028580, S121 $1.148005249, S122 $0.008616480, S123 $0.854134980, S124 $0.423766557). $1.134201227 headroom, 23% of the cap.
S125 — 2026-08-07 (UTC), E-20260807-narrator-unknown (ARM-voice-crossing step 2, arm closed resolved)
| date | session | what | pre-flight worst case | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-07 | S125 | stage 0 — independent pre-run critic, one pass (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider Together) |
0.12 from max_tokens 16,000 |
0.020611200 | per-response usage.cost |
NEEDS-REDESIGN, 9 findings, 3 BLOCKING, all three accepted — two outright, one in a modified form argued on the page. Its BLOCKING 1 is the session's best money: the registered recognition gate measured only whether a seat could name the work from the English, and the confound needs identification on both sides. The amendment added a source-side probe, and that probe is where the run's only true recognition hit came from. Its BLOCKING 2 put back a control the design had dropped to save ten calls against $5.00 of headroom, which was not a reason |
| 2026-08-07 | S125 | stages 1–3, 5 — 48 persona profiles (SOURCE 16, CLOSE 16, FLAT 8, PANEL 8) over four seats (P1 gpt-5.6-terra, P2 gemini-3.6-flash, P3 grok-4.5, P5 deepseek-v4-pro, provider-pinned) |
0.87 from max_tokens 1,200 × 36 and 3,000 × 12, with the 2× routing margin |
0.319… (included in the total below) | per-response usage.cost |
48 of 48 returned, 0 of 48 cells missing, no retries and no truncation. S2's max_tokens 3,000 was applied in advance from S123's measured hidden-reasoning appetite rather than rediscovered — the first time this project has pre-empted note (bhq) instead of paying for it |
| 2026-08-07 | S125 | stage 4 — two PANEL translations (mistralai/mistral-medium-3-5, non-panel, so no profiling seat scores its own prose) |
0.04 from max_tokens 2,500 × 2 |
included below | per-response usage.cost |
Reinstated by critic BLOCKING 2. Matched its source at 8 of 8, the same as the lead's |
| 2026-08-07 | S125 | stages 6–7 — 16 recognition calls, 8 from the English (G2a) and 8 from the source (G2b) |
0.10 → 0.33 after amendment A10 | included below | per-response usage.cost |
G2a 0 of 8, G2b 1 of 8. The gate the whole run turns on |
| 2026-08-07 | S125 | WASTE ROW — one discarded body, recogsrc-karr-S1 at finish_reason: length with zero content characters |
— | 0.004329000 | per-response usage.cost; it is the whole of the key-usage residual |
Note (bhq), and the cap was the variable, not the seat. 600 of 600 completion tokens were reasoning_tokens; S1 had performed this role cleanly at S123 on English text and the new thing was source-language input, so note (bhf) rule (iii) — change the seat — does not apply and the cap was raised instead (amendment A10, max_tokens 600 → 2,500, written into run.py). The dead body is kept in runs/discarded/ |
| 2026-08-07 | S125 | Ten renderings — four R06 drafts, four R04 revisions with frozen logs, two R14 flattenings; the contamination search and measurement; all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). The four texts the entire primary rests on cost nothing, and two of them exist in English nowhere else |
Session total $0.425134240 (key-usage delta 71.060119435 → 71.485253675), against a declared
worst case of $2.23 after amendment A10 ($1.78 before it). 19% of the declared figure. The
per-request sum over the 67 stored bodies is $0.420805240; the residual $0.004329 is
exactly the one discarded body in runs/discarded/, which the analysis glob does not reach. The
reconciliation is exact to 1e-9 and there is no mystery line this session.
S126 — 2026-08-07 (UTC), E-20260807b-botchan-register (ARM-ja-register step 1)
Pre-flight, written before dispatch, built from max_tokens and the worst plausible provider
(note (abc), and the routing caution in config/models.md): 12 rating calls $0.22, one pre-run
critic at 16,000 max_tokens $0.17, one re-dispatch allowance $0.11 — declared worst case $0.50
against $4.574866 of headroom.
| date | session | item | pre-flight (USD) | actual (USD) | source | note |
|---|---|---|---|---|---|---|
| 2026-08-07 | S126 | critic pass 1 (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider Together) |
0.17 from max_tokens 16,000 |
0.059765400 | per-response usage.cost |
finish_reason: length, ZERO content characters — 16,000 of 16,000 completion tokens on hidden reasoning. Findings salvaged from the reasoning field, degraded. Note (bhf), fifteenth firing, and it falsifies (bhf)'s own S123 remedy: this is the slug that note recommended |
| 2026-08-07 | S126 | critic pass 2 (qwen/qwen3.7-max, non-panel reserve, provider Alibaba) |
not in the pre-flight; taken from the re-dispatch allowance | 0.049856470 | per-response usage.cost |
The best money in the session. NEEDS-AMENDMENT, six findings, all applied; its BLOCKING 1 found that the source-side gate G1 re-imported the very content confound the first pass's advisory had just made the design demote on the English side |
| 2026-08-07 | S126 | stage 1 — 4 source-side register calls (18 Japanese loci each; P1, P2, P3, P5 provider-pinned) | included in the 0.22 | 0.027772260 | per-response usage.cost |
4 of 4 stop, 72 of 72 cells |
| 2026-08-07 | S126 | stage 2 — 8 English register calls (two balanced blocks of 27 shuffled items per seat) | included in the 0.22 | 0.078779160 | per-response usage.cost |
8 of 8 stop, 216 of 216 cells, no retries, no truncation, no re-dispatch — the S2 cap was again set in advance from the measured appetite rather than rediscovered |
| 2026-08-07 | S126 | Three renderings (R06, R04 with its frozen log, R21 control), the whole-chapter close reading of both texts, the contamination measurement, all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). The arm the primary rests on — that the low pole is reachable in English narration — is a $0 artifact |
S126 total: $0.216173295, 43% of the declared worst case, and the key-usage delta agrees to 5e-9.
2026-08-07 day total: $0.641307535 of $5.00 — two sessions (S125 $0.425134240, S126 $0.216173295). $4.358692465 headroom, 87% of the cap. $4.574865760 headroom, 91% of the cap.
S127 — 2026-08-07 (UTC), E-20260807c-two-hands (ARM-two-hands step 1)
| date | session | what | pre-flight worst case (USD) | actual (USD) | source | note |
|---|---|---|---|---|---|---|
| 2026-08-07 | S127 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider Together) |
0.15 from max_tokens 16,000 |
0.034219800 | per-response usage.cost |
stop, 11,576 characters — note (bhf) did NOT fire on the slug that produced a zero-content body at S126 at this same cap. NEEDS-REDESIGN, 5 BLOCKING, 11 ADVISORY; ten amendments accepted, two overruled. Its BLOCKING 1 struck the registered primary and its BLOCKING 3 bought the third seat block that carries the replacement primary |
| 2026-08-07 | S127 | stage 1 — 4 translation hands (P1 ×2, P5 ×2, unbriefed, temperature 1.0) | 0.11 from max_tokens 3,000 |
0.036245640 | per-response usage.cost |
4 of 4 stop, 817–848 words, all inside the registered 600–1300 band; FC2 silent |
| 2026-08-07 | S127 | stage 2 — 9 seat blocks as dispatched (3 seats × 3 blocks) | 0.84 from max_tokens 6,000 |
0.518635100 | per-response usage.cost |
6 of 9 stop. Three length bodies, two of them from one seat, are the waste row below |
| 2026-08-07 | S127 | WASTE ROW — three discarded bodies, seat-S3-b1 (zero content characters at cap), seat-S3-b3 (199 characters), seat-S1-b3 (truncated) |
— | 0.303697 | per-response usage.cost; included in the stage-2 figure above |
Note (bhf), sixteenth firing, 41% of the session's spend. The two failures on one seat were treated by rule (iii) — the seat was changed — and the single failure on a seat that had already cleared this exact format twice was treated by note (bhq) — the cap was raised. Both re-dispatches returned stop. Dead bodies preserved in runs/discarded/, never overwritten (note (bhd)) |
| 2026-08-07 | S127 | stage 2 re-dispatch — 4 blocks (seat-S1-b3r at cap 12,000; seat-S4-b1/b2/b3 on qwen/qwen3.7-max, non-panel, amendment A11) |
taken from the re-dispatch allowance | 0.143639750 | per-response usage.cost |
4 of 4 stop. The fourth was seat-S4-b2, added unprompted so the replacement seat would carry control data like the seat it replaced |
| 2026-08-07 | S127 | WASTE ROW — one lost body, a deepseek translation call billed after a two-minute foreground shell timeout killed the runner mid-call |
— | 0.003331230 | key-usage residual; it is the whole of it | The call completed server-side and its body was never written. Reconstructed from the residual and ledgered rather than left as drift |
| 2026-08-07 | S127 | Two lead renderings (R06 draft with the registered 16-locus prediction, R04 close with its frozen log), the whole-chapter reading of the Hungarian, the comparator extraction, the contamination measurement, all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). One of the two is, as far as this project could establish, the only English rendering of this chapter made from the Hungarian since 1900 |
S127 total: $0.736071520, 67% of the declared ceiling of $1.10 (raised from $0.80 mid-run by amendment A3, with the reason written on the design). Reconciliation exact once the lost body is counted.
2026-08-07 day total: $1.377379055 of $5.00 — three sessions (S125 $0.425134240, S126 $0.216173295, S127 $0.736071520). $3.622620945 headroom, 72% of the cap.
S128 — 2026-08-07 (UTC), E-20260807d-marking-work (ARM-marking-work step 1)
Key-usage snapshots. Opening 73.423507026, closing 74.478091001, delta 1.054583975 against a per-request sum of 1.054583975 — exact to 0.000000000, on 23 bodies across seven providers. Both snapshots written to disk before being read (note (bco)).
| date | session | stage | pre-flight worst case | actual (USD) | method | note |
|---|---|---|---|---|---|---|
| 2026-08-07 | S128 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider BaseTen) |
0.15 from max_tokens 16,000 |
0.0206148 | per-response usage.cost |
stop, 12,941 characters. NEEDS-REDESIGN, 5 BLOCKING, 7 ADVISORY, all twelve accepted, two remedies overruled. Its BLOCKING 1 would have voided the primary on its own — the Stage A prompt was supplying the standing the seats were meant to read off the grammar. The best $0.02 in the session |
| 2026-08-07 | S128 | stage A — 3 source-side classification calls (P1, P2, P5; 56 Japanese utterances each) | 0.13 from max_tokens 6,000 |
0.2778331 | per-response usage.cost |
56 of 56 items parsed on all three seats. P2 needed a cap raise (A12) and then a two-block split (A13) |
| 2026-08-07 | S128 | stages C1/C2 — relation statements (P3) and the leak screen (P1) | 0.09 | 0.0163874 | per-response usage.cost |
Both stop. The leak screen fired at 9 of 12 and withheld the primary a second time |
| 2026-08-07 | S128 | stage D — 2 parity calls (P1, P5) | 0.05 | 0.0080076 | per-response usage.cost |
Both stop, 12 of 12 pairs. F4 fired at 12 of 12 — because amendment A6 made the probe the treatment |
| 2026-08-07 | S128 | stage E — 6 grading blocks (P4, qwen/qwen3.7-max, z-ai/glm-5.2) |
0.42 from max_tokens 6,000 |
0.2500050 | per-response usage.cost |
156 of 156 cells, no missing data. Two blocks needed a cap raise (A16); both re-dispatches returned stop |
| 2026-08-07 | S128 | stage F — recognition (qwen/qwen3.7-max) |
0.06 | 0.0447515 | per-response usage.cost |
stop. Work unknown, author unknown, not confident. The second recogniser never returned a body — see the waste row |
| 2026-08-07 | S128 | WASTE ROW — seven discarded bodies | — | 0.436984615 | per-response usage.cost; included in the stage figures above |
Note (bhf), and it is 41.4% of the session. Two bodies returned zero content characters with the whole completion budget on hidden reasoning; one returned 733 characters of 24,000; two were partial. Note (bhq) — raise the cap where the seat has cleared the format — was applied four times and worked three. Note (bhf) rule (iii) — change the seat — was applied once, on recognition, and the replacement model failed the same way at the same cap, which is new. Every dead body preserved in runs/discarded/, none overwritten (note (bhd)) |
| 2026-08-07 | S128 | Two lead renderings of Ōgai's «Saigo no ikku» second half (R06 draft 2,573 words, R04 close 2,606 words with the frozen log), a third rendering of «Takasebune»'s opening for the contamination cell, the whole-story reading in Japanese, the comparator extraction and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Corrected at hand-off: this is not the lead's first rendering of this story. S027 rendered §4 entire on 2026-07-27; the new artifacts are re-filed R04-v2 / R06-v2 and the self-overlap is measured on them (218 / 74 / 35 / run 24 against 160 / 27 / 12 / 24 for two independent published hands). The gate that should have caught it looked only at published comparators |
S128 total: $1.054583975 against a declared $1.10 — 96% of the declaration, the closest this project has run to its own ceiling, and it did not cross it.
2026-08-07 day total after S128: $2.431963030 of $5.00 — four sessions (S125 $0.425134240, S126 $0.216173295, S127 $0.736071520, S128 $1.054583975).
S129 — 2026-08-07 (UTC), E-20260807e-sense-tradeoff (ARM-sense-tradeoff step 1)
Pre-flight, written before dispatch: $0.65, raised to $1.10 by amendment A11 after the pre-run critic's three BLOCKING findings added a source-blind pass, an unbriefed independent hand, and a dose pilot. Worst case built from the caps the requests actually permit (note (abc)), never from expected output.
Key-usage snapshots. Opening 74.485572001, closing 74.837369916, delta 0.351797915 against a per-request sum of 0.351797915 — exact to 0.000000000, on 17 bodies across six providers. Both snapshots written to disk before being read (note (bco)).
| date | session | stage | pre-flight worst case | actual (USD) | method | note |
|---|---|---|---|---|---|---|
| 2026-08-07 | S129 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider BaseTen) |
0.16 from max_tokens 16,000 |
0.0276492 | per-response usage.cost |
stop, 13,626 characters. NEEDS-REDESIGN, 4 BLOCKING, 6 ADVISORY, all ten accepted, one remedy overruled. Its BLOCKING 1 bought the source-blind naturalness pass the primary is now read on; its BLOCKING 2 bought the independent hand that carries §6 of the result. The best $0.03 in the session |
| 2026-08-07 | S129 | independent hand — 2 translations (mistralai/mistral-medium-3-5, non-panel so no seat scores its own prose; amendment A2) |
0.06 from max_tokens 4,000 |
0.0227145 | per-response usage.cost |
Both stop, 6 of 6 segments each, 770 and 864 words. Dispatched R08 before R07 (A10), the reverse of the lead's order |
| 2026-08-07 | S129 | dose pilot — 1 non-panel seat, 18 control items (qwen/qwen3.7-max; amendment A4) |
0.05 | 0.035251025 | per-response usage.cost |
stop, 18 of 18. Δ(accuracy) 5.333 and Δ(naturalness) 2.333 against a 0.50 gate — the critic's BLOCKING 3 refuted before the main run. Pilot scores discarded from every reported figure |
| 2026-08-07 | S129 | stage 1 — 9 scoring blocks (3 seats × 3 balanced blocks of 14, four senses, source visible) | 0.52 from max_tokens 10,000 |
0.19001868 | per-response usage.cost |
9 of 9 stop, 126 of 126 cells, no retries and no truncation |
| 2026-08-07 | S129 | stage 2 — 3 source-blind blocks (42 English-only passages per seat, naturalness only) |
0.21 from max_tokens 12,000 |
0.076173630 | per-response usage.cost |
126 of 126 after one re-dispatch. This is the stage the primary is read on |
| 2026-08-07 | S129 | WASTE ROW — one discarded body, blind-J3 returning 41 of 42 rows |
— | 0.008961870 | per-response usage.cost; included in the stage-2 figure above |
F2 fired. The body was complete and well-formed — it closed cleanly and was not truncated — the seat simply omitted one item. So neither note (bhf) rule (iii) nor note (bhq) applies: this is not a cap failure and not a seat failure, it is an omission, and the registered remedy (re-dispatch the block once) returned 42 of 42. Dead body preserved in runs/discarded/, never overwritten (note (bhd)). 2.5% of the session — the lowest waste rate this week |
| 2026-08-07 | S129 | Three lead renderings of Lu Xun 〈孔乙己〉 (R06, R07 fluency, R08 resistancy, 1,201 Chinese characters each, all three logs frozen and committed before the evaluation was designed), the two operator-derived control arms, the whole-story reading in Chinese, the contamination measurement against two comparators, and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). The project had never rendered this story before — checked mechanically before the source was chosen, which is S128's note (bhb) lesson applied as a gate rather than discovered at hand-off |
S129 total: $0.351797915 against a declared $1.10 — 32% of the declaration.
2026-08-07 day total: $2.783760945 of $5.00 — five sessions (S125 $0.425134240, S126 $0.216173295, S127 $0.736071520, S128 $1.054583975, S129 $0.351797915). $2.216239055 headroom, 44% of the cap.
S130 — 2026-08-07 (UTC), E-20260807f-carriage-or-strangeness (ARM-sense-overlap step 1)
Opening key snapshot 75.195884816, closing 75.548107416, delta 0.352222600 against a per-request sum of 0.393192500. The key delta is the smaller figure, which is the ordinary settling direction (note (bcx)); per-request sums are ledgered, per this page's stated method.
| date | session | what | pre-flight worst case (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-08 | S137 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.06 from max_tokens 16,000 |
0.036928200 | per-response usage.cost |
stop, 6,400 characters. NEEDS-REDESIGN, 3 BLOCKING, 7 ADVISORY; all ten accepted, none overruled — the tenth session running. Its BLOCKING 3 replaced the gate's checker (which is why note (bks) exists); its ADVISORY 8 made the floor rule symmetric and removed the design's last degree of freedom |
| 2026-08-08 | S137 | stage A0 — the LOW-P arm, placeless-constrained low rendering of 17 IT sites (mistralai/mistral-medium-3-5, the model that wrote LOW-A) |
0.05 from max_tokens 6,000 |
0.004473000 | per-response usage.cost |
stop, 17 of 17 |
| 2026-08-08 | S137 | stage A — G3′ probe call, DEAD (qwen/qwen3.8-max, non-panel, new lab) |
— | 0.087184000 | per-response usage.cost |
WASTE. finish_reason: length, zero content, 14,000 reasoning tokens on a 1,586-token prompt. The design's registered live probe caught it on the smallest call: one call at $0.087 instead of five at ~$0.50, and the fallback was already named in the frozen design. Note (bks). Body in runs/discarded/ |
| 2026-08-08 | S137 | stage A — G3′, source-relative parity + locatability, 6 surviving calls, 197 items, 6 arms, blind (nvidia/nemotron-3-ultra-550b-a55b, the registered fallback) |
0.50 from max_tokens 20,000 |
0.185977000 | per-response usage.cost sum |
197 of 197 returned; 5 of 5 planted content errors named. Includes the JA re-dispatch: the single Japanese call died at length with zero content ($0.073362, below), and the same model returned 59 Italian items on the same cap in 5,322 characters — the language drove the reasoning chain, not the batch size; split in two it returned both halves for $0.017156 + $0.016297 |
| 2026-08-08 | S137 | stage A — the JA call, DEAD (nvidia/nemotron-3-ultra-550b-a55b) |
— | 0.073362000 | per-response usage.cost |
WASTE. finish_reason: length, zero content, 20,000-token cap. Body in runs/discarded/. Not included in the 0.185977 line above, which sums the six surviving calls only; the two rows are separate and both are in the session total |
| 2026-08-08 | S137 | stage B — REG coding, 3 seats × 40 items, 5 arms at the 8 IT narration sites |
0.25 from max_tokens 9,000 |
0.053993390 | per-response usage.cost sum |
stop on all three; 120 of 120 cells, no missing data. J1 $0.02237025 · J2 $0.028050 · J3 $0.00357309 |
| 2026-08-07 | S130 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel, provider Together) |
0.06 from max_tokens 16,000 |
0.0297342 | per-response usage.cost |
stop, 22,122 characters. NEEDS-REDESIGN, 8 BLOCKING, six accepted, two remedies overruled. Its BLOCKING (d) bought Stage C, which two of the four findings are read on, and its BLOCKING (c) caught that all six odd prepositional government edits were calques of the French preposition at the site — a third of the STRANGE factor, replaced before dispatch. The best $0.03 in the session, for the third session running |
| 2026-08-07 | S130 | G5 device census (qwen/qwen3.7-max, non-panel; French only, no English) |
0.02 → 0.06 after the cap rise | 0.030699175 | per-response usage.cost |
stop, 3,360 characters. 14 of 18 of the lead's catalogued devices independently named, 0.778 against a 0.60 bar |
| 2026-08-07 | S130 | parity check (mistralai/mistral-medium-3-5, non-panel, 24 pairs) |
0.04 | 0.013158 | per-response usage.cost |
stop. 18 of 18 formal-edit pairs judged propositionally equivalent and 6 of 6 planted WRONG errors caught by name — a positive control at ceiling, and the fact that makes G1's failure a finding rather than a defect |
| 2026-08-07 | S130 | stage A — 3 seats × 30 items, source present, style-correspondence + accuracy |
0.30 from max_tokens 10,000 |
0.103059760 | per-response usage.cost |
90 of 90 cells after one re-dispatch (see the waste row) |
| 2026-08-07 | S130 | stage B — 3 seats × 30 items, source absent, perceived-source-carriage + naturalness |
0.26 | 0.050229440 | per-response usage.cost |
90 of 90, 3 of 3 stop, no retries |
| 2026-08-07 | S130 | stage C — 3 seats × 30 items, source present, perceived-source-carriage alone (amendment A1) |
0.15 | 0.066040200 | per-response usage.cost |
90 of 90, 3 of 3 stop. The stage P5 and P6 are read on, and it did not exist before the critic pass |
| 2026-08-07 | S130 | WASTE ROW — two discarded bodies | — | 0.100271725 | per-response usage.cost; additional to the stage figures above, which count kept bodies only |
25.5% of the session, and the two failures are different animals. census attempt 1 returned zero content characters at max_tokens 4,000 with the whole budget on hidden reasoning — note (bhf), and note (bhq)'s remedy (raise the cap where the format is plausible) worked at 14,000. stageA-J2 attempt 1 was complete and well-formed and simply omitted one of 30 items — the exact shape S129 recorded, so neither (bhf) nor (bhq) applies and the registered F2 remedy returned 30 of 30. Both dead bodies preserved in runs/discarded/, never overwritten (note (bhd)) |
| 2026-08-07 | S130 | Two lead renderings of Schwob «Paroles de Monelle» (R06 draft and R04 close, 519 words each, both logs frozen and committed before any comparator was read), the eighteen-device formal census, the three operator-derived arms, the WRONG control arm, the whole-book reading in French, the contamination measurement against three comparator cells, and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). The project had never rendered Schwob before — checked mechanically before the source was chosen (note (bhb) applied as a gate) |
S130 total: $0.393192500 against a declared $1.20 — 33% of the declaration. The declaration was
raised from $0.90 to $1.20 by amendment A11 before dispatch, to pay for the stage the critic's
BLOCKING (d) required; headroom at that moment was $2.216239055.
2026-08-07 day total: $3.176953445 of $5.00 — six sessions (S125 $0.425134240, S126 $0.216173295, S127 $0.736071520, S128 $1.054583975, S129 $0.351797915, S130 $0.393192500). $1.823046555 headroom, 36% of the cap.
S131 — 2026-08-07 (UTC), E-20260807g-indirect-deference (ARM-ja-register step 2)
Opening key snapshot 75.614366416, closing 76.432350471, delta 0.817984055 against a per-request sum of 0.817984055 — exact to 1e-9, the second such reconciliation in three sessions.
| date | session | what | pre-flight worst case (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-07 | S131 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.06 from max_tokens 16,000 |
0.022087200 | per-response usage.cost |
stop, 6,739 characters. NEEDS-REDESIGN, 5 BLOCKING. Its BLOCKING 3 and 4 took the three manipulated arms out of the lead's hands entirely — the design as frozen had the lead writing every comparison arm in every cell, including G1's own control — and its BLOCKING 2 caught that the rating scale's object is ambiguous in reported speech. The best $0.02 in the session, for the fourth session running |
| 2026-08-07 | S131 | stage 0b — arm generation, IND + INDC + DIRF for 17 utterances (google/gemini-3.6-flash, panel P2 but not a seat in this run; amendment A11) |
0.10 | 0.161670000 | per-response usage.cost |
3 of 3 stop. The amendment A1 stage; without it the run would have been the lead marking its own homework |
| 2026-08-07 | S131 | stage 1 — source-side gate, 3 seats × 24 Japanese items | 0.09 | 0.030353870 | per-response usage.cost |
3 of 3 stop. G2 passes at +4.0896 (4.471 marked against 0.381 unmarked), 3 of 3 seats |
| 2026-08-07 | S131 | stage 2 — the rating stage, 3 seats × 3 blocks × 80 English items | 0.36 from max_tokens 8,000 |
0.076334630 | per-response usage.cost |
240 of 240 cells, 9 of 9 stop, no re-dispatch. G1 passes at +2.0588 |
| 2026-08-07 | S131 | stage 3 — parity gate (mistralai/mistral-medium-3-5, non-panel, 34 real + 4 planted) |
0.05 | 0.010612500 | per-response usage.cost |
stop. G3 FAILS at 27 of 34 = 0.7941 against 0.80 — one item — with the planted-error control at 4 of 4. F1 fires and the primary is withheld; the bar is not weakened after firing |
| 2026-08-07 | S131 | stage 4 — device census (qwen/qwen3.7-max, non-panel) |
0.07 | 0.026799275 | per-response usage.cost |
stop. Descriptive only (amendment A5): 22 of 27 devices classified as unable to survive reporting, against a measured cost of 0.296 of a scale point |
| 2026-08-07 | S131 | WASTE ROW — six discarded bodies | — | 0.490126180 | per-response usage.cost; additional to the stage figures above |
59.9% of the session, and it is ONE failure repeated. The IND generation prompt returned zero content with the whole budget on hidden reasoning — note (bhf) — under z-ai/glm-5.2 at caps 14,000 and 30,000 and under moonshotai/kimi-k3 at 14,000 ($0.212442 for nothing), while the same models completed the INDC and DIRF prompts in the same batch cleanly. Note (bhq)'s cap-raise remedy failed and so did the standing change-the-seat remedy. The two clean glm bodies were discarded deliberately rather than mixed with a second hand's — that would have confounded arm with author. New note (bkh): the failure is prompt-shaped, not seat-shaped. All six preserved in runs/discarded/, never overwritten (note (bhd)) |
| 2026-08-07 | S131 | Two lead renderings of Botchan ch. 1 ¶8–10, 14, 20 (R06 draft frozen and committed at d9f5bab before the R04 revision, per R04 §2a), the translator's log, the eighteen-item locus build, the contamination measurements, and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Note (bhb) applied as a gate before translating: the project had not rendered these five paragraphs before |
S131 total: $0.817984055 against a declared $0.90 — 91% of the declaration, the highest
utilisation this project has recorded, and 60% of it bought nothing. The declaration was raised
from $0.75 to $0.90 by amendment A8 before dispatch to pay for the generation stage the critic
required; headroom at that moment was $1.823046555.
2026-08-07 day total: $3.994937500 of $5.00 — seven sessions (S125 $0.425134240, S126 $0.216173295, S127 $0.736071520, S128 $1.054583975, S129 $0.351797915, S130 $0.393192500, S131 $0.817984055). $1.005062500 headroom, 20% of the cap.
S132 — 2026-08-08 (UTC), E-20260808a-set-forms (ARM-two-hands step 2)
Opening key snapshot 76.432350471 — exactly config/budget.md's closing figure for S131, so
the unexplained $0.533 gap RS-20260807c §9 reported did not recur. Closing 77.165458791, delta
0.733108320 against a per-request sum of 0.733108320 — exact to 1e-9, the third such
reconciliation in four sessions.
| date | session | what | pre-flight worst case (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-08 | S132 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.06 from max_tokens 16,000 |
0.025595400 | per-response usage.cost |
stop, 9,680 characters. NEEDS-REDESIGN, 9 BLOCKING, eight amendments, none overruled. Its BLOCKING 5 moved the primary off a 1900 × 2026 pair whose era gap would have produced the predicted effect on its own; its BLOCKING 4 bought a second, within-paragraph PLAIN control; its BLOCKING 8 bought the second draw per model that the run's main finding depends on entirely. The best $0.03 in the session, for the fifth session running |
| 2026-08-08 | S132 | stage 1 — four unbriefed renderings of the whole chapter (P5 ×2, qwen/qwen3.7-max ×2 after amendment A10) |
0.13 from max_tokens 12,000 |
0.064425938 | per-response usage.cost |
4 of 4 stop, 30 of 30 paragraphs each |
| 2026-08-08 | S132 | stage 2 — the seat blocks, 3 seats × 4 blocks, 144 locus judgments per seat | 0.86 from max_tokens 12,000 / 4,000 |
0.485661050 | per-response usage.cost |
12 of 12 stop after one re-dispatch. Quote verification 0.98–1.00; REPEAT fired on no seat; WRONG 10 of 12 |
| 2026-08-08 | S132 | WASTE ROW — five discarded bodies | — | 0.157425932 | per-response usage.cost; additional to the stage figures above |
21.5% of the session. Note (bhf) fired twice in one stage in both its documented shapes: P5 hit a 4,000 cap with content → cap raised, seat kept, both re-dispatches clean; P4 returned zero content twice at $0.127563 → seat changed per the note's own rule. A third shape the notes do not cover: stage2-A-P3 returned zero content with finish_reason: stop, and a straight re-dispatch of the identical prompt to the identical seat returned 8,334 characters. All five preserved in runs/discarded/, never overwritten (note (bhd)) |
| 2026-08-08 | S132 | Two lead renderings of «Szent Péter esernyője» ch. 2 (R06 draft frozen at b583be6 before the R04 revision, per R04 §2a), both translator's logs, the 49-locus classification, the contamination measurements, and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination measured before the design was written: 12 shared 7-grams, 0 twelve-grams, longest run 11 tokens |
S132 total: $0.733108320 against a declared $1.10 — 67% of the declaration. The ceiling was
raised from $0.80 to $1.10 by amendment A8 before dispatch, to pay for the second draw per
model and the fourth block the critic required; headroom at that moment was $4.974405.
2026-08-08 day total: $0.733108320 of $5.00 — one session (S132 $0.733108320). $4.266891680 headroom, 85% of the cap.
S133 — 2026-08-08 (UTC), E-20260808b-discordant-marking (ARM-marking-work step 2)
Opening key snapshot 77.420865591, which is $0.255406800 above config/budget.md's closing
figure for S132 (77.165458791). The key is used outside this project and per-request costs are
primary; the gap is recorded, not explained. Closing 78.272708963, delta 0.851843372
against a per-request sum of 0.851843373 — exact to 1e-9, the fourth such reconciliation in
five sessions.
| date | session | what | pre-flight worst case (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-08 | S133 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.06 from max_tokens 16,000 |
0.049166 | per-response usage.cost |
stop, 4,096 characters. NEEDS-AMENDMENT, 2 BLOCKING, 6 ADVISORY, all eight accepted and none overruled. Its BLOCKING 1 removed the lead's discretion over where the two generated arms' words begin and end — the design as frozen would have had the lead cutting the spans out of whole-story translations by eye, knowing the hypothesis. The best $0.05 in the session, for the sixth session running |
| 2026-08-08 | S133 | stage 1 — the source census, 3 seats × 26 Russian utterances (P1, P3, P5) | 0.13 from max_tokens 6,000 |
0.045447 | per-response usage.cost |
3 of 3 stop after one re-dispatch. Markedness κ 0.7214; concordance κ 0.0584 |
| 2026-08-08 | S133 | stage 2 — generation, both stories × IND-PLAIN / IND-FORCED (mistralai/mistral-medium-3-5, non-panel, amendment A5) |
0.50 from max_tokens 8,000 |
0.042427 | per-response usage.cost |
4 of 4 stop, paragraph numbering round-tripped 17/17 and 30/30 |
| 2026-08-08 | S133 | stage 3 — the English rating, 3 seats × 4 blocks, 116 items (P2, qwen3.7-max, glm-5.2) |
0.60 from max_tokens 8,000 |
0.291977 | per-response usage.cost |
12 of 12 stop, no re-dispatch. 346 of 348 cells (99.43%) |
| 2026-08-08 | S133 | stage 4 — recognition, non-gating (mistralai/mistral-medium-3-5) |
0.01 | 0.002060 | per-response usage.cost |
stop. Named Chekhov, "very confident", and got the title wrong |
| 2026-08-08 | S133 | WASTE ROW — five discarded bodies | — | 0.420768053 | per-response usage.cost; additional to the stage figures above |
49.4% of the session. Note (bhf), sixth consecutive session: deepseek-v4-pro returned zero content at length on the census, cured by (bhq)'s raised cap at the first attempt for $0.0089. Then moonshotai/kimi-k3 returned zero content on three of four generation bodies at $0.365763 — the whole allowance on hidden reasoning — while the fourth, same model, same prompt shape, same batch, came back clean, which falsifies note (bkh)'s prompt-shaped, not seat-shaped. The seat was changed per (bhf) rule (iii) and the replacement produced all four arms for $0.042, an eighth of the cost of the failures. The fourth kimi body was clean and was discarded deliberately, so that both generated arms come from one hand — mixing hands would have confounded the arm with the author. All five preserved in runs/discarded/, never overwritten (note (bhd)). New note (bkh-corr) |
| 2026-08-08 | S133 | Two complete Chekhov stories rendered twice (R06 drafts frozen at 9dcf063 before the R04 revision, per R04 §2a), both translator's logs with a registered discordance prediction, the 26-site inventory, the contamination cells and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination measured after the freeze and before the design was written, and it changed the design: 27 contiguous tokens with Garnett against 16 for two independent published hands, so the lead was demoted out of every primary |
S133 total: $0.851843373 against a declared $1.60 — 53% of the declaration. The ceiling was set in the frozen design before Stage 0 and was not raised; headroom at that moment was $4.266891680.
S134 — 2026-08-08 (UTC), E-20260808c-sense-tradeoff-de (ARM-sense-tradeoff step 2)
Pre-flight: $1.20 worst case, built from max_tokens and not from expected output (note (abc)):
critic 16,000 · 9 scoring bodies at 10,000 · 3 blind bodies at 12,000 · 2 translate bodies at 4,000 ·
1 audit body at 12,000. Headroom at dispatch: $3.415048307.
| date | session | what | pre-flight worst case (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-08 | S134 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.06 from max_tokens 16,000 |
0.028122600 | per-response usage.cost |
stop, 11,980 characters. NEEDS-REDESIGN, 5 BLOCKING, 5 ADVISORY, 8 accepted and 2 overruled with written reasons. Its BLOCKING 1 replaced P > 0.05 with a real equivalence margin — and the run's primary then failed by 0.027 on a point estimate of exactly zero, which it would have passed under the original wording. The best $0.03 in the session, for the seventh session running |
| 2026-08-08 | S134 | stage 1 — the two independent-hand arms (mistralai/mistral-medium-3-5, non-panel; R07 dispatched first, crossed against S129) |
0.05 from max_tokens 4,000 |
0.019257000 | per-response usage.cost |
2 of 2 stop, 7 of 7 segments each |
| 2026-08-08 | S134 | stage 2 — the blind naturalness pass, dispatched FIRST (amendment A10), 3 seats × 49 passages |
0.30 from max_tokens 12,000 |
0.057799540 | per-response usage.cost |
3 of 3 stop, 147 of 147 cells. Ordered before the source-visible pass so the figure every naturalness claim uses was taken before any seat saw the German |
| 2026-08-08 | S134 | stage 3 — the source-visible pass, 3 seats × 3 blocks, 49 items × 4 senses | 0.75 from max_tokens 10,000 |
0.216869800 | per-response usage.cost |
9 of 9 stop, no re-dispatch. 588 of 588 cells (100%) |
| 2026-08-08 | S134 | stage 4 — rule-compliance audit (amendment A4, non-panel) |
0.05 from max_tokens 12,000 |
0.025724400 | per-response usage.cost |
stop. 42 of 42 claimed sites judged: 30 REQUIRED, 9 LICENSED, 3 NOT-SUPPORTED |
| 2026-08-08 | S134 | WASTE ROW — none | — | 0.000000 | — | ZERO discarded bodies, zero re-dispatches, 16 of 16 stop. Note (bhf) did not fire for the first time in seven sessions. The one in-run defect cost nothing: a KeyError in the audit prompt's brace escaping, caught before any request left the machine |
| 2026-08-08 | S134 | Three lead renderings of «Michael Kohlhaas»'s Lisbeth span (R06 frozen at commit 41e017f before R08 and R07 were begun), three translator's logs, both control arms, and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination measured after the freeze and before the design was written, and it changed the design: R06 shares 110 shared 7-grams and a 21-token run with two published hands that share 21 and 16 with each other, so R06 was demoted out of every primary |
S134 total: $0.347773240 against a declared $1.20 — 29% of the declaration, the lowest fraction of a declared ceiling this project has recorded. Key-usage snapshots 78.278907863 → 78.626681103, delta 0.347773240, reconciliation exact to 0.000000000.
2026-08-08 day total: $1.932724933 of $5.00 — three sessions (S132 $0.733108320, S133 $0.851843373, S134 $0.347773240). $3.067275067 headroom, 61% of the cap.
S135 — 2026-08-08 (UTC), E-20260808d-carriage-decoupled (ARM-sense-overlap step 2)
Pre-flight: $1.00 worst case, built from max_tokens and not from expected output (note (abc)),
then revised to $1.40 when the pre-run critic's BLOCKING 3 bought a fifth stage. Headroom at
dispatch: $3.067275067.
| date | session | what | pre-flight worst case (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-08 | S135 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.06 from max_tokens 16,000 |
0.034903800 | per-response usage.cost |
stop, 11,484 characters. NEEDS-REDESIGN, 5 BLOCKING, 8 ADVISORY; 8 amendments accepted before dispatch, 4 overruled with written reasons. Its BLOCKING 3 bought stage 5 outright — the halo test that largely restores accuracy's oldest sentence — and its BLOCKING 1 bought the enrichment audit the primary's sensitivity check rests on. The best $0.03 in the session, for the eighth session running |
| 2026-08-08 | S135 | stage 1 — blind naturalness, dispatched FIRST, 3 seats × 28 passages, English only |
0.20 from max_tokens 8,000 |
0.061297100 | per-response usage.cost |
3 of 3 stop, 84 of 84 cells. Ordered before any seat saw Danish, so gate G1 was taken blind |
| 2026-08-08 | S135 | stage 2 — Danish competence screen, 3 seats × 5 items | 0.05 from max_tokens 2,000 |
0.007103300 | per-response usage.cost |
3 of 3 stop. G6 5 of 5 at every seat |
| 2026-08-08 | S135 | stage 3 — accuracy ALONE, 3 seats × 28 items, Danish present |
0.25 from max_tokens 8,000 |
0.061120149 | per-response usage.cost |
3 of 3 stop, 84 of 84 cells. RS-20260807f §6.4's explicit requirement |
| 2026-08-08 | S135 | stage 4 — style-correspondence + perceived-source-carriage, 3 seats × 28 items |
0.35 from max_tokens 10,000 |
0.065003244 | per-response usage.cost |
3 of 3 stop, 168 of 168 cells |
| 2026-08-08 | S135 | stage 5 — all four senses in ONE call (amendment A5, the halo test), 3 seats × 28 items |
0.40 from max_tokens 12,000 |
0.097999820 | per-response usage.cost |
3 of 3 stop, 336 of 336 cells. The stage that answers §6.4's second question |
| 2026-08-08 | S135 | C1 — independent propositional-parity and enrichment call (mistralai/mistral-medium-3-5, non-panel) |
0.05 from max_tokens 8,000 |
0.012784500 | per-response usage.cost |
stop. 7 of 7 pairs equivalent, 5 of 5 planted errors named, 2 of 2 error-free pairs cleared, 4 of 7 enrichments flagged |
| 2026-08-08 | S135 | C2 — independent formal-device census from the Danish alone (mistralai/mistral-medium-3-5, non-panel) |
0.04 from max_tokens 6,000 |
0.008688000 | per-response usage.cost |
stop. 14 devices named, covering 12 of the lead's 16 Class A devices (0.75), one clean false positive |
| 2026-08-08 | S135 | WASTE ROW — none | — | 0.000000 | — | ZERO discarded bodies, zero re-dispatches, 18 of 18 stop. Note (bhf) silent for the second session running. The one in-run defect cost nothing: an import-order bug in run.py's parity stage raised ModuleNotFoundError before any request left the machine; fixed and re-invoked |
| 2026-08-08 | S135 | Two lead artifacts — «Ved Vejen» rendered close under R06 with its log frozen and committed at 5999d6f before this design existed, then flattened under R14 by a 24-site operator table — plus the 23-device census, the 19 oddity edits, the 5 error edits and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination could not be measured: no English «Ved Vejen» is reachable outside copyright, so the declaration is a reachability statement and says so on the artifact |
S135 total: $0.348899913 against a declared $1.40 — 25% of the declaration. Key-usage snapshots 78.632095803 → 78.968211216, delta 0.336115413 against a per-request sum of 0.348899913, a gap of $0.012784500 — exactly the C1 parity call's billed cost, to the cent, i.e. the last dispatch had not settled at close. That is the settling lag note (bcx) names, with the cleanest attribution the ledger has recorded: not an estimate of the gap's cause but an exact match to one named body. Per-request sums are ledgered, per this page's stated method.
2026-08-08 day total: $2.281624846 of $5.00 — four sessions (S132 $0.733108320, S133 $0.851843373, S134 $0.347773240, S135 $0.348899913). $2.718375154 headroom, 54% of the cap.
UPDATED 2026-08-08 by S137: the UTC day stands at $3.580596 of $5.00 — seven sessions
(S132 $0.733108320, S133 $0.851843373, S134 $0.347773240, S135 $0.348899913,
S136 $0.857053410, S137 $0.441917740). $1.419404 headroom, 28% of the cap.
S137's key-usage snapshots 79.851203626 → 80.293121366, delta 0.441917740, against a
per-request sum over every body including both dead ones of 0.441917740 — a gap of exactly
zero, which is note (bhd) obeyed after being violated at S136. Waste $0.160546000, 36.3%, both
items finish_reason: length with zero content; one of the two was caught by a registered live
probe on the smallest call of its stage (note (bks)).
| date | session | what | pre-flight worst case (USD) | actual (USD) | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-08 | S136 | stage 0 — pre-run critic (nvidia/nemotron-3-ultra-550b-a55b, non-panel) |
0.06 from max_tokens 16,000 |
0.029126400 | per-response usage.cost |
stop, 14,248 characters. NEEDS-REDESIGN, 8 BLOCKING, 5 ADVISORY; 11 of 13 findings accepted, 2 overruled with reasons. Its BLOCKING 5 caught that the run's only moderator turned on a free-indirect-discourse judgement call by the one party who knew the hypothesis, and its BLOCKING 2 turned a one-model floor into a two-hand floor — which the data then vindicated outright, LOW-B coming back at exactly 0.0000. The best $0.03 in the session for the ninth session running |
| 2026-08-08 | S136 | stage 1 — independent site census B1 (mistralai/mistral-medium-3-5, non-panel) |
0.07 from max_tokens 10,000 |
0.051423000 | per-response usage.cost |
stop. 96 marked stretches; zero Japanese, recorded as a limit |
| 2026-08-08 | S136 | stage 1 — independent site census B2 (x-ai/grok-4.5, panel P3, not a seat in this run) |
0.07 from max_tokens 10,000 |
0.066602400 | per-response usage.cost |
stop. 89 stretches across all four languages |
| 2026-08-08 | S136 | stage 2 — LOW-A, low-instructed rendering of 52 sites (mistralai/mistral-medium-3-5) |
0.05 from max_tokens 8,000 |
0.013702500 | per-response usage.cost |
stop, 52 of 52 sites |
| 2026-08-08 | S136 | stage 2 — LOW-B (x-ai/grok-4.5) |
0.05 from max_tokens 8,000 |
0.024234400 | per-response usage.cost |
stop, 52 of 52. Dispatched under the wrong tag by an operator error and re-filed with the error recorded on the body; the prompt is byte-identical to the LOW prompt |
| 2026-08-08 | S136 | stage 3 — seat competence screen, 3 seats × 8 items, four languages, no English shown | 0.05 from max_tokens 2,000 |
0.014287832 | per-response usage.cost |
3 of 3 stop, 24 of 24 |
| 2026-08-08 | S136 | stage 4 — REG coding, 5 groups × 3 seats, 158 items |
0.72 from max_tokens 9,000 |
0.248639478 | per-response usage.cost |
15 of 15 stop, 474 of 474 cells, no missing data |
| 2026-08-08 | S136 | stage 5 — parity + planted errors + addition audit (nvidia/nemotron-3-ultra-550b-a55b; wrote neither low arm, is not a seat) |
0.10 from max_tokens 24,000 |
0.056550600 | per-response usage.cost |
stop, 45 of 47 items. 4 of the 4 planted errors it returned, named by content. Its equivalence figure failed the registered bar and withheld every primary |
| 2026-08-08 | S136 | WASTE ROW — $0.352486800, 41.1% of the session, the worst this ledger has carried | — | 0.352486800 | — | Two moonshotai/kimi-k3 bodies returned finish_reason: length with ZERO content after spending the whole cap on hidden reasoning — $0.179433000 (census B2) and $0.125778000 (LOW-B) — and one parity call truncated the same way at $0.037906200 before succeeding at a 24,000 cap. Plus $0.009369600 for a body the lead overwrote by hand instead of rotating into runs/discarded/, a note (bhd) violation, which is exactly the key-delta gap. Note (bhf) fires after two silent sessions; new note (bkp) |
| 2026-08-08 | S136 | Three lead artifacts — «I Malavoglia» ch. I opening rendered under R06 and again under R21, both logs frozen and committed at 422ecff before Craig 1890 was opened; plus List A's 52 sites, the alignment script, all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). The contamination gate removed the R06 arm from the experiment before the design existed — 44 / 6 / run 16 against Craig |
S136 total: $0.857053410 against a declared $1.60 — 54% of the declaration. Key-usage snapshots 78.988319016 → 79.845372426, delta 0.857053410, against a per-request sum over surviving bodies of $0.847683810; the gap is exactly $0.009369600, the one body the lead overwrote by hand. Per-request sums plus the named lost body are ledgered, per this page's stated method.
2026-08-08 day total: $3.138678256 of $5.00 — five sessions (S132 $0.733108320, S133 $0.851843373, S134 $0.347773240, S135 $0.348899913, S136 $0.857053410). $1.861321744 headroom, 37% of the cap.
S138 — 2026-08-08 (UTC), E-20260808g-two-pasts (ARM-legend step 1)
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-08 | S138 | stage 0 — pre-run critic, DEAD (nvidia/nemotron-3-ultra-550b-a55b, the registered critic) |
0.11 from max_tokens 16,000 |
0.061347600 | per-response usage.cost |
WASTE. finish_reason: length, zero content, whole cap on hidden reasoning. Note (bhf): zero content means change the seat, not raise the cap. Body in runs/discarded/ |
| 2026-08-08 | S138 | stage 0 — pre-run critic (google/gemini-3.6-flash, panel P2, not a hand in this run) |
0.11 from max_tokens 12,000 |
0.035775000 | per-response usage.cost |
stop, 3,893 characters. NEEDS-REDESIGN, 3 BLOCKING and 1 ADVISORY, all four accepted, none overruled. Its BLOCKING 1 caught that the same-arm pairs were fixed model pairs, so a model-family habit would have been read as an arm effect; its BLOCKING 2 replaced the positive control — the lexical one it had would have licensed nothing about morphology, and the tense control that replaced it is the only reason the null is readable. The best $0.04 in the session, for the eleventh session running |
| 2026-08-08 | S138 | the four translation hands, 17 passages each, arms counterbalanced, tag never mentioned | 0.42 from max_tokens 8,000 |
0.056604427 | per-response usage.cost sum |
4 of 4 stop, 17 of 17 numbered items each, 68 of 68 cells, 0 unextractable. H1 deepseek-v4-pro $0.002214727 · H2 mistral-medium-3-5 $0.012240 · H3 gpt-5.6-terra $0.0143453 · H4 grok-4.5 $0.0278044. Cheapest powered stage this ledger has carried |
| 2026-08-08 | S138 | duplicate H1 dispatch (deepseek/deepseek-v4-pro) |
— | 0.001547499 | per-response usage.cost |
Not waste and not used: the first translate invocation was killed by a 2-minute shell timeout after H1 and H2 had landed, and the re-invocation re-dispatched H1. The first body is the one analysed, deterministically; the duplicate is preserved as translate-H1.1.json |
| 2026-08-08 | S138 | recognition probe, non-gating (mistralai/mistral-medium-3-5) |
0.02 | 0.001519500 | per-response usage.cost |
stop. Named Jókai, «Az ezüst kecske», "very confident", and is wrong |
| 2026-08-08 | S138 | UNRECONCILED — the H3 dispatch the shell killed in flight | — | 0.011892249 | key-usage delta, not a returned body | Note (bid), second firing. The 2-minute timeout killed the first translate run mid-H3; OpenRouter billed the request and no body was ever received. The residual is the size of a gpt-5.6-terra call of this shape ($0.0143 for the one that returned). Ledgered at the key delta, per this page's method |
| 2026-08-08 | S138 | All lead work — «Szent Péter esernyője» ch. III translated whole under R05 (2,511 Hungarian → 3,479 English words, 93 paragraphs, log D1–D20), the register, the Part I collation of two witnesses, the whole-novel tag census, the held-out replication on «A vén gazember», the Worswick alignment, and all analysis and verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination measured after the freeze and before any locus was chosen — clean against Worswick 1900, DEPENDENT against the lead's own S077 rendering at 27 contiguous tokens in 689 words, note (bhb) |
S138 total: $0.168686275 against a declared $0.76 — 22% of the declaration. Key-usage
snapshots 80.634643766 → 80.803330041, delta 0.168686275, against a per-request sum over
every body including the dead one of $0.156794026; the gap is $0.011892249 and is the
killed-in-flight H3 dispatch, note (bid). Waste $0.061347600, 36.4% — one body,
finish_reason: length with zero content, and note (bhf)'s own remedy (change the seat) worked
first time on a model that then returned the most consequential findings of the session.
UPDATED 2026-08-08 by S138: the UTC day stands at $3.749282 of $5.00 — eight sessions (S132 $0.733108320, S133 $0.851843373, S134 $0.347773240, S135 $0.348899913, S136 $0.857053410, S137 $0.441917740, S138 $0.168686275). $1.250718 headroom, 25% of the cap.
S139 — 2026-08-08 (UTC), E-20260808h-trajectory (ARM-trajectory step 1)
| date | session | stage | pre-flight | actual USD | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-08 | S139 | stage 0 — pre-run critic, DEAD (nvidia/nemotron-3-ultra-550b-a55b, the registered critic) |
0.09 from max_tokens 12,000 |
0.046932 | per-response usage.cost |
WASTE. finish_reason: length, zero content, whole cap on hidden reasoning. Note (bhf), second session running for this seat |
| 2026-08-08 | S139 | stage 0 — pre-run critic, DEAD (moonshotai/kimi-k3, the design's declared fallback) |
0.15 from max_tokens 10,000 |
0.168210 | per-response usage.cost |
WASTE. Same failure, same stage, more expensive. Two registered critic seats dead in one session |
| 2026-08-08 | S139 | stage 0 — pre-run critic (mistralai/mistral-medium-3-5, non-panel, not a hand and not a rater) |
0.06 from max_tokens 8,000 |
0.031467 | per-response usage.cost |
stop, 11,198 characters. NEEDS-REDESIGN, 5 BLOCKING and 5 ADVISORY; 9 of 10 accepted, 1 overruled in writing. Its BLOCKING 1 caught that the frozen selection rule returned a POS block of 3 ты and 1 вы, confounding the control with register; its BLOCKING 4 caught that the raters were never told where the manipulated half begins, which forced the whole pipeline onto numbered paragraphs and PART ONE / PART TWO. The best $0.03 in the session for the twelfth session running |
| 2026-08-08 | S139 | stage 1 — independent scene screen, 35 candidate spans (mistralai/mistral-medium-3-5) |
0.03 from max_tokens 4,000 |
0.031142 | per-response usage.cost |
stop, 35 of 35. 29 marked as two-party exchanges |
| 2026-08-08 | S139 | stage 4 — the three hands, 16 spans each, 151 numbered paragraphs, arms counterbalanced | 0.19 from max_tokens 14,000 |
0.106989 | per-response usage.cost sum |
H1 deepseek-v4-pro $0.002415 (after a dead $0.004009) · H2 gpt-5.6-terra $0.049698 · H3 grok-4.5 $0.054876 (after a dead $0.054324 and a 400). 16 of 16 sites complete, 151 of 151 paragraphs, 0 junk lines, from all three |
| 2026-08-08 | S139 | stage 5 — the reading probe, 3 seats × 2 calls × 16 items | 0.14 from max_tokens 4,000–12,000 |
0.134635 | per-response usage.cost sum |
R1 gemini-3.6-flash $0.111878 (its two counted calls) · R2 qwen3.7-max $0.021806 · R3 glm-5.2 $0.000952. 94 English ratings |
| 2026-08-08 | S139 | stage 6 — source-side gate, 2 seats, complementary arms, 16 Russian spans each | 0.05 from max_tokens 4,000 |
0.011772 | per-response usage.cost sum |
deepseek-v4-pro $0.000985 · gpt-5.6-terra $0.010787. 32 of 32. 0 grammaticality flags — and this is the stage that withheld the primary at +0.667 against +1.00 |
| 2026-08-08 | S139 | WASTE ROW — $0.273475, 42.5% of the session | — | 0.273475 | — | Four bodies with zero content after spending their caps on hidden reasoning: two critic seats ($0.046932, $0.168210) and two hands ($0.004009, $0.054324). Note (bhf), fourth and fifth firings. Both hands were recovered on the same seat by disabling reasoning in the request — new note (bkw). The superseded rater call is ledgered separately below |
| 2026-08-08 | S139 | F6 re-dispatch — the superseded rater call (google/gemini-3.6-flash, R1c0) |
— | 0.054597 | per-response usage.cost |
Not waste and not used: it returned 13 of 16 parseable items at finish_reason: length, which is over F6's bar, so the registered criterion required one re-send. R1c0b returned 16 of 16 and is the body analysed; R1c0 is preserved |
| 2026-08-08 | S139 | All lead work — «Дуэль» XV ¶660–700 translated whole under R06 (1,107 → 1,268 words, log D1–D15), the site pool, the whole address manipulation and its override table, all analysis, the census and the verification |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination measured after the freeze: suspected, 23 twelve-grams and run 18 against Garnett 1916, of which 19 of 23 carry a proper name and no fifteen-gram survives name exclusion — and uncalibrated, because no second published English of «Дуэль» is freely reachable |
| 2026-08-09 | S141 | score — 3 seats × 6 blocks, 106 items, four senses, German present | 0.99 from max_tokens 10,000 | 0.468080 | per-response usage.cost sum | 1,272 of 1,272 source-visible cells. J1 $0.109, J2 $0.294, J3 $0.045 |
| 2026-08-09 | S141 | blind — 3 seats × 2 blocks, naturalness on the English alone, dispatched FIRST | 0.40 from max_tokens 12,000 | 0.168185 | per-response usage.cost sum | 317 of 318 cells; one row dropped from J2's first block, carried as limit 6 |
| 2026-08-09 | S141 | M1 retrieval probe — 3 seats + the independent hand, 15 German segments interleaved | 0.03 from max_tokens 3,000 | 0.047682 | per-response usage.cost sum | 27 of 28 Kleist hits, 0 of 32 Arnim. Includes the $0.029877 re-dispatch of the truncated gemini body at a raised cap |
| 2026-08-09 | S141 | independent hands — 4 unbriefed arms (IND-R06 on both sources, IND-R07/IND-R08 on Arnim) | 0.03 from max_tokens 4,000 | 0.036926 | per-response usage.cost sum | 4 of 4 stop, every segment returned. The IND-R06 arm cost $0.0092 and is what note (bkz) is about |
| 2026-08-09 | S141 | pre-run critic — one adversarial pass over the frozen design | 0.06 from max_tokens 16,000 | 0.016118 | per-response usage.cost | NEEDS-REDESIGN, 6 BLOCKING. Dispatched with reasoning: {enabled: false} per note (bkw) from the start |
| 2026-08-09 | S141 | compliance audit — 61 logged rule claims judged by a non-panel reader | included above | 0.010618 | per-response usage.cost | 54 REQUIRED / 6 LICENSED / 1 NOT-SUPPORTED |
| 2026-08-09 | S141 | WASTE ROW — $0.026397, 3.4% of the session | — | 0.026397 | — | One body: gemini-3.6-flash on the probe, length at 264 chars. Note (bkw) was applied prophylactically and gemini 400s on reasoning.enabled, which is (bkw)'s own recorded caveat — so the switch could not be used and the cap was burned on hidden reasoning |
| 2026-08-09 | S141 | All lead work — Arnim «Die Kronenwächter» I.7 rendered three times under R06 → R08 → R07 (694 German words → 807 / 775 / 795 English, 32 + 29 logged decisions), the copy-text collation, the segmentation, the control operators, all analysis and the verifier | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). contamination: none with no comparator to measure against, declared on each artifact rather than left silent |
How the rows add up, stated because they do not add up naively. The stage rows carry live
bodies only: 0.031467 (critic) + 0.031142 (screen) + 0.106989 (hands) + 0.134635 (raters) +
0.011772 (gate) + 0.054597 (the F6 re-send) = $0.370602. The four dead bodies sit in the
waste row alone and are not double-charged anywhere: 0.046932 + 0.168210 (critics, which also
appear as their own rows above) + 0.004009 + 0.054324 (hands) = $0.273475. $0.370602 +
$0.273475 = $0.644077, and the two dead-critic rows above are the same money as their share of
the waste row, shown twice for readability and counted once.
S139 total: $0.644078 against a declared worst case of $0.67. Key reconciliation exact to 1e-9: snapshots 80.804849541 → 81.448927403, delta 0.644077862, per-response sum 0.644077863.
UPDATED 2026-08-08 by S139: the UTC day stands at $4.393360 of $5.00 — nine sessions (S132 $0.733108320, S133 $0.851843373, S134 $0.347773240, S135 $0.348899913, S136 $0.857053410, S137 $0.441917740, S138 $0.168686275, S139 $0.644077862). $0.606640 headroom, 12% of the cap.
S140 — 2026-08-09 (UTC), E-20260809a-trajectory-ja (ARM-trajectory step 2, arm closes)
New UTC day: the ledger opens at $0.00 of $5.00.
| date | session | stage | pre-flight | actual USD | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-09 | S140 | stage 0 — two pre-run critics, two labs, both ALIVE (mistral-medium-3-5; nemotron-3-ultra with reasoning:{enabled:false}) |
0.12 from max_tokens 10,000 |
0.044961 | per-response usage.cost |
mistral $0.033836, 9,691 chars, NEEDS-REDESIGN, 10 findings; nemotron $0.011126, 10,390 chars, NEEDS-REDESIGN, 10 findings. Note (bkw) applied BEFORE the failure — the same seat burned $0.046932 on a zero-content body as S139's registered critic. Twelve amendments came out of these two calls |
| 2026-08-09 | S140 | stage R — the Russian re-gate, 6 counted calls, S139's 16 spans byte-identical | 0.10 from max_tokens 4,000 |
0.095699 | per-response usage.cost sum |
3 seats × 2 complementary arm-sets = 6 ratings per site against S139's 2. Three calls died at the cap that worked at S139 on the identical prompt; cap raised to 12,000 per (bhf), two recovered; the third recovered only on (bkw)'s reasoning switch at $0.0010549. P0R = +1.389 against +1.00 — S139's withheld primary released |
| 2026-08-09 | S140 | stage 1 — three-seat scene screen, 27 candidates (mistral-medium-3-5, gpt-5.6-terra, nemotron-3-ultra) |
0.04 from max_tokens 6,000 |
0.041918 | per-response usage.cost sum |
Majority rule; 4 candidates excluded as MORE-THAN-TWO. nemotron's screen is degenerate (TWO on all 27) and contributes no discrimination |
| 2026-08-09 | S140 | stage 4 — the three hands, 12 spans each, 100 numbered paragraphs, arms counterbalanced 6/6 | 0.30 from max_tokens 16,000 |
0.101073 | per-response usage.cost sum |
H1 deepseek-v4-pro $0.002825 · H2 gpt-5.6-terra $0.026963 · H3 grok-4.5 $0.071284. 12 of 12 sites, 100 of 100 paragraphs, 0 junk lines, from all three, first attempt |
| 2026-08-09 | S140 | stage 5 — the reading probe, 3 seats × 2 calls × 12 items | 0.35 from max_tokens 8,000 |
0.142731 | per-response usage.cost sum |
R1 gemini-3.6-flash $0.081350 · R2 qwen3.7-max $0.050126 · R3 glm-5.2 $0.011256. 72 English ratings. One dead glm-5.2 body recovered by (bkw) for $0.000216 |
| 2026-08-09 | S140 | stage 6 — the Japanese source gate at parity, 6 seats × 2 calls × 12 items | 0.30 from max_tokens 12,000 |
0.165924 | per-response usage.cost sum |
Hand block $0.107 · independent block $0.058. 144 source ratings, 12 per site, 6 per arm — the note (bkv) repair, and it still returned +0.139 against +1.00 |
| 2026-08-09 | S140 | WASTE ROW — $0.049308, 7.2% of the session | — | 0.049308 | — | Five zero-content bodies: glm-5.2 as a mis-seated screen ($0.039174 — a dispatch error, it is an English rater seat), deepseek ×3 on stage R ($0.003526, $0.001914, $0.001772), glm-5.2 rater call ($0.002922). Against S139's 42.5%, and the difference is note (bkw) applied prophylactically |
| 2026-08-09 | S140 | All lead work — «こころ» 上・十三 translated whole under R06 (1,328 chars → 703 words, log D1–D14), the pool, the whole predicate-register manipulation and its 39-rule override table, the selection, the census, all analysis and the verifier |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination suspected and unmeasured — no published English of «こころ» is freely reachable, and the note (bhb) repository check returned zero prior lead renderings before the section was selected |
| 2026-08-09 | S140 | SUPERSEDED-BUT-NOT-EMPTY ROW — $0.045223 | — | 0.045223 | per-response usage.cost | Not waste and not used. screen $0.012110: the first single-seat screen over the 20-candidate pool, superseded when amendment A1 re-pooled to 27 candidates and A2 put the screen on three seats — its body is preserved at run/screen-v1.json beside run/pool-v1.json. RG10 $0.033113: a length body truncated at 129 characters, superseded by RG10b at the raised cap |
How the rows add up, stated because they do not add up naively. The stage rows carry counted bodies only: 0.044961 (critics) + 0.095699 (stage R) + 0.041918 (screen) + 0.101073 (hands) + 0.142731 (raters) + 0.165924 (Japanese gate) = $0.592306. The five dead bodies sit in the waste row alone ($0.049308) and the two superseded non-empty bodies in their own row ($0.045223); nothing is double-charged. $0.592306 + $0.049308 + $0.045223 = $0.686837, which is the sum of six-place roundings; the exact total is $0.686836395.
S140 total: $0.686836 against a declared worst case of $1.60. Key reconciliation to 3e-9: snapshots 81.454738403 → 82.141574795, delta 0.686836392, per-response sum over every body including the five dead ones 0.686836395. Note (bid) did not fire: every dispatch ran in the background.
UPDATED 2026-08-09 by S140: the UTC day stands at $0.686836 of $5.00 — one session. $4.313164 headroom, 86% of the cap.
S141 total: $0.774006 against a declared worst case of $1.60. Counted bodies $0.747609336 + the one dead body $0.026397 = $0.774006336. Key reconciliation EXACT: snapshots 82.147424595 → 82.921430931, delta 0.774006336, against a per-response sum over every body including the dead one of 0.774006336 — 0.000000000 apart. Note (bid) did not fire: every dispatch ran in the background.
UPDATED 2026-08-09 by S141: the UTC day stands at $1.460842 of $5.00 — two sessions (S140 $0.686836, S141 $0.774006). $3.539158 headroom, 71% of the cap.
| 2026-08-09 | S142 | stage 0 — pre-run critic over the frozen design and materials/arms.py (nvidia/nemotron-3-ultra-550b-a55b, reasoning off per (bkw)) | 0.10 from max_tokens 16,000 | 0.018533 | per-response usage.cost | NEEDS-REDESIGN, 3 BLOCKING + 3 SERIOUS + 2 MINOR. All three BLOCKING are factually false (it mis-simulated Python substring containment and invented a carriers entry) and were overruled with demonstrations; six lesser findings accepted as amendments A1–A7. Its one wholly correct catch — that G6's 0.33 bar sat below chance — is what makes §7's withholding meaningful |
| 2026-08-09 | S142 | stage 1 — blind naturalness, 3 seats × 2 blocks, English only, dispatched before any seat saw Czech | 0.11 from max_tokens 3,000 | 0.036199 | per-response usage.cost | J1 gpt-5.6-terra $0.0124 · J3 deepseek-v4-pro $0.0051 · J2 gemini-3.6-flash recovered $0.0187. 84 of 84 cells |
| 2026-08-09 | S142 | stage 2 — blind diction formality, 3 seats × 2 blocks, a separate call so the fluency question is not primed | 0.11 from max_tokens 3,000 | 0.035002 | per-response usage.cost | 84 of 84 cells. This stage is gate G2, and note (bkv) is why it exists at the primary's precision rather than at one seat's |
| 2026-08-09 | S142 | stage 3 — Czech competence screen, 3 seats × 5 gloss items | 0.05 from max_tokens 3,000 | 0.008646 | per-response usage.cost | 5 of 5 for all three seats on pětmecítma, prý, krev a mlíko, funus, nezůstavilť — the project's first Czech and the panel reads it |
| 2026-08-09 | S142 | stage 4 — voice alone, Czech present, 3 seats × 4 blocks, two required quote fields | 0.40 from max_tokens 6,000 | 0.091393 | per-response usage.cost | 84 of 84 cells, 12 of 12 bodies stop, first attempt. The Czech quote is verbatim in 83 of 84 |
| 2026-08-09 | S142 | C1 — independent propositional parity (qwen/qwen3.7-max, non-panel), 21 pairs, A/B seed-shuffled | 0.15 from max_tokens 8,000 | 0.046933 | per-response usage.cost | 14 of 14 formal pairs EQUIVALENT, 5 of 5 planted errors named by the exact word, 0 false positives. The cleanest parity control this project has run |
| 2026-08-09 | S142 | WASTE ROW — $0.051605, 17.9% of the session | — | 0.051605 | per-response usage.cost | Two google/gemini-3.6-flash stage-1 bodies, both having spent their cap on hidden reasoning: one echoed the passage back, one truncated at 8 items of 14. Recovered on the same seat by (bkw)'s other form, reasoning.effort: low — see the note's S142 cell |
| 2026-08-09 | S142 | All lead work — Neruda «O měkkém srdci paní Rusky» ¶1–5 translated whole (630 Czech words → 794 English, R06 draft then R04, log D1–D23), the seven-segment source unit, all three derived arms and their operator tables, every mechanical check, the analysis and the verifier | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). contamination: suspected and unmeasured — no free English of the work exists; note (bhb)'s repository check returned zero before the passage was selected |
S142 total: $0.288309979 against a declared worst case of $1.20. Key reconciliation EXACT: snapshots 83.157617431 → 83.445927410, delta 0.288309979, against a per-response sum over every body including the two dead ones of 0.288309979 — 0.000000000 apart. Note (bid) did not fire: every dispatch ran in the background.
UPDATED 2026-08-09 by S142: the UTC day stands at $1.749152 of $5.00 — three sessions (S140 $0.686836, S141 $0.774006, S142 $0.288310). $3.250848 headroom, 65% of the cap.
S143 — 2026-08-09 (UTC), E-20260809d-forced-choice (ARM-forced-choice step 1, T4)
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-09 | S143 | pre-run critic over the frozen design, the built items and the exact prompt string (mistralai/mistral-medium-3-5, reasoning off per (bkw)) |
0.12 from max_tokens 16,000 |
0.040940 | per-response usage.cost |
stop, 10,775 chars. NEEDS-REDESIGN, 2 BLOCKING / 4 SERIOUS / 3 MINOR. Seven amendments accepted, five findings overruled in writing, all before dispatch. Its SERIOUS 4 is the finding that was overruled and then came true — note (blb) |
| 2026-08-09 | S143 | rating run — 6 seats × 2 blocks (scale forward, scale reversed), 28 items, whole dialogue as context | 0.52 from max_tokens 6,000 |
0.146600 | per-response usage.cost sum |
336 of 336 ratings, 12 of 12 bodies stop, no re-dispatch, no missing item. S1 gpt-5.6-terra $0.01340 · S2 gemini-3.6-flash $0.03897 · S3 grok-4.5 $0.02796 · S4 kimi-k3 $0.04948 · S5 deepseek-v4-pro $0.00073 · S6 qwen3.7-max $0.01606 |
| 2026-08-09 | S143 | WASTE ROW — $0.000000, 0.0% of the session | — | 0.000000 | — | No dead body and no re-dispatch. Note (bkw) applied prophylactically to five of six seats and effort: low to the two that reject the other form |
| 2026-08-09 | S143 | All lead work — Poe's "The Cask of Amontillado" rendered whole into French under R06 then R04 (2,306 English words → 2,522 French, 89 paragraphs, translator's log with a registered 26-site address grid), the seven-hand census, the anchor, the dependence_check Unicode repair and its eleven tests, all analysis and the verifier |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination measured after the freeze and declared high: 65 twelve-grams and a 23-token run against Baudelaire, against a maximum of 15 across ten independent published pairs of the same story |
S143 total: $0.187540 against a declared ceiling of $0.90. Key reconciliation EXACT to 2 × 10⁻⁹: snapshots 84.058177412 → 84.245717057, delta 0.187539645, against a per-response sum over all 13 bodies of 0.187539647. Note (bid) did not fire.
UPDATED 2026-08-09 by S143: the UTC day stands at $1.936692 of $5.00 — four sessions (S140 $0.686836, S141 $0.774006, S142 $0.288310, S143 $0.187540). $3.063308 headroom, 61% of the cap.
S144 — 2026-08-09 (UTC), E-20260809e-legend-layer (ARM-legend step 2, T1)
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-09 | S144 | pre-run critic over the frozen design (openai/gpt-5.6-terra, P1, max_tokens 12,000) |
0.40 declared ceiling, built from the cap and the ×4 routing caution | 0.0288515 | per-response usage.cost |
stop, provider OpenAI, 15,712 chars. NEEDS-REDESIGN, 3 BLOCKING / 8 SERIOUS / 3 MINOR, all 14 accepted, nothing overruled. Two findings replaced the experiment (opportunity-set null; ≥2-distinct-lemma statistic) and a third — exclude the hypothesis-generating part — is what made the primary fail. Notes (bld), (ble) |
| 2026-08-09 | S144 | WASTE ROW — $0.000000, 0.0% of the session | — | 0.000000 | — | One call, one clean body, no re-dispatch. Note (bkw) was not needed: P1 has no dead-body record |
| 2026-08-09 | S144 | All lead work — Mikszáth ch. IV ¶141–¶210 rendered whole (2,389 HU → 3,376 EN, log D21–D42), chapter IV's ⚑ readings pinned and a ninth divergence found, the whole-novel census (2,133 paragraphs, 721 candidate types adjudicated, 992 opportunity tokens), the permutation, the verifier and the contamination measurement | 0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination measured after the freeze and declared none: 1 twelve-gram, 12-token run against Worswick 1900 ch. IV |
S144 total: $0.0288515 against a declared ceiling of $0.40. Key reconciliation EXACT:
snapshots 84.309129757 → 84.337981257, delta 0.028851500, against the single response's
usage.cost of 0.0288515.
Observation, not a discrepancy: the pre-session snapshot (84.309129757) sits 0.0634127 above
S143's closing snapshot (84.245717057). No lit-trans session ran between them; the key's all-time
usage includes non-project spend (CLAUDE.md), which is why per-request cost is primary and the
delta is the sanity check.
UPDATED 2026-08-09 by S144: the UTC day stands at $1.965544 of $5.00 — five sessions (S140 $0.686836, S141 $0.774006, S142 $0.288310, S143 $0.187540, S144 $0.028852). $3.034456 headroom, 61% of the cap.
S145 — 2026-08-09 (UTC), E-20260809g-device-cross (ARM-register-devices, T5)
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-09 | S145 | pre-run critic over the frozen design and the frozen R23 rule set (openai/gpt-5.6-terra, reasoning off, max_tokens 16,000) |
part of the $1.80 declared ceiling; $0.40 line | 0.038737250 | per-response usage.cost |
stop, 23,980 chars. NEEDS-REDESIGN, 10 BLOCKING / 13 SERIOUS / 1 MINOR; 22 accepted, 2 overruled in writing. Its BLOCKING 1+2+6 split the primary into a policy contrast and a conditional one |
| 2026-08-09 | S145 | generation, 16 calls — 2 hands × 2 cells × {∅, A, B, AB}, x-ai/grok-4.5 and mistralai/mistral-medium-3-5, cap 3,000 |
0.20 | 0.081855200 | per-response sum | 16 of 16 stop, every site returned, no re-dispatch. Two HTTP 400s cost nothing: grok-4.5 and gemini-3.6-flash both reject reasoning: {enabled: false} ("Reasoning is mandatory for this endpoint") and take effort: low |
| 2026-08-09 | S145 | loc, 2 judges — nvidia/nemotron-3-ultra-550b-a55b and z-ai/glm-5.2, 158 ⟨site, arm⟩ items, cap 16,000 |
0.30 | 0.081623758 | per-response sum | 4 surviving bodies. z-ai/glm-5.2 was a live slug probe on the smallest call and returned clean JSON on both cells |
| 2026-08-09 | S145 | REG rating, 3 seats × 2 cells, 158 items, caps 16,000 / 12,000 |
0.45 | 0.084495800 | per-response sum | 6 of 6 stop; 473 of 474 expected cells returned. The one miss is ⟨JA-05, PUB⟩ from J2, and amendment A15 dropped that whole site from every primary |
| 2026-08-09 | S145 | WASTE ROW — $0.119569500, 29.4% of the session | — | 0.119569500 | — | Two bodies, both nemotron-3-ultra on the JAPANESE loc call. (a) $0.0597804, finish_reason: length, zero visible content, all 16,000 tokens spent on hidden reasoning — RS-20260808f §7's failure reproduced exactly, same model, same language; the registered split-by-site fallback re-ran the same 54 items for $0.0237612 + $0.0151782. (b) $0.0597891, a billed request whose body was never received: the first dispatch ran in a foreground shell the harness killed at its two-minute limit while the request was open. The lesson is about the client, not the model — note (blf) |
| 2026-08-09 | S145 | All lead work — R23 frozen, the Malavoglia crossed quadruple (A, B, AB over the whole 595-token span) with its log and dependence measurement, the frozen compliance protocol code.py and its 27 self-tests, all analysis and the verifier |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). Contamination measured after the freeze and declared high by design (the cells are a manipulation); clean against Craig 1890 on all three cells — 0 twelve-grams, longest run 8 |
S145 total: $0.406281508 against a declared ceiling of $1.80. Key-usage snapshots
84.337981257 → 84.744262765, delta 0.406281508; per-response sum over the 29 bodies in runs/
is $0.346492408. The gap of $0.059789100 is NOT rounding and is not unexplained: it is
waste-row item (b), the billed request whose body a shell timeout killed, and its size matches the
re-dispatch of the same call to within $0.0000087.
UPDATED 2026-08-09 by S145: the UTC day stands at $2.371826 of $5.00 — six sessions (S140 $0.686836, S141 $0.774006, S142 $0.288310, S143 $0.187540, S144 $0.028852, S145 $0.406282). $2.628174 headroom, 53% of the cap.
S146 — 2026-08-09 (UTC), E-20260809h-rule-execution (ARM-rule-execution, T3)
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-09 | S146 | M1 retrieval probe, 7 hands over the 8 new German segments, cap 4,000, reasoning off with a documented fall-through |
part of the $1.85 declared ceiling | 0.033340493 | per-response sum | 7 of 7 clean. 6 of 7 answered NONE on all 8; mistralai/mistral-medium-3-5 claimed a published English for 4 of 8 and the void condition fired — amendment A8 quarantined it |
| 2026-08-09 | S146 | pre-run critic over the frozen design, the rubric, the R08 rule block, the operators and both frozen lead logs (openai/gpt-5.6-terra, reasoning off, cap 16,000) |
0.116 line | 0.060871250 | per-response usage.cost |
stop, 37,996 chars. NEEDS-REDESIGN, 8 BLOCKING / 16 SERIOUS / 4 MINOR; 24 accepted, 4 overruled in writing before any generation call. Its BLOCKING 1 replaced the execution bar with a within-run negative control, which is what withheld the primary |
| 2026-08-09 | S146 | generation, 14 arms — 7 hands × {R06, R08}, cap 4,000, temperature 0.3, arm order counterbalanced |
0.362 line | 0.127228 | per-response sum | 9 of 14 clean first time |
| 2026-08-09 | S146 | generation re-dispatch, 5 arms at cap 8,000 with reasoning suppressed (amendment A19) | — | 0.051515 | per-response sum | 5 of 5 clean, all eight segments each. The identical prompt to the identical slug; no substitute model (amendment A11) |
| 2026-08-09 | S146 | scoring, 3 seats × 8 blocks, 144 items, 4 senses, cap 10,000 | 1.290 line | 0.575896678 | per-response sum | 24 of 24 stop; 1,716 of 1,728 cells returned, 99.31%. 12 missing across H6-R06 (8) and H2-R06 (4) |
| 2026-08-09 | S146 | WASTE ROW — $0.158510249, 15.7% of the session | — | 0.158510249 | — | Five generation bodies, finish_reason: length with ZERO visible content, 4,000/4,002 completion tokens each spent entirely on hidden reasoning: moonshotai/kimi-k3 ×2 ($0.065232 + $0.066294), qwen/qwen3.7-max ($0.0201942), z-ai/glm-5.2 ($0.0014662), minimax/minimax-m3 ($0.0053238). Notes (bfb), (bga), (bkw). The cap was the defect, not the hands — all five returned clean at 8,000 for $0.0515. Dead bodies preserved under runs/discarded/ per note (bhd) |
| 2026-08-09 | S146 | All lead work — the source span selected, fetched and collated; T-kronenwaechter-abschied-R06-v1 and -R08-v1 with their frozen logs (41 coded sites); the operators; build_items.py, analyse.py, verify.py, collate.py |
0.00 | 0.000000 | — | No API call. Lead translation is free and is never ledgered (charter §3, A4). contamination: none, with the basis stated on the artifact because no comparator exists to run dependence_check.py against |
S146 total: $1.007361311 against a ceiling declared at $1.85 and raised mid-run to $2.10 (amendment A19, written before any scoring call, to pay for the re-dispatch of five reasoning-truncated bodies). Key-usage snapshots 85.017990275 → 86.035870165, delta 1.017879890; per-response sum over all 50 bodies 1.007361311.
The gap of $0.010518579 (1.03%) is named, not absorbed. The probe stage was first launched with
nohup … & inside a tool call — untracked by the harness — and that process was still alive when
the tracked relaunch began. Both walked the same job list with only a skip-if-exists check between
them, both dispatched the same seat's probe, and the body that lost the write race is billed in the
key delta and stored nowhere. Note (blj); converse of note (blf), which fired at S145.
Observation, not a discrepancy: the pre-session snapshot (85.017990275) sits 0.273727510 above
S145's closing snapshot (84.744262765). No lit-trans session ran between them; the key's all-time
usage includes non-project spend (CLAUDE.md).
UPDATED 2026-08-09 by S147: the UTC day stands at $4.119253 of $5.00 — eight sessions (S140 $0.686836, S141 $0.774006, S142 $0.288310, S143 $0.187540, S144 $0.028852, S145 $0.406282, S146 $1.007361, S147 $0.740065642). $0.880747 headroom, 17.6% of the cap.
S147 — 2026-08-09 (UTC), E-20260809i-terminology-drift (ARM-terminology-drift, T2)
Declared ceiling $1.20 (design.md §8, rebuilt at §9.6 after the pre-run critic). Spent
$0.740065642, 62% of the ceiling.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-09 | S147 | pre-run critic over the frozen design and the full text of every arm in two blocks (openai/gpt-5.6-terra, cap 16,000) |
0.10 line | 0.052093250 | per-response | Returned NEEDS-REDESIGN, 12 findings, 9 BLOCKING. The best-value call of the session: its BLOCKING 1 replaced an interpolation null that could have fired with no effect present |
| 2026-08-09 | S147 | G1 parity gate, 4 calls to moonshotai/kimi-k3 (P4, not a jury seat), caps 6,000 and 10,000 |
0.14 line | 0.235045650 | per-response sum | 8 of 8 pairs equivalent, 4 of 4 planted errors named, reversed-order replicate agreeing on all 11 items |
| 2026-08-09 | S147 | stage C, consistency, English alone, one arm per call — 68 calls, 4 seats × 3 blocks × 6 arms, cap 4,000 |
0.34 line | 0.201083 | per-response sum | 68 of 68 stop, zero discarded, zero retried |
| 2026-08-09 | S147 | stage V, voice, Ukrainian shown, one arm per call — 60 calls, 4 seats × 3 blocks × 5 arms, cap 4,000 |
0.34 line | 0.251844 | per-response sum | 60 of 60 stop, zero discarded, zero retried |
| 2026-08-09 | S147 | All lead work — the Ukrainian copy-text fetched and cleaned from three Wikisource Сторінка: pages; T-tini-polonyna-R06-v1 and -R04-v1 with both translator's logs, frozen at b007b18 and 8145531; template.txt, arms.py, gen.txt, spell.txt, run.py, analyse.py, verify.py; the result page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.00. No dead body, no re-dispatch, no truncation, no discarded call. One arm per call makes every jury answer a single small JSON object, so the 4,000-token cap was never approached.
Key reconciliation: EXACT. Snapshots 86.236088715 → 86.976154357, delta 0.740065642, against a per-response sum over 136 bodies of 0.740065642. Gap $0.000000000 — note (blf)'s lost-body signature is absent.
Where the estimate was wrong, and in which direction. The critic's amendment turned 54 large
calls into 128 small ones and the worst case was rebuilt at $1.15; the outturn was $0.74, because a
per-call worst case built from max_tokens assumes a 4,000-token answer and no jury answer exceeded
150 characters. The parity gate is the one line that overran — $0.235 against $0.14 — because
kimi-k3 reasons at length over a bilingual equivalence task; note (abc)'s rule held everywhere
else.
UTC day 2026-08-10
Spent $2.409270920 of $5.00 — eight sessions (S148 $0.042213320991, S149 $0.0293675, S150 $0.510728020, S151 $0.285568128, S152 $1.153321442, S153 $0.182315829, S154 $0.03313625, S155 $0.172620430). $2.590729 headroom, 51.8% of the cap unused. S148's run stopped at its registered admission gate on call 6 of a planned 205; S149's study limb needed no API call at all beyond its pre-run critic; S150 came in at 21.3% of its declared ceiling, S151 at 11.9% and S152 at 40.3% of its own.
S155 — 2026-08-10 (UTC), E-20260810z-idiom-reach (ARM-idiom-reach step 1, T5)
Pre-flight, written before dispatch and restated after the critic's A11. Base: 1 critic call
at a 16,000 cap, 6 generation calls at 6,000, 8 rating calls at 16,000. Worst case built from
max_tokens and not from an expected answer length (note (abc)): $0.89, plus a maximum
permitted retry path of one full re-dispatch at $0.86 — ceiling $1.75, declared against
headroom of $2.763 before the critic and $2.736 after it. Spent $0.172620430, 9.9% of the
ceiling.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-10 | S155 | pre-run critic over the frozen design, R24, code.py and every prompt string that would be sent (openai/gpt-5.6-terra, reasoning off, cap 16,000, finish_reason: stop) |
$0.12 line | 0.02704075 | per-response | NEEDS-REDESIGN, 16 findings, 7 BLOCKING; 12 accepted, 4 accepted-in-part, 5 individual remedies overruled in writing. Its BLOCKING 1 added the contemporaneous −I arm and moved the primary by 0.167; its BLOCKING 5/6/14 replaced a self-fulfilling positive control with the one that caught the finding |
| 2026-08-10 | S155 | generation — 2 hands × {Nsrc, Asrc, Lsrc}, 30 sites per call, temperature 0 |
$0.30 line | 0.05532580 | per-response sum | 180 of 180 items returned first time; 0 respellings in 180 cells at the −S purity gate |
| 2026-08-10 | S155 | rating — 2 judges × 4 site-blocks, 14 arms of a site kept in one call | $0.56 line | 0.09025388 | per-response sum | 840 of 840 cells returned. One dead body — glm-5.2 block c returned finish_reason: error with a JSON array truncated at 424 completion tokens — billed $0 to the project and re-dispatched once under F5 |
| 2026-08-10 | S155 | All lead work — T-botchan-R24-v1 (3,482 Japanese characters → 2,067 English words) and its 64-row W5+ table; the R24 regime; the site-aligned lead spans; code.py, run.py, analyse.py, verify.py; the contamination measurements; the result page; framework/v0.2 §7.3 and v0.2.5; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.00 billed. The one dead body cost the project nothing — the provider returned
error after 424 completion tokens and charged $0, absorbing an upstream $0.0144502 that is
recorded here because it happened, not because it was paid. Third session running with an exact
reconciliation and no paid waste.
Key reconciliation: EXACT. Snapshots 89.871470094 → 90.044090524, delta 0.172620430, against a per-response sum over 15 live bodies of 0.172620430. Agreement to 1e-15.
Where the estimate was wrong. Nowhere that cost money, and the direction is the usual one: the
worst case priced 16,000-token rating answers and the judges returned 4,300–7,500 characters, so
the rating line came in at 16% of its estimate. The one thing the estimate could not have priced is
the critic's own amendments, which added two generation calls and two rating calls after the
ceiling was first written — A11 exists because the critic asked for that arithmetic to be shown,
and it was shown before dispatch rather than after.
S154 — 2026-08-10 (UTC), E-20260810x-legend-again (ARM-legend step 4, T1, arm closed)
Pre-flight, written before dispatch. One post-hoc critic call, P1 openai/gpt-5.6-terra, a
~5,600-token page at a 7,000-token cap. List worst case built from max_tokens per note (abc):
$0.006 in + $0.042 out = $0.048; ceiling $0.19 for the 4× routing caution in
config/models.md. No other call was planned and none was made — the re-rendering, the
copy-text gate, the diff, the causal coding, the merge-rule sensitivity, the Károli fetch and the
phrase check are all lead work.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-10 | S154 | post-hoc critic over the finished result page (openai/gpt-5.6-terra, cap 7,000, provider OpenAI, finish_reason: stop, 6,004 in / 4,272 out) |
$0.19 ceiling | 0.03313625 | per-response | NEEDS-REDESIGN, 14 findings, 6 BLOCKING; all 14 accepted, none overruled. It withdrew three headline claims: a "replicates at the same rate" framing, "the register is not what moved it", and a 23-of-24 convergence statistic that was selection on the outcome — replaced by a frame fixed inside the freeze, on which the figure is 12 of 22. Note (blw) |
| 2026-08-10 | S154 | All lead work — chapters I–II re-rendered whole (1,963 English words); the punctuation-inclusive re-collation of ¶1–¶47; diff_sites.py, sensitivity.py, decision_frame.py, karoli_phrases.py; the Károli Bible re-fetched (1,195 chapter files); the result page, the craft report and every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.00. One call, one usable body, no truncation, no re-dispatch.
Key reconciliation: EXACT. 89.832125844 → 89.865262094, delta 0.033136250 against the per-response cost of 0.03313625 — residual 0.000000000.
Where the estimate was wrong. Nowhere on money. The declared deviation is elsewhere and it is free: there was no pre-run critic, because there was no design page — the predictions were frozen inside the translation log instead. P-C's registered band was arithmetically wrong (a per-English-word rate applied to a Hungarian word count) and is the visible price of the step that was skipped.
S153 — 2026-08-10 (UTC), E-20260810w-recurrence-census (ARM-recurrence step 1, T4)
Pre-flight, written before dispatch: worst case $1.44 from max_tokens, declared ceiling $1.60
(note (abc): built from the cap the request permits, not from an assumed answer length —
48 seat calls at a 4,000 cap plus one critic call at 8,000, each seat priced at its config/models.md
rate and P5 priced at the S022 routing caution rather than list). Outturn $0.182315829, 11.4% of
the ceiling. The estimate was wrong in the conservative direction and for the usual reason: the
answers are lists of short quoted strings, and no body came near its cap.
| date | session | line | estimate | actual | source | note |
|---|---|---|---|---|---|---|
| 2026-08-10 | S153 | pre-run critic over the frozen design + one complete seat prompt (qwen/qwen3.7-max, reserve slug outside the jury, reasoning off, cap 8,000, finish_reason: stop) |
$0.10 line | 0.010653920 | per-response | NEEDS-REDESIGN, 5 findings, 2 BLOCKING; 4 accepted, 1 accepted as a written limitation, none overruled. Its BLOCKING 1 split every quantity by site type and its BLOCKING 2 moved P1 off the region where two published hands measure dependent |
| 2026-08-10 | S153 | census — 3 seats × 4 hands × 2 regions × {network, control}, one ⟨hand, region, task⟩ per call | $1.34 line | 0.171661909 | per-response sum | 48 of 48 bodies returned, all finish_reason: stop, zero dead bodies, zero re-dispatch. Reasoning disabled by default per note (bkw), and google/gemini-3.6-flash run at effort: low because it and x-ai/grok-4.5 both return HTTP 400 "Reasoning is mandatory for this endpoint and cannot be disabled" — measured this session, new |
| 2026-08-10 | S153 | All lead work — T-shinel-znachitelnoe-R06-v1 and -R04-v1 (611 Russian words rendered twice, both logs frozen), the two source spans and the copy-text reduction, log_counts.py, build_items.py, run.py, analyse.py, verify.py, the contamination gate, the anchor page and its five stored texts, the result page, every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Key reconciliation. 89.496841727 → 89.679243144, delta 0.182401417 against a per-response sum of 0.182315829: residual $0.000085588 over 49 bodies, 0.047%, consistent with per-response rounding and not with a billed-but-lost body (note (blt)'s signature — a whole response priced upstream — is absent).
Waste: $0.00. No dead body, no failed task, no re-dispatch. First such session since S149.
S152 — 2026-08-10 (UTC), E-20260810t-drift-source-present (ARM-terminology-drift step 2, arm closed)
| date | session | line | estimate | actual | source | note |
|---|---|---|---|---|---|---|
| 2026-08-10 | S152 | pre-run critic, DEAD BODY — moonshotai/kimi-k3, cap 16,000, finish_reason: length with zero content, the whole cap spent on hidden reasoning |
$0.10 line | 0.266061 dead | per-response | note (bkw)'s remedy, fifth occurrence. My cap choice, not the API's |
| 2026-08-10 | S152 | pre-run critic, re-dispatched with reasoning disabled | — | 0.053280 | per-response | NEEDS-AMENDMENT, 12 findings, 3 BLOCKING; 11 accepted, 1 in part, none overruled. Its BLOCKING 3 bought the FALSE arm and MIXED paid for it |
| 2026-08-10 | S152 | G1 parity, FIRST SEAT — task failed (mistralai/mistral-medium-3-5, 3 calls) |
$0.06 line | 0.030216 dead | per-response sum | Returned non-equivalent for all ten pairs including the four with planted errors, on exactly the lexical grounds the prompt excludes. config/models.md already calls this slug the panel's weakest |
| 2026-08-10 | S152 | G1 parity, second seat (moonshotai/kimi-k3, 3 calls, reasoning off) |
— | 0.047766 | per-response sum | 4 of 4 planted errors caught, 5 of 6 real pairs equivalent; all three MATCH/DRIFT pairs equivalent with empty difference lists |
| 2026-08-10 | S152 | scoring — 2 conditions × 4 seats × (4 arms × 6 blocks + SPELL × 3), one arm per call, interleaved |
$2.40 line | 0.748792042 | per-response sum | 216 of 216 cells returned and parsed, zero dead bodies, zero retries, zero finish_reason: length in either condition |
| 2026-08-10 | S152 | All lead work — T-tini-budz-R06-v1 and T-tini-budz-R04-v1 (769 Ukrainian words rendered twice, both logs frozen), the source unit and its manifest, build_template.py, arms.py, false_arm.py, run.py, analyse.py, verify.py, the result page, both goodness-senses.md entries, every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Session total $1.153321442 against a declared ceiling of $2.86. Stored per-response sum $1.146115042; key-usage snapshots 88.343520285 → 89.496841727, delta $1.153321442.
Residual $0.007206400, and it is attributed rather than waved at. Seven retries in the dispatch
log: six 429 (no body, no charge) and one IncompleteRead(319 bytes read) — a response billed
upstream whose body never reached the client and which was re-dispatched and billed again. The
figure is one source-present scoring call at the top of the observed range. New note (blt): a
reconciliation on a runner that retries on exception should predict its residual from the retry
log, not expect zero.
Waste $0.296277, 25.9% of the per-response sum — $0.266061 of it the one dead critic body, and both waste lines are my choices (a 16,000 cap with reasoning left on; a parity seat the panel page already describes as its weakest).
S151 — 2026-08-10 (UTC), E-20260810d-carriage-naming (ARM-rule-execution step 2, arm closed)
Declared worst case $2.40, built from max_tokens and multiplied 4× for provider routing (note
(abc); config/models.md's S022 caution). Actual $0.285568128 — 11.9% of the ceiling.
| date | session | what | pre-flight | actual | method | note |
|---|---|---|---|---|---|---|
| 2026-08-10 | S151 | pre-run critic over the frozen design + the frozen translation (x-ai/grok-4.5, cap 16,000, finish_reason: stop) |
$0.20 line | 0.035686400 | per-response | NEEDS-REDESIGN, 16 findings, 4 BLOCKING; 13 accepted, 2 in part, 1 overruled in writing. Three site rows were rebuilt and an equivalence reading struck before dispatch |
| 2026-08-10 | S151 | Pass A, the naming task — 3 seats × 40 items, one item per call | $1.60 line | 0.196772128 live + 0.033348000 dead | per-response sum | 120 of 120 cells returned. Eight of J2's bodies came back finish_reason: length with the cap spent on hidden reasoning — seven of them readable but truncated, which the parse rule accepted silently (note (bls)) |
| 2026-08-10 | S151 | Pass B, the markedness check — 3 seats × 10 sites, no German shown | $0.40 line | 0.013047600 live + 0.006714000 dead | per-response sum | All ten of J2's first attempts died at the 120 cap; all ten recovered at 4,000 with reasoning at effort: low |
| 2026-08-10 | S151 | All lead work — T-kronenwaechter-heilung-R06-v1 (604 German words translated whole, 684 English words) and its frozen log; the source span and its collation against the 1857 witness; build_items.py, run.py, analyse.py, verify.py, sensitivity.py, collate.py; the forty hand-built minimal-pair spans; the result page; the goodness-senses.md entry; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Totals: $0.245506128 live + $0.040062000 dead = $0.285568128. Waste 14.03%, all of it one seat's hidden-reasoning truncations at caps I set too low.
Key reconciliation: EXACT to 0. 88.047314257 → 88.332882385, delta 0.285568128000 against a
per-response sum of 0.285568128 — residual 0.000000000000. (The close snapshot was taken
before the F5 repair stage ran and is superseded by keysnapshot-final.json; both are kept.)
S150 — 2026-08-10 (UTC), E-20260810c-register-reach (ARM-register-reach, T5, arm closed)
Pre-flight ceiling $2.40, built from max_tokens and not from expected output (note (abc)): 1
critic at 16,000, 3 census at 8,000, 8 generation at 6,000, 4 loc at 16,000, 9 REG at 12,000,
plus input tokens and a re-dispatch contingency. Spent $0.510728020, 21.3% of the ceiling.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-10 | S150 | pre-run critic over the frozen design + R23 (openai/gpt-5.6-terra, reasoning off, cap 16,000) |
$0.40 line | 0.03250275 | per-response | NEEDS-REDESIGN, 20 findings, 4 BLOCKING; 17 accepted, 3 accepted-in-part, 6 individual remedies overruled in writing. Its BLOCKING 1 turned the estimand from a device effect into a permission-policy effect |
| 2026-08-10 | S150 | stage 1 census — 3 independent annotators over 3,482 Japanese characters | $0.20 line | 0.065800475 live + 0.070250740 dead | per-response sum | All three first attempts failed: two zero-content finish_reason: length, one stop with a truncated JSON array. Repaired by disabling reasoning and raising the cap to 16,000 |
| 2026-08-10 | S150 | stage 2 generation — 2 hands × 4 R23 cells × 30 sites |
$0.35 line | 0.093885425 live + 0.028517650 dead | per-response sum | qwen/qwen3.7-max returned zero content at cap 6,000; recovered on the same seat with reasoning disabled at $0.0061 against $0.0285 for the dead call — note (bkw)'s remedy, fourth slug |
| 2026-08-10 | S150 | stage 3 loc — 2 judges × 2 blocks |
$0.40 line | 0.074806 | per-response sum | 240 ⟨site, arm⟩ items, no failures |
| 2026-08-10 | S150 | stage 4 REG — 3 seats × 3 blocks, all 8 arms of a site kept in one call |
$0.75 line | 0.14496498 | per-response sum | 720 of 720 cells returned, no re-dispatch |
| 2026-08-10 | S150 | All lead work — T-botchan-R22-v1 (12 paragraphs, 3,482 characters → 1,961 English words) and its frozen 21-row W5 table; the source span; run.py, code.py, analyse.py, verify.py; the contamination measurement; the result page; framework/v0.2 §7.2; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.098768390, 19.3%. Four dead bodies, three of them zero-content hidden-reasoning
truncations on three slugs at once (glm-5.2, nemotron-3-ultra, qwen3.7-max) at caps of
8,000 and 6,000, and one truncated-array stop. Note (bkw)'s prophylactic covered gpt-5.6-terra,
deepseek-v4-pro, gemini-3.6-flash and grok-4.5 and had never been extended to these three,
because they had not previously been used in a role with a small cap. All four recovered on the
same seat; no re-seating was needed and no bar moved.
Key reconciliation: EXACT. 87.530082937 → 88.040810957, delta 0.510728020000002 against a per-response sum of 0.510728020 — agreement to 1e-15, the tightest this ledger has recorded.
Where the estimate was wrong. Nowhere that cost money — the run came in at 21.3% of its
ceiling. The three census and generation caps were too small to hold hidden reasoning, which is
note (abc)'s standing blind spot restated: pricing the worst case from max_tokens says nothing
about whether max_tokens is large enough to hold an answer.
S149 — 2026-08-10 (UTC), E-20260810-legend-lexis (ARM-legend span E, T1)
Pre-flight, written before dispatch. One pre-run critic call, P1 openai/gpt-5.6-terra, a
~4,000-token design at a 6,000-token cap. List worst case built from max_tokens per note
(abc): $0.040; ceiling raised to $0.16 for the 4× routing caution in config/models.md.
No other call was planned and none was made — the corpora, the census reuse, the masking, the
permutation, the verifier and the contamination measurement are all lead work.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-10 | S149 | pre-run critic over the frozen design (openai/gpt-5.6-terra, cap 6,000, provider OpenAI, finish_reason: stop) |
$0.16 ceiling | 0.0293675 | per-response | NEEDS-REDESIGN, 12 findings, 6 BLOCKING; 11 accepted, 1 accepted-in-part with the overrule written. Its BLOCKING 9 cancelled an erratum the translator's frozen log had already promised |
| 2026-08-10 | S149 | All lead work — span E translated whole (2,437 English words); the span-E re-collation; the Károli Bible, five Jókai works and five historical novels fetched and normalised; run.py, verify.py; the result page; every state page |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.00. One call, one usable body, no re-dispatch, no truncation.
Key reconciliation: EXACT. Snapshots 87.364332537 → 87.393700037, delta 0.029367500, against the per-response cost of 0.0293675 — agreement to 1e-9, and the opposite of S148's lagging delta (note (abf)).
Where the estimate was wrong. Nowhere. The 6,000-token cap held a 4,097-token answer with room, which is the defect S148 recorded in the other direction.
S148 — 2026-08-10 (UTC), E-20260810-source-asymmetry (ARM-forced-choice step 2, arm closed)
Pre-flight, written before dispatch. Stage A 200 calls at a 400-token cap and ~450 prompt
tokens, stage R 5 calls at ~4,000 prompt tokens, one critic pass. Worst case built from
max_tokens and not from an expected answer length (note (abc)): $1.00, raised to $1.20
before dispatch when critic amendment A2 added 30 control calls. Spent $0.042213320991, 3.5% of
the ceiling.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-10 | S148 | pre-run critic over the frozen design and the built items (mistralai/mistral-medium-3-5, cap 8,000) |
0.05 line | 0.0139305 | per-response | NEEDS-REDESIGN, 8 findings, 4 BLOCKING; 4 accepted, 4 overruled in writing. Its BLOCKING 1 is the amendment that stopped the run |
| 2026-08-10 | S148 | stage R, the recognition admission gate — 5 seats, the whole anonymised dialogue, cap 500 |
0.05 line | 0.028282820991 | per-response sum | G2 FIRED: 2 of 5 named the story, 5 of 5 named Harte. One body truncated at the 500 cap (google/gemini-3.6-flash) — my cap, recorded as a defect |
| 2026-08-10 | S148 | stage A — 5 seats × 40 items, one item per call |
1.10 line | $0.00 — NOT DISPATCHED | — | Barred by G2. 200 calls not made |
| 2026-08-10 | S148 | All lead work — the Harte source and the Polish translation fetched, cleaned and stored; T-brown-calaveras-R06-v1 whole with its frozen log; census.py, build_items.py, run.py, verify.py; the anchor and the result page; both framework rows |
— | $0.00 | charter §3, A4 | Lead translation is free and is never ledgered |
Waste: $0.00. No dead body, no re-dispatch, no discarded call. The 200 calls that were not made are not waste — they are the gate doing the only thing a gate is for.
Key reconciliation: the lagging direction. Snapshots 87.314456017 → 87.335714017, delta 0.021258, against a per-response sum over 6 bodies of 0.042213320991. The delta is smaller than the sum, which is the accounting-lag signature this ledger has recorded before (note (abf)); per-request costs are primary and are what is ledgered here, and a lagging delta is a loose lower bound rather than a contradiction.
Where the estimate was wrong. Nowhere that cost money — but the $0.05 stage-R line was built
from a 500-token cap that turned out to be too small for one seat's hidden reasoning, truncating
its body. The estimate was right and the cap was wrong; note (abc) prices the worst case from
max_tokens and says nothing about whether max_tokens is large enough to hold an answer.
2026-08-12 — S164 (E-20260812b-emphasis-carriage)
Pre-flight: $0.22. One pre-run critic call at a 12,000 cap with a declared single re-dispatch at
20,000 if finish_reason == "length" (32,000 output tokens worst case at $6.00/M = $0.192, plus two
prompts of ~10k at $1.00/M), and two blind type-coding calls at a 4,000 cap. Worst case built from
max_tokens, note (abc). Today's other session (S163) spent $0.0235655; the day's total is
$0.09337865 of $5.00.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-12 | S164 | pre-run critic over the frozen design, the full type coding and the sealed prediction (openai/gpt-5.6-terra, cap 12,000) |
0.20 line | 0.03457825 | per-response | NEEDS-REDESIGN, 12 findings, 8 BLOCKING, all 12 accepted. finish_reason: stop — the S163 defect of a cap too small to hold the answer did not recur, and the declared re-dispatch was not needed |
| 2026-08-12 | S164 | blind type coder 1 (google/gemini-3.6-flash, cap 4,000, reasoning: low) |
0.02 line | 0.0129165 | per-response | 68 of 68 parsed |
| 2026-08-12 | S164 | blind type coder 2 (x-ai/grok-4.5, cap 4,000, reasoning: low) |
0.02 line | 0.0223184 | per-response | 68 of 68 parsed; three-way agreement with the lead 65 of 68 |
| 2026-08-12 | S164 | All lead work — the three-witness corpus built and stored; T-black-cat-R04-v3 and its log; the 148 correspondence calls; census, verifier, anchor, result, framework §7.8 |
— | $0.00 | charter §3, A4 | Lead translation and lead reading are free and are never ledgered |
Spent $0.06981315 of a declared $0.22 — 31.7%. Waste $0.00: no dead body, no re-dispatch, no discarded call.
Key reconciliation: EXACT. 96.71398668 → 96.78379983, delta 0.06981315 against a per-response sum of 0.06981315, agreeing to 1e-8. Between-session drift into this session's opening snapshot was +0.063303022 (S163 closed at 96.650683658), an order of magnitude smaller than the +1.07 S163 recorded and back in the range where the delta is usable as a cross-check.
2026-08-12 — S165 (E-20260812c-grade-shift)
Pre-flight: $0.50. Built from max_tokens, note (abc). One pre-run critic call at an 8,000 cap
(openai/gpt-5.6-terra, $6.00/M out → $0.048, prompt ~7k at $1.00/M → $0.007). Twelve translation
calls at a 1,200 cap (moonshotai/kimi-k3 6 × 1,200 × $15.00/M = $0.108 plus prompts;
deepseek/deepseek-v4-pro 6 × 1,200 × $0.87/M = $0.006 plus prompts). Seventy-two arbiter calls at a
400 cap across openai/gpt-5.6-terra, google/gemini-3.6-flash and x-ai/grok-4.5 — 24 each,
worst case 24 × 400 × $7.50/M = $0.072 at the dearest seat, $0.20 across the three, plus prompts.
Today's earlier sessions spent $0.09337865 (S163 $0.0235655, S164 $0.06981315); the day's headroom
is $4.90662135. Opening key snapshot 96.792474230.
The pre-flight above is the one written before dispatch and is left standing as the record. The
arbiter calls it prices were design v1's, and design v1 was killed by its critic before any of them
were made; design v2 re-itemised the remainder at ≈$0.19 (design.md §8) and its own critic then
cancelled that too. The ceiling never moved; what was bought was two refusals.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-12 | S165 | pre-run critic, design v1 (openai/gpt-5.6-terra, cap 8,000) |
0.055 line | 0.04057375 | per-response | NEEDS-REDESIGN, 35 findings, 27 BLOCKING, all decisive ones accepted. finish_reason: stop |
| 2026-08-12 | S165 | pre-run critic, design v2 (openai/gpt-5.6-terra, cap 8,000) |
0.055 line | 0.03854425 | per-response | NEEDS-REDESIGN, 25 findings. Its finding 16 cancelled the data collection: the primary was already legible in the frozen artifact. finish_reason: stop |
| 2026-08-12 | S165 | 12 translation calls + 2 blind coder calls | 0.13 line | CANCELLED, $0.00 | — | Not dispatched. Note (bml) |
| 2026-08-12 | S165 | All lead work — the two-stage copy-text gate over 22 page images, span B of T-dakghar-R05-v1 (2,986 English words), log D24–D36, register amendment, census, verifier, result page |
— | $0.00 | charter §3, A4 | Lead translation and lead reading are free and are never ledgered |
Spent $0.0791180 of a declared $0.50 — 15.8%. Waste $0.00: no dead body, no re-dispatch, no discarded call; the unspent 84% is a run that was cancelled by its own critic rather than a run that came in cheap.
Key reconciliation. 96.792474230 → 96.871592230, delta 0.079118000 against a per-response sum of 0.079118000, agreeing to 1e-9. Between-session drift into this session's opening snapshot was +0.008674400 (S164 closed at 96.783799830), the smallest recorded since the drift began to be tracked.
UTC day 2026-08-12: $0.17249665 of $5.00 across three sessions — S163 $0.0235655, S164 $0.06981315, S165 $0.0791180. 96.6% of the day's cap unspent.
2026-08-12 — S166 (E-20260812d-slot-or-carrier)
Pre-flight, written before dispatch. Five stages, 295 calls at the dearest admissible assignment, declared ceiling $1.60, raised to $2.10 on the pre-run critic's finding 5 before any Stage-0 call, with a registered hard stop: if cumulative spend after Stage S exceeds $0.60, Stage T runs with one hand only and every two-hand claim is withdrawn. The stop did not fire — spend after Stage S was $0.1199255. Today's earlier sessions had spent $0.17249665; the day's headroom was $4.82750335. Opening key snapshot 96.891523430.
One declared deviation, made after Stage 0 and before Stage S. The §8 table priced Stage T at a
700-token output cap while run.py carried 900–1,600, so the stated worst case was not the worst
case the code permitted — note (abc), the failure mode of the only estimate this project has
overrun. The two admitted hands' caps were brought down to the priced figures and the reason was
written into run.py at the point of change. The ceiling did not move.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-12 | S166 | pre-run critic over the frozen design, items.json and build_items.py (qwen/qwen3.7-max, reserve slug, cap 14,000) |
0.075 line | 0.027901100 | per-response | NEEDS-AMENDMENT, 6 findings, 2 BLOCKING; 5 accepted in full, 1 in part with the overrule written. finish_reason: stop |
| 2026-08-12 | S166 | Stage 0 — admission, 5 candidate seats × 6 control items | 0.09 line | 0.031868044 | per-response | Excluded deepseek/deepseek-v4-pro on a rule registered before dispatch |
| 2026-08-12 | S166 | Stage S — source side, 2 admitted arbiters × 20 candidate windows | 0.28 line | 0.060156400 | per-response | Every parity gate passed; selection applied before any hand was paid |
| 2026-08-12 | S166 | Stage T — hands, 2 hands × 39 | 0.35 line | 0.103793850 | per-response | 0 empty bodies, 0 re-dispatches |
| 2026-08-12 | S166 | Stage E — English side, 2 arbiters × 42 items | 0.58 line | 0.137764200 | per-response | 154 of 154 rating bodies parsed |
| 2026-08-12 | S166 | post-hoc byte-identical English null, unregistered, licensing nothing | — | 0.010312800 | per-response | Added after G3 failed, to separate arbiter false-alarm (0 of 12) from hand variation |
| 2026-08-12 | S166 | All lead work — «祝盃» whole and its log; the 26 built items; the design, analysis, verifier and result page; six NDL page images read and collated | — | $0.00 | charter §3, A4 | Lead translation and lead reading are free and are never ledgered |
Spent $0.371796394 of a declared $2.10 — 17.7%. Waste $0.00: 245 bodies, every one
finish_reason: stop, no dead body, no re-dispatch, no error, no discarded call. The run was
withheld by a gate, not truncated by one: every call bought a number that is printed.
Key reconciliation. 96.891523430 → 97.263319822, delta 0.371796392 against a per-response sum of 0.371796394, agreeing to 2e-9 — the fourth exact reconciliation in a row. Between-session drift into this session's opening snapshot was +0.019931200 (S165 closed at 96.871592230), the largest recorded since drift began to be tracked; noted for watching, not acted on.
UTC day 2026-08-12: $0.544293044 of $5.00 across four sessions — S163 $0.0235655, S164 $0.06981315, S165 $0.0791180, S166 $0.371796394. 89.1% of the day's cap unspent.
2026-08-12 — S167 (E-20260812e-dose)
Declared ceiling $2.00, printed by run.py --dry-run from max_tokens and the exact prompt
length of every pair (worst case $1.947759). Opening key snapshot 97.476969602; day spend before
this session $0.17249665 + $0.371796394 = $0.544293044, headroom $4.455706956.
| date | session | what | pre-flight | actual | source | notes |
|---|---|---|---|---|---|---|
| 2026-08-12 | S167 | pre-run critic over the frozen design, code.py, the prompt template and one built ladder (qwen/qwen3.7-max, reserve slug, not a judging seat, cap 16,000) |
0.096 line | 0.083555800 | per-response | NEEDS-REDESIGN, 3 findings, 1 BLOCKING; 2 accepted in full, 1 overruled in writing. finish_reason: stop |
| 2026-08-12 | S167 | latency probe of x-ai/grok-4.5 at reasoning: {effort: low}, 1 call |
— | 0.006184400 | per-response | Bought the measurement that motivated the deviation; licenses nothing |
| 2026-08-12 | S167 | 18 P3 bodies at effort: low, DELETED in the repair |
— | 0.074160000 | per-response, measured before deletion | Waste. §7.1 of the result page; note (bmp) |
| 2026-08-12 | S167 | the study — 100 pairs + 8 duplicate pairs × 3 seats | 1.70 line | 1.001306300 | per-response | 324 bodies, 0 dead, 0 truncated, 0 unparsable, 0 re-dispatches |
| 2026-08-12 | S167 | billed with no surviving record — two dispatch processes racing on one output directory | — | 0.407307700 | key-usage residual | Waste. Both processes billed the same slots; only the later writer's record survives |
| 2026-08-12 | S167 | All lead work — «Hastrman» whole and its 33-site log; the design, builder, runner, analyser, verifier and result page; the framework and sense-page writing | — | $0.00 | charter §3, A4 | Lead translation and lead reading are free and are never ledgered |
Spent $1.572514200 of a declared $2.00 — 78.6%. Waste $0.481467700, 30.6% of billed spend, and
it is the worst fraction since S162's 21.7%. Unlike that one it bought no finding at all: it is
the deleted low-effort bodies plus the unrecorded duplicate dispatches of a race between two runners.
The 324 bodies the result rests on were themselves clean — every one finish_reason: stop, no dead
body, no re-dispatch, no error.
Key reconciliation, and it is NOT exact, by design of the defect. 97.476969602 → 99.049483802, delta 1.572514200 against a per-response sum with surviving records of 1.165206500; the residual +0.407307700 is the race and is ledgered as spend rather than written off, because the records that would price it were destroyed by the process that made it. The ledgered session total is therefore the key-usage delta, not the per-response sum — the first time this project has had to prefer the delta, and the reason is on the result page and in note (bmp).
Between-session drift into this session's opening snapshot was +0.213649780 (S166 closed at 97.263319822) — an order of magnitude above the +0.0199 S166 recorded as its largest. Noted for watching, not acted on; the key's all-time figure includes non-project spend.
UTC day 2026-08-12: $2.116807244 of $5.00 across five sessions — S163 $0.0235655, S164 $0.06981315, S165 $0.0791180, S166 $0.371796394, S167 $1.572514200. 57.7% of the day's cap unspent.
2026-08-12 (UTC) — S168, E-20260812f-affect-unprompted
Declared ceiling $1.40; key-usage delta $1.424037082 — OVERRUN by $0.024037082, 1.7%, and ledgered as spent. Opening key snapshot 99.134932802, closing 100.558969884.
| live bodies, per-response sum | $0.688137460 (170 bodies, 0 dead, 1 truncated and superseded) |
| superseded bodies still on disk | $0.316460693 (51 of 63; see below) |
pre-run critic (qwen/qwen3.7-max) |
$0.054144300 |
| reconciled subtotal | $1.058742454 |
| unreconciled residual | $0.365294628 |
47.9% of billed spend bought nothing, the worst fraction this ledger has recorded. The cause is note (bmb), which existed and was not applied: two reasoning-capable judging seats were given a 300-token cap with no effort pinned and returned 26 of 30 and 20 of 30 empty bodies. The residual has a second cause — superseded bodies were moved twice into one flat directory and the second move overwrote 12 of the first move's records, destroying their costs — which is a new clause on note (bmp)'s remedy (iii).
Lead translation of 677 Bulgarian words into three English renderings: $0, not ledgered (charter §3, A4).
UTC day 2026-08-12 running total: $3.540844326 of $5.00 across six sessions — S163 $0.0235655, S164 $0.06981315, S165 $0.0791180, S166 $0.371796394, S167 $1.572514200, S168 $1.424037082. 29.2% of the day's cap unspent.
2026-08-12 (UTC) — S169, E-20260812g-foreign-italic
Declared ceiling $0.30; key-usage delta $0.160294400 — 53.4% of it, and the reconciliation is EXACT. Opening key snapshot 100.558969884, closing 100.719264284.
| cost | basis | ||
|---|---|---|---|
| pre-run critic over the frozen design | openai/gpt-5.6-terra, cap 14,000, 1 call, finish_reason: stop |
$0.032721 | per-response |
| blind correspondence coder 1 | google/gemini-3.6-flash, cap 4,000, effort: low, 25/25 parsed |
$0.067905 | per-response |
| blind correspondence coder 2 | x-ai/grok-4.5, cap 4,000, effort: low, 25/25 parsed |
$0.059668400 | per-response |
| total | $0.160294400 |
3 bodies, 0 dead, 0 truncated, 0 re-dispatches, no waste. Key reconciliation exact: delta
0.160294400 against a per-response sum of 0.160294400, residual 0. Note (bmb) was applied
before dispatch rather than after — both reasoning-capable coder seats had effort pinned and a
4,000-token cap in the first dispatch.
One pre-flight miss, inside the ceiling and recorded: the coders' prompt was estimated at ~10k tokens and was ~30k, because each of the 25 items carries a whole Russian paragraph. The declared ceiling was not touched; the estimate's method was wrong even though its total held.
Lead translation of 1,533 source words into Russian — the project's first — is $0 and is not ledgered (charter §3, A4). So are the corpus build, the census, the 74 correspondence calls and the verifier.
UTC day 2026-08-12 running total: $3.701138726 of $5.00 across seven sessions — S163 $0.0235655, S164 $0.06981315, S165 $0.0791180, S166 $0.371796394, S167 $1.572514200, S168 $1.424037082, S169 $0.160294400. 26.0% of the day's cap unspent.
2026-08-12 (UTC) — S170, RS-20260812h-dakghar-grade
$0.00. Lead translation of section ৩ of «ডাকঘর» and a whole-play census over frozen artifacts and a free public-domain comparator. No call was planned and none was made; no pre-flight was written.
2026-08-12 (UTC) — S171, E-20260812i-footing-channel
Declared pre-flight $0.19; key-usage delta $0.187374200 — 98.6% of it, and the reconciliation is NOT exact. Opening key snapshot 101.030785984, closing 101.218160184.
| cost | basis | ||
|---|---|---|---|
| pre-run critic over the frozen design | qwen/qwen3.7-max (reserve, not a judging seat), cap 12,000, 1 call, finish_reason: stop, provider Alibaba |
$0.0431054 | per-response |
| contrast hand, «Унтер Пришибеев» | openai/gpt-5.6-terra, cap 5,000, stop |
$0.0162297 | per-response |
| contrast hand, «Унтер Пришибеев» | x-ai/grok-4.5, cap 5,000, stop |
$0.0151348 | per-response |
| contrast hand, «Злоумышленник» | openai/gpt-5.6-terra, cap 5,000, stop |
$0.01641975 | per-response |
| contrast hand, «Злоумышленник» | x-ai/grok-4.5, cap 5,000, stop |
$0.0342964 | per-response |
| per-response sum | $0.125186050 | ||
| unattributed residual | billed with no surviving record | $0.062188150 | key-usage delta |
| ledgered total | $0.187374200 | key-usage delta |
The residual has a known cause and it is waste. The first run_hands.py dispatch was launched
in the foreground, hit the harness's 120-second tool timeout and was killed after it had already
billed at least one translation call whose response was never written to disk — the same failure
mode as S167's racing dispatch processes, arriving by a different route. Rule taken forward: an
API dispatch script is launched detached from the first attempt, never in the foreground. 5 bodies
survive, 0 dead, 0 truncated, 0 re-dispatches among them.
Lead translation of «Унтер Пришибеев» whole, twice — 1,225 Russian words to 1,777 and 1,785 English — is $0 and is not ledgered (charter §3, A4). So are the source builds, both censuses, all seven hand-story codings, the dependence checks and the 331-check verifier.
UTC day 2026-08-12 running total: $3.888512926 of $5.00 across nine sessions — S163 $0.0235655, S164 $0.06981315, S165 $0.0791180, S166 $0.371796394, S167 $1.572514200, S168 $1.424037082, S169 $0.160294400, S170 $0.00, S171 $0.187374200. 22.2% of the day's cap unspent.
2026-08-13 (UTC) — S172, E-20260813a-world-dose
Declared ceiling $2.30 (design §7.1, amendment A1); key-usage delta $1.319019290 — 57.3% of the
ceiling. Opening key snapshot 101.298362335, closing 102.617381625.
| cost | basis | ||
|---|---|---|---|
| pre-run critic over the frozen design | qwen/qwen3.7-max (reserve, not a judging seat), cap 16,000, finish_reason: stop, provider Alibaba |
$0.058617090 | per-response |
| 240 study + 12 duplicate judgments | openai/gpt-5.6-terra · google/gemini-3.6-flash (effort low) · x-ai/grok-4.5 (default effort), caps 350 / 1,000 / 500 |
$1.237719200 | per-response, 252 of 252 recorded |
| per-response sum | $1.296336290 | ||
| unattributed residual | billed with no surviving record | $0.022683000 | key-usage delta |
| ledgered total | $1.319019290 | key-usage delta |
The residual has a known cause and it is note (blf) in a third mode. The first pre-run critic
dispatch was launched with nohup … & inside a shell the harness had already backgrounded; the
wrapper returned immediately, the Python process died with its parent, and no body was written.
nohup inside a backgrounded shell is not detachment. The second attempt, launched through the
harness's own background mechanism, completed in four minutes and is the critic on record. Every
other call this session has a body on disk.
What the lock file bought. run.py refuses to start while another instance holds runs/_lock.
No second dispatcher ever started and there is no racing-process residue of the kind that cost
RS-20260812e 30.6% of its spend; the waste here is 1.7%.
The pre-flight was conservative in the right direction for the first time. P3 was priced at
2,000 completion tokens rather than its max_tokens of 500, because RS-20260812e §7.1 measured
that the cap does not bind reasoning on xAI. Observed completions on that seat ran 984–1,413 tokens,
so the ceiling was real and not a formality — note (abc), third mode.
Lead translation of «Ἡ Σταχομαζώχτρα» whole — 2,974 Greek words to 3,570 English — is $0 and is not ledgered (charter §3, A4). So are the source build, the 40-row site table, the contamination measurement, the build (100 checks) and the 1,679-check verifier.
UTC day 2026-08-13 running total after S172: $1.319019290 of $5.00 across one session — S172 $1.319019290.
2026-08-13 (UTC) — S173, E-20260813b-affect-yardstick
Declared ceiling $3.00 (design §7.1, raised from $2.00 → $2.20 → $3.00 as the pre-flight was re-priced and the pre-run critic's BLOCKING finding added a third document — each step recorded before dispatch, none after). Key-usage delta $1.482121752 — 49.4% of the ceiling. Opening key snapshot 102.635888025, closing 104.118009777.
| cost | basis | ||
|---|---|---|---|
| two pre-run critic passes over the frozen design | qwen/qwen3.7-max (reserve, not a judging seat), cap 16,000, both finish_reason: stop |
$0.107611575 | per-response |
| 329 study bodies across ten stages | openai/gpt-5.6-terra · google/gemini-3.6-flash (effort low) · x-ai/grok-4.5 · deepseek/deepseek-v4-pro (effort low), caps 400–4,000 |
$1.081059356 | per-response, 329 of 329 recorded |
| per-response sum | $1.188670931 | ||
| unattributed residual | billed with no surviving record | $0.293451 | key-usage delta |
| ledgered total | $1.482121752 | key-usage delta |
The residual has one cause and it is a defect in the runner, quantified rather than estimated.
377 dispatches produced 329 bodies: 48 calls were truncated on the first attempt and re-dispatched
at double the cap, and dispatch() records the successful attempt's cost, overwriting the
discarded one. 48 × roughly $0.006 accounts for $0.29 of a $0.293451 residual. The remedy is one
line — accumulate cost across attempts — and it is note (bmz).
Note (bmb) fired in a new mode, and the design had already applied its remedy. The source-side seat's effort was pinned low in the first dispatch, exactly as the note prescribes, and the cap was still too small: pinned-low reasoning ran 640–1,158 tokens against an 800-token cap, so 35 of 43 yardstick calls needed a second dispatch. Pinning the effort does not size the cap.
The pre-flight was conservative in the right direction again. Worst case was priced at
$2.807070 including a 4× routing surcharge on deepseek/deepseek-v4-pro (config/models.md, the
2026-07-25 caution); billed spend came in at 52.8% of it, and that seat routed across 16
providers inside this one run without ever exceeding its priced per-call worst case.
A gap between sessions, recorded rather than absorbed. S172's closing snapshot is
102.617381625 and this session's opening snapshot is 102.635888025 — $0.018506400 billed between
the two, with no record in either session. The key's all-time figure includes non-project spend
(CLAUDE.md), so this is noted and not attributed to either session's ledger.
Lead translation of «Căldură mare» whole, twice — 916 Romanian words to 1,142 and 1,190 English — is $0 and is not ledgered (charter §3, A4). So are the source build, the segment cut, the contamination measurement and the 4,267-check verifier.
UTC day 2026-08-13 running total: $2.801141042 of $5.00 across two sessions — S172 $1.319019290, S173 $1.482121752. 44.0% of the day's cap unspent.
2026-08-13 (UTC) — S174, D-20260813-17 ratification + E-20260813c-ennoblement-direction
Declared session ceiling $1.40 (gate ≤ $0.75, run ≤ $1.20 as designed). Key-usage delta $0.398367107 — 28.5% of the ceiling. Opening key snapshot 104.118009777, closing 104.516376884.
| cost | basis | ||
|---|---|---|---|
D-20260813-17 adversarial review, first dispatch VOID |
qwen/qwen3.7-max (non-panel), cap 6,000, finish_reason: length, 6,000 of 6,002 completion tokens spent on hidden reasoning, empty body |
$0.049151425 | per-response |
D-20260813-17 adversarial review, re-dispatch |
same slug, --reasoning-effort low --max-tokens 9000 |
$0.041858140 | per-response |
D-20260813-17 routed panel vote |
moonshotai/kimi-k3 (P4), effort low |
$0.069864000 | per-response |
pre-run critic over the frozen E-20260813c design |
x-ai/grok-4.5 (P3, no role in the jury) |
$0.017736400 | per-response |
| 42 study bodies, 54 dispatches | openai/gpt-5.6-terra · google/gemini-3.6-flash · deepseek/deepseek-v4-pro, all effort low, caps 1,500→6,000 |
$0.219757145 | per-response, 54 of 54 recorded |
| sum of parts | $0.398367110 | ||
| key-usage delta | $0.398367107 | ||
| unattributed residual | $0.000000003 |
The residual is three billionths of a dollar, and that is the point of the entry.
RS-20260813b carried a 19.8% unattributed residual because dispatch() recorded the
successful attempt's cost and overwrote the discarded one. This runner accumulates cost across
attempts, which is note (bmz)'s one-line remedy, and on 54 dispatches — 12 of them
re-dispatches — the ledger reconciles to the ninth decimal. (bmz)'s money half is discharged.
Waste this session: $0.049151425, 12.3% of spend, fully attributed. One void review dispatch, note (b)/(bmb) firing at default reasoning effort. The reasoning trace was substantive and nearly complete and was not read for a verdict — an interrupted deliberation is not a verdict. The remedy that worked is the note's own paired one: pin the effort and size the cap to the pinned effort's observed output.
(bmb) fired a second time, on a different seat, and cost the run its primaries.
deepseek/deepseek-v4-pro returned finish_reason: length on 5 of its 14 sites, still
unparseable after re-dispatch at a 6,000-token cap with effort pinned low. Those 5 dead bodies are
11.90% of 42 and fired F3. The cheapest seat on the panel was not cheap; it was unusable, and
it voided a run that had already passed every other gate.
Pre-flight conservative again. Worst case priced at ≈$1.10 for the run with a 4× routing surcharge on every seat; billed spend came in at 20.0% of it. Twelve providers were routed across in 42 calls (OpenAI, Google, Google AI Studio, GMICloud, StreamLake, Novita, DigitalOcean, CoreWeave, Alibaba, Parasail, BaseTen, Baidu) with no call exceeding its priced per-call worst case.
No gap between sessions this time. S173's closing snapshot 104.118009777 is exactly this session's opening snapshot — the first day this month with no unexplained inter-session billing.
Lead translation of «Flipperne» whole, twice — 767 Danish words to 873 and 1,070 English — is $0
and is not ledgered (charter §3, A4). So are the R25 regime, the materials dependence gate, the
contamination measurement, and the 1,057-check verifier.
UTC day 2026-08-13 running total: $3.199508149 of $5.00 across three sessions — S172 $1.319019290, S173 $1.482121752, S174 $0.398367107. 36.0% of the day's cap unspent.
2026-08-13 (UTC) — S175, E-20260813d-first-span-again-dakghar
Declared session ceiling REVISED from $0.70 to $1.20 before any seat call. Key-usage delta $0.251101700 — 20.9% of the revised ceiling. Opening key snapshot 104.834555784, closing 105.085657484.
| cost | basis | ||
|---|---|---|---|
| pre-run critic over the frozen design, Amendment 1, the completed coding and the completed measurement | x-ai/grok-4.5 (P3, no role in the seat run), cap 8,000, effort not pinned |
$0.071492400 | per-response |
| 84 seat bodies, 84 dispatches | openai/gpt-5.6-terra · google/gemini-3.6-flash · x-ai/grok-4.5, all effort low, cap 1,500 |
$0.179609300 | per-response, 84 of 84 recorded |
| sum of parts | $0.251101700 | ||
| key-usage delta | $0.251101700 | ||
| unattributed residual | $0.000000000 |
The ceiling was revised upward, before spending, because note (abc) fired on this design's own cost table. Frozen §9 priced the seats at "800 output tokens each" — an expected length, which is precisely what the note forbids. The revision is built from the cap actually sent (1,500) plus a global allowance of ten re-dispatches: 84 × 1,500 + 10 × 3,000 output tokens at the dearest seat rate. Zero re-dispatches were used, and billed spend came in at 20.9% of the revised figure.
Waste this session: $0.00. 85 bodies, 0 dead, 0 re-dispatched, 0 void, 84 of 84
finish_reason: stop — the first clean completeness record in four sessions. P5
deepseek/deepseek-v4-pro was not asked to do anything, per NEXT.md's standing finding after
S174's five dead bodies, and the seat question was plain text rather than structured output.
Effort pinning, both ways in one session. Note (b)/(bmb) pins reasoning effort low for cheap
mechanical seat tasks, and that is what the 84 one-letter verdicts got. The critic was left at the
provider default on purpose — a critic pass is the one call in a run whose whole value is depth — and
it returned ten findings, six of them blocking, for 28% of the session's spend. It changed the
result: two of the six blocking findings struck salvage labels the lead had already written into its
own analysis.
Between-session drift: +0.318178900. S174's closing snapshot is recorded as 104.516376884 and this session opened at 104.834555784. Recorded, not explained; the second largest such figure this month after S159's +0.293605. Within the session the reconciliation is EXACT to the ninth decimal.
Providers across 85 calls: OpenAI 28, xAI 29, Google 17, Google AI Studio 11 — P2 split across two
Google endpoints inside a single stratum. No call billed above its declared per-call worst case.
Lead re-translation of «ডাকঘর» span A whole — 1,147 Bengali words to 1,638 English — is $0 and is
not ledgered (charter §3, A4). So are the whole-play ক্ষেপ- census, the three-way overlap
measurement, the causal coding, and the 156-check verifier.
UTC day 2026-08-13 running total: $3.450609849 of $5.00 across four sessions — S172 $1.319019290, S173 $1.482121752, S174 $0.398367107, S175 $0.251101700. 31.0% of the day's cap unspent.
2026-08-13 (UTC) — S176, E-20260813e-slot-typology-ja
Declared session ceiling $0.60, built from the caps actually sent. Key-usage delta $0.502906385 — 83.8% of the ceiling. Opening key snapshot 105.090588884, closing 105.593495269.
| cost | basis | ||
|---|---|---|---|
| pre-run critic over the frozen design | qwen/qwen3.7-max (off-panel; produced no hand and did no coding), cap 12,000, effort not pinned |
$0.053251925 | per-response |
| 2 whole-story hands | openai/gpt-5.6-terra · x-ai/grok-4.5, cap 6,000, effort not pinned |
$0.066706650 | per-response, 2 of 2 |
| coding, first seat, 6 dispatches for 5 usable bodies | google/gemini-3.6-flash, cap 4,000 → 10,000, effort low |
$0.163075500 | per-response |
| coding, cross seat, 5 bodies | openai/gpt-5.6-terra · x-ai/grok-4.5, cap 5,000, effort low |
$0.129110400 | per-response |
| mutation positive control, 2 bodies | google/gemini-3.6-flash · openai/gpt-5.6-terra |
$0.054786750 | per-response |
| 1 dead body | deepseek/deepseek-v4-pro, cap 6,000, effort low |
$0.018416160 | per-response |
| sum of parts (16 stored bodies) | $0.459385385 | ||
| key-usage delta | $0.502906385 | ||
| unattributed residual | $0.043521000 |
The residual is fully identified and is not a mystery: it is the first code_SHAW_P2 call, which
truncated on a 4,000-token cap and whose stored body was overwritten by its re-dispatch. Counting it,
reconciliation is exact to the ninth decimal. This is note (bmz) part (ii) recurring in a
milder form — a re-dispatch overwriting the discarded attempt's record — in a hand-run script rather
than in a dispatcher.
Waste this session: $0.061937160, 12.3% of the spend. Two items, both known failure modes and both now carrying method notes:
deepseek/deepseek-v4-proreturnedcontent: nullafter spending its whole 6,000-token cap on 21,497 characters of hidden reasoning, witheffort: lowset and ignored,finish_reason: length, provider DigitalOcean.NEXT.md's standing finding against this seat was scoped to structured output; this was plain text and it failed identically. New note (bne): the seat is not usable for judging or coding on any task shape, and an effort pin is not a cost bound.google/gemini-3.6-flashtruncated twice at a 4,000-token cap, its hidden reasoning consuming the cap before the answer began. One was re-dispatched at 10,000 (the residual above); the other, on a hand excluded from every primary, was left partial.
The dead seat changed the design rather than the budget. With P5 gone and every remaining
affordable seat also a translator in the same run, the second coder was cross-assigned — each
seat codes every hand but its own — which cost nothing extra and is a stricter arrangement than the
frozen design had.
Lead translation is free and is not ledgered (charter §3, A4): Akutagawa's 「煙管」 sections 三–八, 4,617 Japanese characters to 2,442 English words, completing the work. So are the 51-site source census, the overlap testing, the coder-independent vocative count, and the 25-check verifier.
Between-session drift: +0.004931400. S175's closing snapshot is recorded as 105.085657484 and this session opened at 105.090588884 — two orders of magnitude below last session's +0.318178900. Recorded, not explained.
UTC day 2026-08-13 running total: $3.953516234 of $5.00 across five sessions — S172 $1.319019290, S173 $1.482121752, S174 $0.398367107, S175 $0.251101700, S176 $0.502906385. 20.9% of the day's cap unspent.
2026-08-13 (UTC) — S177, E-20260813f-affect-confound step 1 (pre-run critic only)
Declared session ceiling $0.35, built from the cap actually sent (9,000 completion at
moonshotai/kimi-k3's $15.00/M, doubled for one re-dispatch). Key-usage delta $0.095881200 —
27.4% of the ceiling. Opening key snapshot 105.610437869, closing 105.706319069.
| cost | basis | ||
|---|---|---|---|
independent pre-run critic on E-20260813f |
moonshotai/kimi-k3 (P4), effort low, cap 9,000, provider Together, finish_reason: stop |
$0.095881200 | per-response usage.cost |
| sum over stored bodies (1) | $0.095881200 | ||
| key-usage delta | $0.095881200 | ||
| unattributed residual | $0.000000000 |
Reconciliation is exact to the ninth decimal, with no re-dispatch and no dead body. The
F6 guard did not fire; the paired remedy of note (bmb) — pin the effort and size the cap to
the pinned effort's observed output — was applied at dispatch and the single attempt returned a
complete 4,445-character body. Waste this session: $0.00, 0%.
The session's principal spend was deliberately not made. E-20260813f's judging run has a
cap-built worst case of $2.684 after the critic's remedies, against $1.046484 of headroom
on this UTC day. continue-prompt.md §7 gives three options — split, scale down, defer — and
defer was taken over scale-down because every bar in the design is registered against 15
segments and shrinking the material would have moved them. The run is frozen and dispatchable;
step 2 declares $2.80 and, if that day's rows do not have it, splits at the stage boundary
(stages 1–4 worst case $1.39) rather than scaling.
Lead translation is free and is not ledgered (charter §3, A4): Kunikida Doppo's 「星」, 2,157 Japanese characters rendered whole twice — 1,534 and 1,481 English words — with both translator's logs, the segmentation of all three texts, and the arm-vs-arm dependence measurement.
Between-session drift: +0.016942600. S176's closing snapshot is recorded as 105.593495269 and this session opened at 105.610437869. Larger than last session's +0.004931400 and smaller than S175's +0.318178900. Recorded, not explained.
UTC day 2026-08-13 running total: $4.049397434 of $5.00 across six sessions — S172 $1.319019290, S173 $1.482121752, S174 $0.398367107, S175 $0.251101700, S176 $0.502906385, S177 $0.095881200. 19.0% of the day's cap unspent.
2026-08-13 (UTC) — S178, E-20260813g-register-quadrants, ARM-ennoblement step 2
Declared session ceiling $0.55 after the pre-run critic's amendment, built from the caps actually sent with this project's standing 4× routing surcharge. Key-usage delta $0.219395200 — 39.9% of the ceiling. Opening key snapshot 105.714590869, closing 105.933986069.
| cost | basis | ||
|---|---|---|---|
pre-run critic on E-20260813g |
moonshotai/kimi-k3 (P4), cap 6,000, effort low, provider Modal, stop |
$0.081607200 | per-response usage.cost |
LAT etymology classification, 758 types |
openai/gpt-5.6-terra (P1), cap 6,000, effort low, length on BOTH attempts — DEAD |
$0.076043850 | accumulated across 2 attempts |
CMP-1 blind formal coding, 144 cells |
openai/gpt-5.6-terra (P1), cap 4,000, effort low, stop, 618 chars |
$0.024289750 | per-response usage.cost |
CMP-2 blind formal coding, 144 cells |
x-ai/grok-4.5 (P3), cap 4,000, effort low, stop, 550 chars |
$0.037454400 | per-response usage.cost |
| sum over stored bodies (4 calls, 5 dispatches) | $0.219395200 | ||
| key-usage delta | $0.219395200 | ||
| unattributed residual | $0.000000000 |
Reconciliation is exact to the ninth decimal, and the cost-accumulation repair held across the one re-dispatch.
Waste this session: $0.076043850 — 34.7%, and it is a single body. LAT asked one seat to
classify all 758 union types and return two lists; both attempts hit the cap. The cap was sized
from a guess about how long the lists would be, on a task shape the project had never run. F6
struck the measure it supplied, F3 reduced Axis E to one measure, and F1 then withheld half the
run's registered predictions — one mis-sized cap cost an entire axis. Note (bng) is written
from it and its remedy is a closed output form, which is what the two sibling calls used: 618 and
550 characters at a 4,000 cap over the same six texts.
The four RS-20260813c §10 jury repairs were NOT bought and are not deferred silently: they
need roughly $1.40 and the day had $0.95. They are a live row in wiki/backlog.md, opened S178.
Lead translation is free and is not ledgered (charter §3, A4): Andersen's «Flipperne», 756
Danish words rendered whole a third time under R08 resistancy, with the translator's log, the
BRAK OCR reconstruction, the fifteen-pair dependence table, the stage-1 census, the lead audit
column and the 169-check verifier.
Between-session drift: +0.008271800. S177's closing snapshot is recorded as 105.706319069 and this session opened at 105.714590869. Smaller than S177's +0.016942600. Recorded, not explained.
UTC day 2026-08-13 running total: $4.268792634 of $5.00 across seven sessions — S172 $1.319019290, S173 $1.482121752, S174 $0.398367107, S175 $0.251101700, S176 $0.502906385, S177 $0.095881200, S178 $0.219395200. 14.6% of the day's cap unspent.
2026-08-13 (UTC) — S179, E-20260813h-carriage-or-elevation, ARM-carriage-direction step 1
Pre-flight, written into the frozen design before dispatch and built from max_tokens, not from
an expected answer length (note (abc)). Declared ceiling $0.70, sized to the $0.731207366 of
headroom seven earlier sessions left on the day. Worst case at the pilot-confirmed caps: critic
P4 cap 4,000 $0.072; 15 P1 cells cap 1,200 $0.119; 15 P2 cells cap 2,500 $0.152; 15 P3 cells
cap 1,200 $0.131 — $0.474 for 46 calls. P2's cap is higher than its siblings' because
config/models.md records ~1.7k tokens of hidden reasoning on that seat, and sizing it at 1,200
would have been note (bng)'s defect committed the day after it was written.
Note (bng)'s remedy was applied as a measurement rather than a guess, which is the whole point of it. Three pilot cells were dispatched first and their completion tokens read — 444, 243 and 273 against caps of 1,200, 2,500 and 1,200 — before the remaining 42 were sent. The pilot cells were kept as real cells rather than discarded.
| item | seat, cap, outcome | cost | method |
|---|---|---|---|
| pre-run critic | moonshotai/kimi-k3 (P4), cap 4,000, effort low, stop |
$0.071944200 | per-response usage.cost |
| 45 jury cells | P1 cap 1,200 · P2 cap 2,500 · P3 cap 1,200, effort low, 45 of 45 stop |
$0.109177650 | accumulated per cell |
| sum over stored bodies (46 calls, 46 dispatches) | $0.181121850 | ||
| key-usage delta | $0.181121850 | ||
| unattributed residual | $0.000000000 |
Opening snapshot 105.940286869, closing 106.121408719. Reconciliation exact to the ninth decimal, and no cell needed a second dispatch, so the cost-accumulation repair was not exercised.
Waste this session: $0.000000000. No dead body, no truncation, nothing discarded — the first session in four days with a waste figure of zero, and the reason is the pilot rather than luck. Spent 25.9% of the declared ceiling.
Between-session drift: +0.006300800. S178's closing snapshot is recorded as 105.933986069 and this session opened at 105.940286869. Smaller than S178's +0.008271800. Recorded, not explained.
Lead translation is free and is not ledgered (charter §3, A4): Andersen's «Flipperne» rendered
whole a fourth time, 756 Danish words under R07 fluency with its translator's log, together
with the four-arm dependence table, the S1 audit column, the distance-from-plain predictor, the
five-segment split and the 1,335-check verifier.
UTC day 2026-08-13 running total: $4.449914484 of $5.00 across eight sessions — S172 $1.319019290, S173 $1.482121752, S174 $0.398367107, S175 $0.251101700, S176 $0.502906385, S177 $0.095881200, S178 $0.219395200, S179 $0.181121850. 11.0% of the day's cap unspent.
UTC day 2026-08-14
S180 — 2026-08-14 (UTC), ARM-alf-layla step 1 (T1), E-20260814-saj-carriage
$0.000000000 spent. No API request was made, so no key snapshot was taken.
The session's principal unit had nothing for the API to do. The translation limb is the lead's own — «ألف ليلة وليلة» span A, 642 Arabic words into 1,201 English — which is free and is never ledgered (charter §3, A4). The study limb is a census over three stored texts: two of them public-domain translations fetched from Project Gutenberg over plain HTTPS, the third the lead's own. The coding, the tallies, the quotation verification and the dependence check are all arithmetic over stored bytes.
A $0 session is a normal outcome (continue-prompt.md §7), and this is the first of a fresh UTC
day, so the whole $5.00 cap stands unspent for whatever comes next. Two runs remain deferred on
money and not on design and both would fit: E-20260813f (~$2.80, ARM-affect-confound step 2) and
the ennoblement jury re-run (~$1.40, the backlog's one live row).
UTC day 2026-08-14 running total: $0.000000000 of $5.00 across one session.
S181 — 2026-08-14 (UTC), ARM-honorific-hands step 1 (T5), E-20260814b-honorific-hands
$1.109045850 spent. Key snapshot before 106.456584319, after 107.565630169, delta 1.109045850 — reconciled to the API-reported per-request costs plus the two truncated critic calls and one call billed but lost to a network error, and exact.
| what | cost |
|---|---|
pre-run critic P4 attempt 1 — max_tokens 4,000, all of it hidden reasoning, empty content |
$0.079782 |
pre-run critic P4 attempt 2 — max_tokens 14,000, all of it hidden reasoning, empty content |
$0.233043 |
pre-run critic P4 attempt 3 — dispatched and billed, response lost to IncompleteRead |
$0.213193 |
pre-run critic P4 attempt 4 — reasoning: {"effort": "low"}, full verdict, 8 findings |
$0.083293 |
two machine translation hands (P1, P3) |
$0.045102 |
gloss audit (P3, 67 items) |
$0.205360 |
ten blind coder calls (P2 ×5, P1c ×4, P3c ×1), 77 items each |
$0.249272 |
| total | $1.109046 |
Waste: $0.526018, 47.4%, all of it on one seat, and the cause is now method note (bnk): on a
reasoning seat max_tokens caps the answer plus the thinking and the thinking will take all of
it. Declared $0.60 in the frozen design; revised in writing to $1.60 during the run, before the
coder calls, with the reason recorded on the design page. Actual came in under the revised figure.
Lead translation is free and is not ledgered (charter §3, A4): 『源氏物語』「蓬生」 §§2‑3–2‑4, 1,841 characters of classical Japanese into 1,200 English words, an R06 single-pass draft and an R04 revision frozen in separate commits before the design existed, with a translator's log of eight decisions.
UTC day 2026-08-14 running total: $1.109045850 of $5.00 across two sessions — S180 $0.000000000, S181 $1.109045850. 77.8% of the day's cap unspent.
S182 — 2026-08-14 (UTC), ARM-affect-confound step 2 (T3), E-20260813f-affect-confound
$0.840025 spent on per-request costs — $0.816705 for the run (261 bodies, including $0.044710
of two discarded attempt sets) and $0.023320 on eleven transport probes. Key snapshot before
107.830757299, after 108.692808434, delta 0.862051135, leaving $0.022027
unreconciled in the over-counting direction. Per-request costs are primary (CLAUDE.md); the
residual is recorded, not chased.
Two reconciliation residuals in one day, both the same shape and both in the same direction. The snapshot read at the start of this session was 107.830757299 against S181's closing 107.565630169 — $0.265127130 that settled after S181 closed and after it had reconciled its own delta as exact. So the key figure lags: a delta taken immediately after a run under-counts, and one taken later picks up the previous session's stragglers. Treated conservatively for the day's cap, i.e. the key deltas are what is charged.
| what | cost |
|---|---|
yardsticks YP + YM, 30 calls (qwen, reasoning disabled) |
$0.037468 |
gates GJ 15 (glm), GA 15, GY 15 (qwen) |
$0.029941 |
directionality GD, 96 calls (three judges) |
$0.210418 |
style match GS, 90 calls (three judges) |
$0.494167 |
two discarded attempt sets, yp__S01 and yp__S14 |
$0.044710 |
eleven transport probes (enabled:false, effort:minimal, cap sizing) |
$0.023320 |
| total | $0.840025 |
Declared ceiling $2.80 in the frozen design; revised in writing to $3.65 worst case before
dispatch (design §12 A6 — caps re-sized from measured reasoning-token use), and the §8 stage
split was then taken because the day's headroom was $3.56: stages 1–4 (worst case $2.26)
dispatched, stage 5 (worst case $1.39) never dispatched, because F2b fired at 12 of 15 and
voids the whole run including the primary stage 5 feeds. Actual came in at 37% of the stage-1–4
worst case, the gap being the reasoning-disable in A6.
$0.044710 — 5.5% — bought nothing, both discards on the yardstick seat: note (bnk) firing a
second time on a third seat with its own remedy applied and ignored ($0.043586), and a bare-prose
body the runner's first usability test wrongly accepted ($0.001124). Notes (bnl), (bnm)
and (bnn) are what came out of it.
Lead translation is free and is not ledgered (charter §3, A4). This session translated nothing new — the arm's two renderings of 「星」 were made and frozen at S177, before the design that measures them existed — but it did read 15 segments of Meiji 擬古文 against the yardstick documents by hand to adjudicate the source-fidelity gate's twelve verdicts, at $0.
UTC day 2026-08-14 running total: $2.236224 of $5.00 across three sessions — S180 $0.000000, S181 $1.109046 plus $0.265127 settling late, S182 $0.862051 (key delta). 55.3% of the day's cap unspent.
S183 — 2026-08-14 (UTC), ARM-elevation-resolution step 1 (T4), E-20260814d-elevation-resolution
$0.420346 spent. Key snapshot before 108.880988254, after 109.301333954, delta 0.420345700 — exactly the per-request sum, residual 0.000000000. (The snapshot read at the start of this session was $0.188180 above S182's closing figure; that is S182's own lag note firing again and is not this session's spend.)
| what | cost |
|---|---|
pre-run critic, one call (moonshotai/kimi-k3, effort low, cap 6,000) |
$0.112383000 |
| stage 1 — the site list, 3 source-only annotation calls | $0.067790400 |
| stage 2 — 36 height bodies over 12 sites × 3 seats, plus 5 re-dispatches | $0.240172300 |
| total | $0.420345700 |
Declared ceiling $1.50, raised in writing from $1.42 before dispatch on the pre-run critic's finding 4, which showed the re-dispatch tail was unbounded and therefore that $1.42 was not a worst case. Actual 28.0% of it. Note (bnl)'s probe is why: one real call per seat measured 422 / 579 / 809 reasoning tokens and the stage cap was set from the measurement rather than from the 2,500 the pre-flight had priced.
$0 bought nothing. All 36 stage-2 bodies and all 3 stage-1 bodies were usable; the 5
truncations of a permitted 6 re-dispatches all recovered, and their first attempts are inside the
figures above. The cost worth naming as a lesson is not waste but a cap that did not bind: P3
returned 6,516 reasoning tokens against a stage-1 max_tokens of 4,000 and still delivered a
complete body, billing $0.045616 against a per-call worst case of $0.0272 built from the cap.
Note (abc)'s worst case is not a bound on a reasoning seat.
Lead translation is free and is not ledgered (charter §3, A4). This session rendered
«Flipperne» whole a fourth time (756 Danish → 866 English words) under the new R26, and ran the
copy-text re-fetch, the 21-pair dependence table, the mechanical span pool and the verifier, all at
$0.
UTC day 2026-08-14 running total: $2.656570 of $5.00 across four sessions — S180 $0.000000, S181 $1.109046 plus $0.265127 settling late, S182 $0.862051, S183 $0.420346. 46.9% of the day's cap unspent.
S184 — 2026-08-14 (UTC), ARM-carriage-direction step 2 (T2), E-20260814f-carriage-elevation-2
$0.263575750 spent (per-request sum over stored bodies, which is what is ledgered).
| what | cost |
|---|---|
pre-run critic, one call (moonshotai/kimi-k3, effort low, cap 8,000) — truncated at finish: length |
$0.144837000 |
| critic continuation call (cap 6,000), for the unfinished finding, further findings and the verdict | $0.035175000 |
45 jury cells, P1 cap 1,200 · P2 cap 1,400 · P3 cap 1,500, effort low, 45 of 45 stop |
$0.083563750 |
| total | $0.263575750 |
Declared ceiling $0.75; actual 35.1% of it. Waste $0.000000000 — no dead body, no re-dispatch, nothing discarded.
The reconciliation does NOT close, and the residual was measured rather than hunted. Key
snapshot before 109.335483504, after 109.716567314, delta 0.381083810 against a
per-request sum of 0.263575750 — unattributed $0.117508060. A second key read taken
immediately afterwards with no calls dispatched between the two returned 109.718302034, i.e.
the figure moved $0.001734720 while nothing was spent. That is a direct demonstration of the
late settlement NEXT.md recorded of S181, and it is now note (bns): when the residual is
non-zero, take the second read before writing anything about it. CLAUDE.md's rule governs —
per-request cost is primary, the delta is the sanity check.
The money lesson of the session is note (bnr). Raising the pre-run critic's cap from S183's 6,000 to 8,000 did not stop it truncating: the seat spent 6,085 of the 8,000 on hidden reasoning, so the visible budget was ~1,915 tokens at both caps. A $0.035 continuation call bought the missing finding, five further findings and the verdict line — and two of the five became accepted amendments. Budget a continuation; do not raise the cap.
Lead translation is free and is not ledgered (charter §3, A4). This session rendered the vision
of Bécquer's «El Miserere» whole, twice (668 Spanish → 725 and 712 English words) under R08
and R07, and ran the two-witness collation, the 24-site device census, six dependence
measurements, the elaboration predictor and the verifier, all at $0.
UTC day 2026-08-14 running total: $2.920146 of $5.00 across five sessions — S180 $0.000000, S181 $1.109046 plus $0.265127 settling late, S182 $0.862051, S183 $0.420346, S184 $0.263576. 41.6% of the day's cap unspent.
S185 — 2026-08-14 (UTC), ARM-alf-layla step 2 (T1), E-20260814g-supplied-sound
$0.469247200 spent (per-request sum over stored bodies, which is what is ledgered).
| what | cost |
|---|---|
pre-run critic, one call (x-ai/grok-4.5, cap 6,000) — returned stop, no continuation needed |
$0.048448400 |
144 census bodies, P1 cap 400 · P2 cap 1,400 · P3 cap 800, 144 of 144 stop, 2 re-dispatched |
$0.420798800 |
P1 $0.026969 · P2 $0.168219 · P3 $0.225611 |
|
| total | $0.469247200 |
Declared ceiling $0.77; actual 60.9% of it. Waste $0.000000000 — no dead body, 0 invalid
passages, nothing discarded. The two re-dispatches (P2 on PT1|lead and PT5|lead) both returned
clean and both attempts are stored.
The continuation budgeted under note (bnr) was not needed and is the note working. S184 raised
the critic cap to 8,000 and still truncated; this run set 6,000 on a different seat and budgeted
a continuation call that never had to be made. One call, finish_reason: stop, eleven findings and a
verdict for $0.048.
Note (bgk) fires a third time, on a third vendor family, and cost nothing. x-ai/grok-4.5
returned 2,167 completion tokens against a declared 800 cap, finish_reason: stop, body clean.
The seat's mean was 702, so the run landed at $0.4208 against a cap-literal worst case of $0.642 —
but the per-body worst case built from the cap was exceeded 2.7×. Nothing sees this except the
after-the-fact check.
Reconciliation, with the second read note (bns) requires. Key snapshot before
109.734332934; first read after 110.214326134; second read, taken immediately with no calls
between, returned 110.214326134 — identical to the digit, so nothing was settling in that window
and the residual is real rather than an artifact of timing. Delta 0.479993200 against a
per-request sum of 0.469247200: unattributed $0.010746000, 2.24%. It is the size of one
P3-shaped body at this run's rates and is recorded, not hunted. CLAUDE.md's rule governs —
per-request cost is primary, the delta is the sanity check.
Lead translation is free and is not ledgered (charter §3, A4). This session rendered span B of «ألف ليلة وليلة» — 582 Arabic words → 1,138 English, including 12 bayts of classical Arabic verse in three poems, the work's first — and ran the two-witness collation, the register update, the locus extraction and every verifier, all at $0.
UTC day 2026-08-14 running total: $3.389393 of $5.00 across six sessions — S180 $0.000000, S181 $1.109046 plus $0.265127 settling late, S182 $0.862051, S183 $0.420346, S184 $0.263576, S185 $0.469247. 32.2% of the day's cap unspent.
S186 — 2026-08-14 (UTC), ARM-honorific-hands step 2 (T5), no experiment
$0.000000000 spent. No API request was made, so no key snapshot was taken.
The session's principal unit had nothing for the API to do. The study limb is the rewrite of
framework/v0.2 §10 on RS-20260814b, run and paid for at S181 — transcription, not
re-derivation. The translation limb is the lead's own, and is free and never ledgered (charter §3,
A4): 『源氏物語』 ch. 15 「蓬生」 §3‑4 rendered twice, 1,227 characters → 836 English words under
R06 and 1,012 under the newly minted R27 footing-max, with a 52-site mechanical extraction and
every count recomputed by counts.py over stored bytes.
Nothing was deferred on money. The two designs framework/v0.2 §10.6 now names — a blind
panel pricing the R06/R27 pair, and a registered test of the volitionality account against the
two published hands already censused in RS-20260814b — are deferred on design, not on
budget; neither exists yet.
UTC day 2026-08-14 running total: $3.389393 of $5.00 across seven sessions — S180 $0.000000, S181 $1.109046 plus $0.265127 settling late, S182 $0.862051, S183 $0.420346, S184 $0.263576, S185 $0.469247, S186 $0.000000. 32.2% of the day's cap unspent.
S187 — 2026-08-14 (UTC), ARM-footing-price step 1 (T3), E-20260814h-footing-price
$0.760331700 spent, against a declared run ceiling of $0.90 and a day headroom of $1.610607. Key-usage snapshot before 110.410537934, after 111.170869634, delta $0.760331700 — exact to 1e-9 against the sum of the three stages.
| stage | seat(s) | calls | cost | bought |
|---|---|---|---|---|
| pre-run adversarial critic | P3 x-ai/grok-4.5 |
1 (stop, no continuation) |
$0.068592400 | NEEDS REDESIGN, 20 findings, a corrected primary |
| grading, first dispatch — KILLED by a 10-minute harness wall clock | P1/P2/P3 |
~57 | $0.312882500 | nothing |
| grading, second dispatch | P1/P2/P3 |
72 bodies (3 redispatched, 1 dead) | $0.378856800 | the run |
41% of the spend bought no data, and the cause was the runner, not the API: all 72 bodies were
held in memory and written once at the end, so a wall clock the script could not see destroyed
everything it had paid for. Fixed in the same session — bodies are appended to bodies.jsonl as
they return and a re-invocation resumes — and written up as method note (bnx). Nothing was
deferred on money.
Lead translation is free and is not ledgered (charter §3, A4): T-genji-yomogiu-R28-v1, the
third rendering of the same 1,227 characters under the newly minted R28 footing-selective, and the
ornament-matched decoy control, both cost $0.
UTC day 2026-08-14 running total: $4.149724700 of $5.00 across eight sessions — S180 $0.000000, S181 $1.109046 plus $0.265127 settling late, S182 $0.862051, S183 $0.420346, S184 $0.263576, S185 $0.469247, S186 $0.000000, S187 $0.760332. 17.0% of the day's cap unspent.
S188 — 2026-08-15 (UTC), ARM-elevation-resolution step 2 (T4), E-20260815-register-room
$0.321810800 spent, against a declared run ceiling of $1.20 and a fresh UTC-day cap of $5.00 with nothing spent before this session. Key-usage snapshot before 111.439242034, after 111.761052834, delta $0.321810800 — exact to 1e-9 against the sum of the five stages.
| stage | seat(s) | calls | cost | bought |
|---|---|---|---|---|
| pre-run critic, first call | P4 moonshotai/kimi-k3 |
1 (length, 0 characters) |
$0.100636350 | nothing |
| pre-run critic, reasoning-capped retry | P4 moonshotai/kimi-k3 |
1 (+1 HTTP 400, free) | $0.075334200 | NEEDS-AMENDMENT, 9 findings, 2 BLOCKING |
| stage 0 probe (stage 1 caps) | P1 P2 P3 |
3 | $0.012030900 | measured reasoning 165/681/825 |
| stage 1, site list from the Russian alone | P1 P2 P3 |
3 | $0.032604650 | 57 spans coded, 55 unanimous |
| stage 0 probe (stage 2 caps) | P1 P2 P3 |
3 | $0.010012650 | measured reasoning 312/390/494 |
| stage 2, blind register-height grading | P1 P2 P3 |
36 | $0.091192050 | the run, 36 of 36 usable |
31.3% of the spend bought no output token, and it is note (bnk) firing a third time with its
own remedy sent and ignored: reasoning: {"effort": "low"} at max_tokens 6,000 returned
finish_reason: length with zero visible characters. Note (bnr)'s continuation remedy was
inapplicable — a continuation needs a truncated critique to continue from. The retry that worked sent
an explicit reasoning: {"max_tokens": 2000} with the total at 9,000. Written up as note
(bny). Nothing was deferred on money.
Lead translation is free and is not ledgered (charter §3, A4): three complete renderings of
«Злоумышленник» under R06, R26 and R25 — 4,711 English words from 1,074 Russian — plus
counts.py, build_pool.py, the dependence table, the post-hoc analysis and the 1,475-check
verifier, all $0.
UTC day 2026-08-15 running total: $0.321810800 of $5.00 across one session — S188 $0.321811. 93.6% of the day's cap unspent.
S189 — 2026-08-15 (UTC), ARM-fluent-carriage step 1 (T2), E-20260815b-fluent-carriage
$0.694796750 spent, against a declared run ceiling of $1.50 and a day headroom of $4.678178. Key-usage snapshot before 111.761052834, after 112.455849584, delta $0.694796750.
| stage | seat(s) | calls | cost | bought |
|---|---|---|---|---|
| pre-run adversarial critic | P4 moonshotai/kimi-k3 |
1 (stop, no continuation) |
$0.135123000 | NEEDS-REDESIGN, 10 findings, 2 BLOCKING, a rebuilt arm |
AB carriage, FS vs ODD |
P1 P2 P3 |
22 (21 cells + 1 re-dispatch) | $0.119702550 | the primary, indeterminate at 13/21 |
AC carriage, FS vs FF |
P1 P2 P3 |
21 | $0.063195300 | the co-primary, 21/21 |
AD carriage, ODD vs FF |
P1 P2 P3 |
21 | $0.068926550 | the false-positive, 18/21 |
QB quality, FS vs ODD |
P1 P2 P3 |
21 | $0.061520800 | G1 21/21 |
QC quality, FS vs FF |
P1 P2 P3 |
21 | $0.062107600 | G2 parity passes at 8/21 |
QD quality, ODD vs FF |
P1 P2 P3 |
5 of 21 — stopped | $0.007759900 | a partial G3 |
LG language inference |
P4 |
9 | $0.060558000 | G5: all three arms named Japanese, no differential |
AU blind carriage audit |
P4 |
2 | $0.055085400 | G6 16 of 20, against the translator's own 81 of 83 |
PAR + PLANT content parity |
P4 |
10 | $0.055895250 | 3 of 3 plants caught, 0 of 7 segments SAME — the run's carriage half withheld |
$0.004922400 was billed and never received, and the cause is named rather than absorbed: the
judging runner was killed mid-call to protect session time, and the sixth QD request was in
flight. Note (bnx)'s append-and-resume saved everything already returned — the run survived
four interruptions with zero loss — but a call in flight at the moment of the kill is billed and
lost, which the note did not previously say. 0.7% of the spend.
The AS condition (21 cells) was not dispatched and QD was stopped at 5 of 21, both deliberate
wall-clock decisions taken before the affected bodies existed, both declared on the result page.
Nothing was deferred on money — the run finished at 46% of its declared ceiling.
Lead translation is free and is not ledgered (charter §3, A4): «熱い砂の上» rendered whole under
R29 (3,198 characters → 1,485 English words), the matched-flattened arm, the oddity arm, the 84-row
site table, the alignment, the 342-check pre-run verifier and the 435-check post-run verifier, all $0.
UTC day 2026-08-15 running total: $1.016607550 of $5.00 across two sessions — S188 $0.321811, S189 $0.694797. 79.7% of the day's cap unspent.
S190 — 2026-08-15 UTC
Baseline key read 112.783646882, which is $0.327797298 above the 112.455849584 that S189
recorded at its close, with a second read taken immediately and no movement (note (bns)).
The movement therefore happened between the two sessions and not while this one was reading. It is
not attributed to this session, and per CLAUDE.md the ledger below is built from per-request
costs, which are primary; the key delta is the sanity check. The key's all-time figure includes
non-project spend, which is the standing reason the delta is not the ledger.
| stage | seats | bodies | cost | note |
|---|---|---|---|---|
| pre-run adversarial critic | P3 |
1 | $0.042382 | NEEDS REDESIGN, 3 BLOCKING of 12 findings, 13,065 chars, finish_reason: stop, no continuation. Note (bny)'s explicit reasoning: {"max_tokens": 2000} inside max_tokens: 9000 worked first time |
| blind sound census | P1 P2 P3 |
165 | $0.440712 | 55 passages × 3 seats; 0 not stop, 1 re-dispatched, 0 invalid |
full-extract check |
P1 P2 P3 |
6 | $0.016619 | the two cells where minimal and full differ, so both codes could be printed as the design required |
| total | 172 | $0.499713 | against a declared ceiling of $0.81 — 62% |
The judging runner was killed once by a wall clock at 93 of 165 bodies and resumed; note (bnx)'s append-and-resume held again — every returned body survived and none was re-bought. Note (bnz) predicts a residual of about one call's cost from the request in flight at the kill.
Reconciliation. Key 112.783646882 → 113.276593632, delta $0.492946750 against a ledgered $0.499713400 — the key reading $0.006766650 LOW, i.e. behind rather than ahead. Note (bns)'s second read, with no calls between, returned 113.284516932, a further $0.007923300 — settlement lag, in the direction the note predicts and large enough to cover the gap. Against that read the residual is $0.001157 and needs no hunt.
Lead translation is free and is not ledgered (charter §3, A4): span C of «ألف ليلة وليلة», 1,036 Arabic words → 1,965 English, the frame tale completed, its collation against the second witness, the Būlāq/Calcutta II search, and all four verification scripts, all $0.
UTC day 2026-08-15 running total: $1.516321 of $5.00 across three sessions — S188 $0.321811, S189 $0.694797, S190 $0.499713. 69.7% of the day's cap unspent.
S191 — 2026-08-15 UTC
Baseline key read 113.538439332, which is $0.253922400 above the 113.284516932 that S190
recorded at its close. As at S190, the movement happened between sessions, with no session of
this project in between; it is not attributed to this session, and the ledger below is built
from per-request costs, which CLAUDE.md makes primary. This is the second consecutive session to
open on a between-session movement of this size (S190: $0.327797298), which is now a pattern rather
than an incident and is recorded as such — the key's all-time figure includes non-project spend.
| stage | seats | bodies | cost | note |
|---|---|---|---|---|
| pre-run adversarial critic | P1 |
1 | $0.046800 | NEEDS REDESIGN, 6 BLOCKING of 13 findings, 15,461 chars, finish_reason: stop, no continuation. Note (bny)'s explicit reasoning: {"max_tokens": 2000} inside max_tokens: 9000 worked first time, third session running |
stage 1 — Q1 instrument controls |
P1 P2 P3 |
12 | $0.027907 | four third-party controls, all unanimous on the expected code; the gate on everything after it |
stage 2 — Q1 manipulation check |
P1 P2 P3 |
156 | ~$0.30 | 52 lead cells × 3 seats |
stage 3 — Q2 source-inference |
P1 P2 P3 |
156 | ~$0.57 | 52 lead cells × 3 seats |
stage 4 — Q2 published panel |
P1 P2 P3 |
48 | $0.071739 | Lane 1839 and Burton 1885 on 16 stored cells, bought last because secondary |
| one-body repair | P1 |
1 | $0.001246 | P2.base/Q2 died twice at finish_reason: length with empty content — note (bnk); third attempt with an explicit reasoning cap, note (bnl) |
| total | 373 | $1.022045 | against a declared ceiling of $1.79 — 57% |
Reconciliation, and it names a defect rather than shrugging at one. Baseline 113.538439332 → 114.586249032, a key delta of $1.047809700 against the ledgered $1.022045450: the key reads $0.025764250 HIGH. Note (bns)'s second read, taken immediately with no calls between, returned exactly the same figure, so this is not settlement lag.
It is the runner's re-dispatch ladder throwing away the cost of the attempt it replaced.
run.py re-dispatches a body whose reply is unusable and then records only the second response's
usage; the first response was billed and is not in the ledger. Ten bodies were re-dispatched,
nine on P1 and one on P2, and every one of them failed by truncation, which means the
discarded attempt burned its whole max_tokens: 9 × (400 out at $6/M + ~230 in at $1/M) + 1 ×
(1,400 out at $3.75/M + input) ≈ $0.029, against a measured residual of $0.0258. The
arithmetic closes. Method note (boe).
Lead translation is free and is not ledgered (charter §3, A4): «باب السائح والصائغ» rendered
whole (998 Arabic words → 1,862 English), its collation against the 1937 Amīriyya witness, the
R31 and R32 arms, the contamination check, cells.py, analyse.py and the 1,286-check
verifier, all $0.
UTC day 2026-08-15 running total: $2.538366 of $5.00 across four sessions — S188 $0.321811, S189 $0.694797, S190 $0.499713, S191 $1.022045. 49.2% of the day's cap unspent.
S192 — 2026-08-15 UTC — ARM-footing-price step 2, E-20260815c-footing-direction
Baseline key read 114.892335232, which is $0.306086 above S191's closing read with no session of this project in between — the third consecutive session to open on such a movement, and this session found out why (below).
| stage | seats | bodies | cost | note |
|---|---|---|---|---|
| pre-run critic, attempt 1 | P4 |
1 | billed, unrecoverable | killed by a 2-minute harness wall clock while the request was in flight — note (bnz). Re-run in the background |
| pre-run critic, attempt 2 | P4 |
1 | $0.050216 | NEEDS-AMENDMENT, 3 BLOCKING of 7 findings, 14,801 in / 2,968 out, finish_reason: stop. Note (bny)'s explicit reasoning: {"max_tokens": 2000} inside max_tokens: 9000 worked first time, fourth session running |
| gate — 2 manufactured controls + 3 duplicated cells and their twins | P1 P2 P3 |
25 rows | $0.073056 | the gate FAILED on C2 and every registered primary was withheld |
diagnostic CTL2 — the same scene with the events flattened |
P1 P2 P3 |
6 | $0.018187 | registered before dispatch; holds unanimously at the ends of the scale |
| main | P1 P2 P3 |
115 rows | $0.400568 | 108 kept + 7 discarded re-dispatch attempts |
| of which discarded attempts, ledgered rather than thrown away | 8 | $0.025236 | note (boe) fixed in this runner, and the fix is visible in the total | |
| total | 138 kept | $0.542027300 | against a declared ceiling of $0.90 — 60% |
The reconciliation is retired rather than performed, and that is the session's method finding. Closing key read 115.900417122, a delta of $1.008081890 against the ledgered $0.542027300. This is neither settlement lag nor the runner leak: two reads twenty seconds apart with no request in between differ by $0.005285, and a 60.9-second idle window moves the figure $0.013748 — roughly $0.81 an hour while this session dispatched nothing. Something outside these sessions bills the same key continuously. Per-request billed costs are the ledger and are exact; the key delta is provenance and is no longer a cross-check. Method note (bof).
Lead translation is free and is not ledgered (charter §3, A4): T-genji-yomogiu-R33-v1 — the
censused span re-rendered whole under the new regime R33, twenty paragraph gradings, one declined
mark — plus extract.py, predictors.py, build.py, analyse.py and the 219-check verifier, all
$0.
UTC day 2026-08-15 running total: $3.080393 of $5.00 across five sessions — S188 $0.321811, S189 $0.694797, S190 $0.499713, S191 $1.022045, S192 $0.542027. 38.4% of the day's cap unspent.
S193 — 2026-08-15 UTC — ARM-supplied-footing step 1 (T4), E-20260815d-supplied-footing
Pre-flight, rebuilt after the pre-run critic's amendments, from the max_tokens cap each request
permits and not from expected length (note (abc)): worst case $0.632 including the critic already
spent, against a declared ceiling of $1.00 and headroom of $1.919607.
| stage | seats | bodies | cost | note |
|---|---|---|---|---|
pre-run critic, NEEDS-REDESIGN, 10 findings, 5 BLOCKING |
P4 |
1 | $0.098946 | 9 accepted, 1 overruled in writing; the acceptances changed which prediction is primary |
GOLD — gate G1b, the coder's labels ratified off the lead |
P3 P4 |
20 | $0.063886 | 1 disagreement of 10, against a withholding threshold of 3 |
FULL / ISO / ISOL / RATE |
P1 P2 P3 |
100 | $0.227871 | includes 19 bodies re-bought after a max_tokens truncation, both attempts ledgered (note (boe)) |
| total | 120 kept, 141 rows | $0.390703 | 39% of the declared ceiling |
No reconciliation against the key-usage delta is performed or reported, per note (bof)
(S192): the key is billed continuously by something outside these sessions, at roughly $0.81 an
hour while idle, so the delta is not an instrument. Per-request billed cost from
"usage": {"include": true} is the ledger and is exact.
A defect found inside the run and repaired, reported rather than smoothed. The reasoning budget
was set at or above the content cap for the short stages, and 21 bodies came back truncated
(finish_reason: "length"). They were re-bought selected mechanically on finish_reason, never
on content; the verifier now asserts no kept body was truncated; no verdict changed. Note
(bph).
Lead translation is free and is not ledgered (charter §3, A4): T-silver-blaze-ja-R06-v1 — the
Mapleton scene rendered into Japanese under R06, twelve logged decisions and three registered
predictions, frozen at fb58476b before the design existed — plus code_politeness.py,
extract.py, analyse.py and the 848-check verifier, all $0.
UTC day 2026-08-15 running total: $3.471096 of $5.00 across six sessions — S188 $0.321811, S189 $0.694797, S190 $0.499713, S191 $1.022045, S192 $0.542027, S193 $0.390703. 30.6% of the day's cap unspent.
S194 — 2026-08-15 UTC — ARM-fluent-carriage step 2 (T2), E-20260815e-fluent-carriage-2
Pre-flight declared ceiling $1.35; billed $1.215929050, 90% of ceiling. Every figure
API-reported with "usage": {"include": true}. No key-usage delta was read and none is reported
as a cross-check — note (bof).
| stage | calls | billed |
|---|---|---|
parity gate (PAR + PLANT), three builds, 2 re-dispatches |
31 | $0.236037 |
pre-run adversarial critic (P4, one call, stop, 18,550 chars) |
1 | $0.180201 |
| pre-run audits — blind carriage audit (4), language signature (1), recognition (3) | 8 | $0.135946 |
| judging, 7 conditions × 7 segments × 3 seats | 147 | $0.663745 |
| 187 | $1.215929050 |
Judging cost by seat, and it is the session's cost lesson. P1 openai/gpt-5.6-terra
$0.085254, P2 google/gemini-3.6-flash $0.105163, P3 x-ai/grok-4.5 $0.473328 —
a 10× spread inside one run, on identical prompts, which appeared for the first 23 cells
(P3 at $0.0222/cell) and receded for the remaining 124 (P3 at ~$0.004/cell). This is
config/models.md's 2026-07-25 routing caution firing intra-run: a pre-flight built from the price
table cannot see it. A hard spend guard that reads billed cost off disk before each call was
installed mid-run, and is what made the two conditions cut on money affordable to restore.
UTC day 2026-08-15 running total: $4.687025050 of $5.00 across seven sessions — S188 $0.321811, S189 $0.694797, S190 $0.499713, S191 $1.022045, S192 $0.542027, S193 $0.390703, S194 $1.215929. $0.312975 unspent; the next session on this UTC day should plan a $0 unit.
S195 — 2026-08-15 UTC — ARM-alf-layla step 4 (T1), E-20260815f-night-formula
Pre-flight built from the max_tokens cap the request permits, not from expected length (note
(abc)): worst case $0.062 for one critic call and $0.124 if the truncation continuation fired,
against a declared ceiling of $0.13 and the day's remaining headroom of $0.312975. The session
planned a $0 unit and spent only the gate.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS REDESIGN, 15 findings, 4 BLOCKING |
P1 openai/gpt-5.6-terra, provider OpenAI |
1 | $0.030287 | finish_reason: stop, 3,144 in / 4,393 out; all 15 accepted; findings 1, 3 and 4 were caught by checking the design against the stored Arabic, which is why the critic was shown the source and no English |
| the census itself | — | — | $0.00 | two Project Gutenberg files and a string count; lead work is free and is never ledgered (charter §3, A4) |
| total | 1 | $0.030287 | 23% of the declared ceiling |
No reconciliation against the key-usage delta is performed or reported, per note (bof).
Per-request billed cost from "usage": {"include": true} is the ledger and is exact.
The cost lesson of this session is the cheap one. A single frontier critic call on an
8,000-character design cost three cents and returned four BLOCKING findings, two of which were
errors of fact in the design's own baseline that the lead had computed and believed. P1 remains
the cheapest frontier seat in practice, as its config/models.md row says.
UTC day 2026-08-15 running total: $4.717312 of $5.00 across eight sessions — S188 $0.321811, S189 $0.694797, S190 $0.499713, S191 $1.022045, S192 $0.542027, S193 $0.390703, S194 $1.215929, S195 $0.030287. $0.282688 unspent; a further session on this UTC day still has essentially no headroom and should plan a $0 unit.
S196 — 2026-08-16 UTC — ARM-answering-figure step 2 (T5), E-20260816b-invented-figure
A fresh UTC day: the ledger opens at $0.00 of $5.00. Pre-flight built from the max_tokens caps
the requests permit, not from expected length (note (abc)): worst case $2.028 for 402 census
bodies plus $0.184 for the re-dispatch ladder plus $0.145 for two critic calls, against a
declared ceiling of $2.40 with a stop-loss in the runner.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS REDESIGN, 22 findings, 13 BLOCKING |
P1 openai/gpt-5.6-terra |
1 | $0.065025 | all 22 accepted; the arms were rebuilt from scratch on BLOCKING 1 |
second, focused critic pass on the rebuilt arms, RUN WITH AMENDMENTS, 14 findings |
P1 |
1 | $0.031697 | six loci rewritten, six refused on a mechanical ground, one accepted as irreducible |
the census — Q1 and Q2 on 52 PLAIN cells, Q2 on 26 SITE cells, 4 controls |
P1 / P2 / P3 |
402 | $1.091013 | providers: OpenAI 134, xAI 134, Google 76, Google AI Studio 58 |
| the translation limb and the Arabic collation | — | — | $0.00 | lead work is free and is never ledgered (charter §3, A4) |
| total | 404 | $1.187735 | 49% of the declared $2.40 ceiling |
Realised cost per census body $0.002714, against S191's $0.002624 on the same three seats — the estimate that mattered was the per-body one and it was right to three decimal places.
Note (boe) fired again and is reported rather than hidden. Eleven bodies were re-dispatched by
the runner's usability test and two were repaired on a third attempt; the ladder overwrites the
discarded attempt's usage, so the $1.091013 is the cost of the bodies that were kept, not of
every request billed. The thirteen discarded attempts were mostly truncations at cap and are worth
roughly $0.03 more. The runner still carries the defect the note describes; fixing it is
instrument work and is not this session's unit.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
Per-request billed cost from "usage": {"include": true} is the ledger.
UTC day 2026-08-16 running total after S196: $1.187735 of $5.00 across one session.
S197 — 2026-08-16 UTC — ARM-invented-ornament step 1 (T3), E-20260816c-checked-ornament
The day opens for this session at $1.187735 of $5.00, $3.812265 unspent. Pre-flight built from
the max_tokens caps the requests permit, not from expected length (note (abc)): worst case
$2.40 for 579 bodies, against a declared ceiling of $2.40 with a stop-loss in the runner. Note
(bpq) is new and is about exactly that: a stop-loss equal to the worst case is not a margin.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS REDESIGN, 28 findings, 8 BLOCKING |
P1 openai/gpt-5.6-terra |
2 | $0.121977 | twelve implemented, two overruled in writing; the design was rebuilt as v2 and v1 was never dispatched |
aborted run v2a — the four Q1 instrument controls, dispatched with no marked span |
P1 / P2 / P3 |
12 | $0.042105 | the abort gate fired correctly; P1: "N; no marked span is present". Note (bpp). Rows kept in run-v2a-aborted-controls.jsonl |
| the run — controls, the Arabic-only gate arm, and the two source-blind arms | P1 / P2 / P3 |
351 | $1.209261 | providers: OpenAI 117, xAI 117, Google 96, Google AI Studio 21 |
| not bought: the four source-visible cells (252 bodies) | — | 0 | $0.00 | MC1 failed, design §8 withholds P2/P3/P4, and those cells existed only to compute them — RS-20260816c §10.2 |
| the translation limb, the Arabic collation and the contamination check | — | — | $0.00 | lead work is free and is never ledgered (charter §3, A4) |
| total | 363 | $1.373343 | 50% of the declared $2.40 ceiling |
Note (boe) is FIXED in this session's runner and the fix changes the ledger's meaning: a re-dispatched body now records the cost of every attempt, not only the attempt kept, so the figure above is the cost billed and not the cost retained. 35 of 351 of the bodies were re-dispatched. Previous sessions' figures are the retained cost and are not comparable to this one at the cent level.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
Per-request billed cost from "usage": {"include": true} is the ledger.
UTC day 2026-08-16 running total: $2.561078 of $5.00 across two sessions. $2.438922 unspent.
S198 — 2026-08-16 UTC — ARM-supplied-footing step 2 (T4), E-20260816d-lexical-channel
The day opens for this session at $2.561078 of $5.00, $2.438922 unspent. Pre-flight built from
the max_tokens caps the requests permit, not from expected length (note (abc)). v1's
stop-loss sat BELOW its own worst case and the pre-run critic caught it; v2 was re-costed until
worst case $1.97 < stop-loss $2.05 < declared ceiling $2.15 — note (bpq), now two-sided.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS-REDESIGN, 18 findings, 15 BLOCKING |
P1 |
1 | $0.045802 | thirteen implemented, four overruled; v1 was never dispatched |
GATE — 8 contestable addressee labels |
P4 QR |
20 | bought before P4 was dropped; stands |
|
RATE — 70 English lines, two seats that never translate |
P1 QR |
142 kept | ||
P4 bodies bought and DISCARDED |
P4 |
36 | $0.174286 | the seat was dropped mid-run on a token-budget defect, note (bps); ledgered in full, note (boe) |
ISO — 70 lines, three hands, line alone |
P1 P2 P3 |
212 kept | complete census | |
ISOL + SWAP |
P1 P2 P3 |
77 kept | narrowed twice after CTL1 failed — RS-20260816d §7.2 |
|
not bought: the scene-A ISOL cells and 16 of scene B's |
— | 0 | $0.00 | they existed only to compute Q1/Q2/Q3, which the gate withheld |
| the translation limb, both collations, 森田's coding and every contamination measurement | — | — | $0.00 | lead work is free and is never ledgered (charter §3, A4) |
| total | 508 bought, 470 kept | $1.004306 | 49% of the declared $2.15 ceiling |
Note (bof) honoured: no key-usage reconciliation was performed or reported. Per-request billed cost is the ledger and is exact.
UTC day 2026-08-16 running total after S198: $3.565384 of $5.00 across three sessions. $1.434616 unspent.
S199 — 2026-08-16 UTC — ARM-mimetic-carriage step 1 (T2), E-20260816e-mimetic-carriage
The day opens for this session at $3.565384 of $5.00, $1.434616 unspent. Pre-flight built from
the max_tokens caps the requests permit, not from expected length (note (abc)). The first
cut of the estimate came out at a worst case of $1.27 against its own stop-loss of $0.95 — note
(bpq)'s exact failure, caught in the design rather than by the critic — and the item cap was cut
from 1,200 to 450 and re-costed until worst case $0.83 < stop-loss $0.92 < declared ceiling
$1.05.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS-REDESIGN, 15 findings, 15 BLOCKING |
P1 |
1 | $0.048823 | none overruled; v1 was never dispatched |
probe — reasoning budget per seat, note (bps) |
QR GL |
2 | $0.001253 | probed on the longest classify item; the note says the longest census item, and the shortfall cost GL (limits §3) |
G4 parity — 15 pairs + 4 planted content errors |
P2 |
19 | $0.019358 | the gate that failed and stopped the run |
G2 census — the Japanese chapter, no English shown |
QR GL P3 |
5 | $0.085941 | of which $0.017566 was 2 GL bodies bought and discarded on finish_reason: length; P3 bought as the replacement |
G3 class — 15 sites, sound vs manner |
QR GL |
30 | $0.020180 | |
G4 replication — the same 19 pairs, second rater |
QR |
19 | $0.025066 | unregistered; bought because a headline resting on one seat is not a headline |
| not bought: the decoy stage (15 cells) and the main stage (39 cells) | — | 0 | $0.00 | G4 failed, design §6 withholds P1/P2/P3, and those 54 cells existed only to compute them — RS-20260816e §8 |
| the translation limb, the contamination check, the census and the Morri coding | — | — | $0.00 | lead work is free and is never ledgered (charter §3, A4) |
| total | 76 bought, 68 kept | $0.200621 | 19% of the declared $1.05 ceiling |
No reconciliation against the key-usage delta is performed or reported, per note (bof).
Per-request billed cost from "usage": {"include": true} is the ledger and is exact. Every attempt
is billed to the row, including discarded ones, per note (boe).
UTC day 2026-08-16 running total after S199: $3.766005 of $5.00 across four sessions. $1.233995 unspent.
S200 — 2026-08-16 UTC — ARM-alf-layla step 5 (T1), E-20260816f-night-seam
The day opens for this session at $3.766005 of $5.00, $1.233995 unspent. Pre-flight built from
the max_tokens caps the requests permit (note abc). Note (bpq) honoured in the design:
worst case $0.691 < stop-loss $0.75 < declared ceiling $0.90, and the stop-loss is
enforced inside run.py rather than by intention.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS REDESIGN, 14 findings, 7 BLOCKING |
P1 |
1 | $0.043261 | 11 taken, 3 refused on the record; v1 never dispatched |
| stage 0 — memory assay, no text (the critic's BLOCKING 8) | P1 P2 P3 QR |
4 | $0.017691 | P1 returned length with an empty body and is unmeasured |
| stage 1 — probe on the longest arm, note (bps) | P1 P2 P3 QR |
4 | $0.051121 | all four stop, all four parsed; kept as BURTON replicate 1 |
| stage 2 — the remaining bodies | P1 P2 P3 QR |
28 | $0.257558 | 2 dead (P2 ×2), reported dead, never re-rolled |
| session total | 37 | $0.369632 | 41% of the declared ceiling |
Two panel facts learned before any scored body was bought, and both are cheap to repeat.
QR (qwen/qwen3.7-max) returns HTTP 400 — "max_completion_tokens must be greater than
thinking_budget" — whenever max_tokens ≤ the reasoning cap; and P1 returns finish_reason:
length with an empty body when the reasoning cap is above max_tokens. Caps of 1,500/900 and
1,400/700 work for all four seats.
UTC day 2026-08-16 running total after S200: $4.135637 of $5.00 across five sessions. $0.864363 unspent.
S201 — 2026-08-16 UTC — ARM-device-function step 1 (T5), E-20260816g-device-function
The day opens for this session at $4.135637 of $5.00, $0.864363 unspent. Pre-flight built from
the max_tokens caps the requests permit (note abc). Note (bpq) honoured: worst case
$0.361 < stop-loss $0.39 < declared ceiling $0.42, and the stop-loss is enforced inside
run.py. Note (bpv) honoured: max_tokens 2000 > reasoning cap 900 everywhere, and no seat
returned an HTTP 400.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS REDESIGN, 9 findings, 4 BLOCKING |
P1 |
1 | $0.021874 | seven taken, two refused in part; v1 never dispatched |
| stage 0 — the plain alternatives checked before dispatch (the critic's BLOCKING 2) | P2 P3 |
2 | $0.017952 | English gate cleared; event preservation 12 and 11 of 12, reported not gated |
| stage 1 — probe, replicate 1, both arms, note (bps) | P1 P2 P3 QR |
8 | $0.084583 | P1 returned finish_reason: length on FLAT and was dropped by the frozen rule, taking its one usable body with it |
| stage 2 — replicates 2 and 3, three live seats | P2 P3 QR |
12 | $0.128920 | all 12 stop, all 12 parsed |
| session total | 23 bought, 21 usable | $0.253329 | 60% of the declared $0.42 ceiling |
One dead body, P1 on FLAT replicate 1, reported dead, counted in nothing, never
re-rolled. Providers seen: OpenAI, Google, Google AI Studio, xAI, Alibaba.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-16 running total after S201: $4.388966 of $5.00 across six sessions. $0.611034 unspent.
S202 — 2026-08-16 UTC — ARM-invented-ornament step 2 (T3), E-20260816h-target-set
The day opens for this session at $4.388966 of $5.00, $0.611034 unspent. Pre-flight built from
the max_tokens caps the requests permit (note abc): input ≤ 750, output capped at 500, three
seats, 31 items. Note (bpq) honoured: worst case $0.365 < stop-loss $0.39 < declared
ceiling $0.42, and the stop-loss is enforced inside run.py. Note (bpv) honoured:
max_tokens 500 > reasoning cap 150 everywhere, and no seat returned an HTTP 400.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS REDESIGN, 11 findings, 4 BLOCKING¹ |
P1 |
1 | $0.033371 | nine taken in six amendments, two refused in writing; v1 never dispatched |
| stage 1 — the three third-party controls, probe and gate in one, note (bps) | P1 P2 P3 |
9 | $0.020099 | all nine parsed; the gate passed on soundwork, owed and the label |
| stage 2 — the 28 study loci | P1 P2 P3 |
84 | $0.299365 | 4 dead (finish_reason: length), reported dead, never re-rolled |
| session total | 93 bought, 89 usable | $0.352836 | 84% of the declared $0.42 ceiling |
¹ Three BLOCKING in the critic's own count; the fourth-order finding BLOCKING 1 was refused in
part. Providers seen: OpenAI, Google AI Studio, xAI.
Four dead bodies — P1 on F4 and F19, P2 on F3 and F22 — counted in nothing, reported
on the result page at §9, never re-rolled. Method note (bph) fires again despite being honoured:
the reasoning cap was 150 against a 500-token content cap and four bodies still spent the content
cap on reasoning.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-16 running total after S202: $4.741802 of $5.00 across seven sessions. $0.258198 unspent.
S204 — 2026-08-17 UTC — ARM-mimetic-carriage step 2 (T2), E-20260817e-mimetic-reading
The day opens at $0.00 of $5.00; this is its first session. Pre-flight built from the
max_tokens cap each request permits (note abc), including the one mechanical re-dispatch at
double cap. Note (bpq) honoured two-sided: worst case $0.74 < stop-loss $0.82 < declared
ceiling $0.95, and the stop-loss and a hard ceiling check are both enforced inside run.py.
Note (bpv) honoured: max_tokens 600 against a reasoning cap of 120 on every item call, 3,000
against 400 on the two census calls.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic v1, NEEDS-REDESIGN, 18 findings, 12 BLOCKING |
P1 |
1 | $0.049393 | truncated by its own cap at 6,000 completion tokens; findings 19+ never seen; v1 never dispatched |
pre-run adversarial critic v2, NEEDS-REDESIGN, 16 findings, 12 BLOCKING |
P1 |
1 | $0.065116 | cap 14,000, finish_reason: stop; its BLOCKING 2 invented the control that ran; v2 never dispatched |
census — G0, the chapter-3 mimetic census |
P2 P3 |
2 | $0.044062 | both parsed at cap 3,000; 12 words each |
classify — G0b, the four chapter-3 sites |
P2 P3 |
8 | $0.015414 | 8 of 8 parsed, both seats agreeing at 4 of 4 |
reading — the primary, the deletion condition and the identity diagnostic |
P2 P3 QR |
103 | $0.479929 | 81 jobs; 4 dead, all QR; P2 truncated at the 600 cap on 4 items and all four were recovered at 1,200 |
build — the blind second hand at the nine unbuildable sites |
P2 P3 |
18 | $0.052674 | 18 of 18 parsed |
| session total | 133 bought, 105 usable | $0.706588 | 74% of the declared $0.95 ceiling |
Four dead bodies, all QR, at N01, X-M03, X-M04, X-N06 — one mechanical re-dispatch
each, then written to run.jsonl as dead and never re-bought, which is the v2 critic's BLOCKING 12
implemented. They are reported on the result page at §5 and §9 and cost $0.0 in the ledger only
because the dead marker row carries no usage; their two paid attempts are inside the reading line.
16% of the session went on two critic passes that killed two designs, the second on arithmetic the lead had got wrong: a registered 0.20 contrast threshold smaller than the 0.219 standard error of the difference it was testing.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-17 running total after S204: $0.706588 of $5.00 across one session. $4.293412 unspent.
UTC day 2026-08-20 — S205
The day opens at $0.00 of $5.00; this is its first session. Pre-flight built from the max_tokens
cap each request permits (note abc). The first pre-flight was wrong and the stop-loss caught
it — see note (bqk) and the revision recorded in run.py.
| stage | seat | bodies | cost | note |
|---|---|---|---|---|
pre-run adversarial critic v1, NEEDS-REDESIGN, 6 BLOCKING |
P1 |
1 | $0.0749395 | cap 14,000, finish_reason: stop; v1 never dispatched |
pre-run adversarial critic v2, NEEDS-REDESIGN, 5 BLOCKING |
P1 |
1 | $0.0829490 | its BLOCKING 4 removed the design's inference and left a description; v2 never dispatched |
census — four renderings × three seats, E-20260820-rhyme-bearer v3 |
P1 P2 P3 |
16 bought, 12 usable | $0.263144 | 4 discarded, 0 dead — all four P1 truncating at the 1,500 cap, all four recovered at 3,000 |
| session total | 18 bought, 14 usable | $0.4210325 | 37% of it killed two designs |
The stop-loss fired mid-design at $0.243000, with Lane holding one seat, because the estimate
priced re-dispatches for a fraction of the calls at an average seat rate and the dearest seat
truncated on every hand. The threshold was revised with the arithmetic written into run.py and
the two remaining registered calls — whose content was frozen and whose outcome could not be steered
— were completed. Note (bqk) records the rule this changed.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-20 running total after S205: $0.4210325 of $5.00 across one session. $4.5789675 unspent.
S206 — the same UTC day, second session
The day opened for this session at $0.4210325 of $5.00, leaving $4.5789675. Pre-flight built
from the max_tokens cap each request permits (note abc) and — per note (bqk), which fired
yesterday — assuming every call needs the doubled-cap re-dispatch, priced at the dearest
seat's rate. Worst case $1.144 < stop-loss $1.35 < declared ceiling $1.50 (note (bpq): the
stop-loss sits above the worst case). The stop-loss was never approached: actual $0.469376, 41%
of the worst case and 31% of the ceiling.
| stage | seat | calls | cost | note |
|---|---|---|---|---|
pre-run adversarial critic pass 1, NEEDS REDESIGN, 3 BLOCKING |
P1 |
1 | $0.063292 | v1 never dispatched |
pre-run adversarial critic pass 2, NEEDS REDESIGN, 3 BLOCKING |
P1 |
1 | $0.067593 | none a restatement; it broke v2's own noise band. v2 never dispatched |
pre-run adversarial critic pass 3, NEEDS REDESIGN, 2 BLOCKING |
P1 |
1 | $0.085206 | both new, both fixed free under a stopping rule fixed before the pass was bought. v3 never dispatched |
| A — the independent hand writes fourteen replacements | P1 |
1 | $0.010494 | finish_reason: stop at the 3,000 cap; no re-dispatch |
| A″ — what the hand says it removed (process documentation) | P1 |
1 | $0.009198 | gates nothing |
| B — gates and measurements on the replacements | P2 P3 |
3 | $0.0595704 | P2 truncated at 3,000 and was recovered at 6,000, as (bqk) requires the estimate to assume |
| C — the reading run, 3 replicates × 3 seats × 2 arms | P2 P3 QR |
18 | $0.174023 | 18 bought, 18 usable, 0 dead, every finish_reason: stop |
| session total | 26 calls | $0.469376 | 46% of it on three critiques, and each removed a claim the design was not entitled to make |
Note (bqk) worked as intended on its first outing. Its remedy — price the worst case as though
every call needed the doubled-cap re-dispatch, at the dearest seat's rate — produced a worst case of
$1.144 against an actual of $0.469376. Exactly one call did need the re-dispatch (P2 in stage B),
and the guard never fired.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-20 running total after S206: $0.8904085 of $5.00 across two sessions. $4.1095915 unspent.
S207 — the same UTC day, third session
The day opened for this session at $0.8904085 of $5.00, leaving $4.1095915. Pre-flight built per note (bqk) — every call priced as though it needed the doubled-cap re-dispatch, at the dearest seat's rate. Worst case ≈ $1.05 < stop-loss $1.35 < declared ceiling $1.50 (note (bpq)). The stop-loss was never approached: actual $0.427041, 41% of the worst case and 28% of the ceiling.
| stage | seat | calls | cost | note |
|---|---|---|---|---|
| A — the independent hand writes nine plain-Japanese substitutes | P1 |
1 | $0.034866 | finish_reason: stop at the 3,000 cap; no re-dispatch |
| B — G3 comparison of the two substitute sets | none | 0 | $0.000000 | free code; wrote stage_b.json |
pre-run adversarial critic pass 1, NEEDS REDESIGN, 5 BLOCKING, 4 MAJOR |
P1 |
1 | $0.075640 | all nine answered in critic-response.md; primary restricted from 9 to 7 sites, criterion rewritten as directional, per-item continuity added |
| C — reading run, 9 sites × 2 English × 2 source × 3 seats | P2 P3 QR |
108 | $0.316535 | 108 bought, 108 usable, 0 dead, every finish_reason: stop |
| session total | 110 calls | $0.427041 | 18% of it on the one critic pass; the stopping rule fixed at S206 held and no second round was bought |
Note (bqk) held again. Its worst-case pricing ($1.05) came in at 41% actual with no re-dispatch fired. Note (bpq)'s stop-loss above the worst case stayed dormant, as it is meant to.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-20 running total after S207: $1.3174495 of $5.00 across three sessions. $3.6825505 unspent.
UTC day 2026-08-21 — S208
A new UTC day: the ledger opens at $0.00 of $5.00. Pre-flight built per note (bqk) — every
call priced as though it needed the doubled-cap re-dispatch, at the dearest seat's rate — and per
note (abc), from max_tokens rather than from an assumed output length.
Declared ceiling $2.50; stop-loss in the runner $2.20; actual $2.097425 — 84% of the ceiling and the closest this project has run to a stop-loss without touching it.
| stage | seat | calls | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS REDESIGN, 3 BLOCKING / 7 MAJOR / 1 MINOR |
P1 |
1 | $0.109895 | all eleven accepted, none overruled; no second round bought under note (bqp) |
| stage A — blind census, 54 loci × 3 seats, whole chapter in every prompt | P1 P2 P3 |
162 | $1.98753035 | 161 parsed, 1 dead after re-dispatch; 36 first attempts hit length and were re-dispatched at the doubled cap |
| session total | 163 calls | $2.097425 | 5% of it on the critic pass, and the critic pass rebuilt the design |
Why this session cost four times the recent average, stated so the number is not read as drift. The pre-run critic's first BLOCKING finding abolished the design's English windows, so every one of the 162 bodies carries a whole 2,300–2,600-word chapter in its prompt instead of a 120-word window. That is roughly a 4× increase in input tokens per call, taken deliberately: the windows could not distinguish the translator omitted this from the lead's window does not reach it, and the census's ABSENT verdicts were circular without the change.
Note (bqk) held, and only just. The declared worst case was $2.20 and the realised total was $2.097425 — 95% of it. The re-dispatch rate (36 of 162 first attempts) is the highest this project has recorded and is what consumed the margin; a future design with whole-document prompts should raise the output cap rather than rely on the doubled-cap retry.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-21 running total after S208: $2.097425 of $5.00 across one session. $2.902575 unspent.
UTC day 2026-08-21 — S209 — ARM-matched-shape-heard step 1 (T2), E-20260821b-matched-heard
Declared ceiling $1.20; stop-loss in each runner $1.00; actual $1.00069965 — the stop-loss fired, three of ninety main cells unbought, and the ceiling was never approached.
| stage | seat | calls | cost | note |
|---|---|---|---|---|
pre-run adversarial critic, NEEDS-REDESIGN, 8 BLOCKING / 3 MAJOR / 1 MINOR |
P1 |
1 | $0.096713 | eleven accepted, one overruled on a stated ground; no second round bought under note (bqp) |
| eligibility screen (22 loci × 2 versions) + parity screen (22 real + 4 planted) | P3, P2 |
70 | $0.148395 | 0 dead; this is the stage that withheld the primary |
| main run — 15 passages × 2 versions × 3 seats | P1 P2 P3 |
87 of 90 | $0.755592 | 1 dead after re-dispatch; stop-loss at cell 87 |
| session total | 158 calls | $1.00069965 | 25% of it on the two stages that decided the result before the main run |
Why the stop-loss fired, stated so it is not read as drift. The pre-flight priced the worst case
from max_tokens alone and did not add reasoning.max_tokens, which is billed at the completion
rate — $0.008685 realised per main body against ≈$0.0039 assumed, a factor of 2.2. Note
(bqs). (bqk) held only because the stop-loss had been set below the declared ceiling, so the
arithmetic error cost three cells rather than a breach.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-21 running total after S209: $3.09812465 of $5.00 across two sessions. $1.90187535 unspent.
UTC day 2026-08-21 — S210 — ARM-alf-layla step 7 (T1), E-20260821c-echo-availability
Declared ceiling $1.40 for the whole experiment; runner ceiling $1.30, runner stop-loss $1.15; actual $0.84032408. Neither guard fired and the ceiling was never approached.
| stage | seat | calls | cost | note |
|---|---|---|---|---|
| pre-run adversarial critic, 4 BLOCKING / 4 MAJOR / 1 MINOR | P1 |
1 | $0.087891 | every finding taken; no second round bought, note (bqp) — every remedy narrowed the primary |
| stage A — availability, 39 members × 2 seats, four sweeps | P2, QR |
171 | $0.12270673 | 78 cells, 0 dead. The sweep count is the cost of note (bqu): a cap raise and two parser repairs, five cells recovered from stored bodies at $0 |
| stage B — echo detection, 87 items × 3 seats | P1 P2 P3 |
325 | $0.62972635 | 261 cells, 13 dead, 11 of them P1 truncating at the 120-token cap; 64 re-dispatch calls |
| session total | 496 calls | $0.84032408 |
Why stage B cost 1.7× its pre-flight, stated so it is not read as drift. The worst case was built correctly per notes (bqs) and (bqk) — $0.373 for one clean pass, $0.909 assuming every call re-dispatches — and the realised $0.630 sits between them, which is what a 64-call re-dispatch overhead looks like. The arithmetic was right and the guard was never tested.
One operational fact worth recording. A background dispatch in this container advances only while a foreground command is running; the run was completed by holding the session with long foreground waits. It is not a cost fact and it changes no estimate, but it doubles the wall-clock plan for any run of a few hundred serial calls.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-21 running total after S210: $3.93844873 of $5.00 across three sessions. $1.06155127 unspent.
UTC day 2026-08-22 — S211 — ARM-synonym-reach step 1 (T5), E-20260822-synonym-reach
Declared experiment ceiling $1.60; runner ceiling $1.40, stop-loss $1.20; actual $1.067599850. Neither guard fired.
| stage | seat | calls | cost | note |
|---|---|---|---|---|
| pre-run critic round 1, 3 BLOCKING / 2 MAJOR | P1 |
1 | $0.0523085 | all five accepted |
| pre-run critic round 2, on the amendments, 1 BLOCKING / 2 MAJOR | P1 |
1 | $0.0321075 | bought under note (bqp)'s second clause — round 1's answer to its own BLOCKING 3 was a component the critic had not seen, and round 2 killed the reading that component was being asked to carry |
| stage G — source-side gate, 36 loci × 3 seats | P1 P2 P3 |
108 | $0.289644 | 108 of 108 live; 24/24 rhymed, 11/12 control |
| stage A — the first ordinary word, 76 members × 2 seats | P2 P1 |
159 | $0.244977 | 152 cells, 7 re-dispatches |
| stage B — the enumeration, 76 members × 2 seats | P2 P1 |
153 | $0.379350 | 152 cells, 1 re-dispatch |
| stage R — cross-call repeat control | P2 P1 |
16 | $0.035400 | mean Jaccard 0.611 |
| stage N — recognition probe | P2 P1 |
20 | $0.098683 | 1 dead; 6 of 14 named the work |
| stage C — the meaning screen + 4 planted errors | P2 P1 |
16 | $0.019545 | 4 of 4 planted caught |
| session total | 472 bodies | $1.067599850 | 458 parsed, 14 dead (3.0%) |
One runner repair, made after stage G and before any stage-A cell was bought, with the arithmetic
written down per note (bqk). Measured over the 108 gate calls, P3 cost $0.00660/call at ~1,015
completion tokens against P1's $0.00127 and P2's $0.00079 — the provider honours neither
max_tokens 400 nor the reasoning cap of 120. Left in stages A/B/C/R it would have cost about
$2.38, above the experiment ceiling on its own. P3 was replaced by P1 in those stages; it
keeps the gate cells it had already returned, nothing bought was re-bought, and no stage-A/B/C
cell existed when the change was written, so no outcome could be steered by it. This is the second
time a panel seat's real cost has been discovered only by dispatching it (cf. the S022 routing
caution): a stage's first twenty calls are the only reliable price list there is.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-22 running total after S211: $1.067599850 of $5.00 across one session. $3.932400150 unspent.
S212 — 2026-08-22 — E-20260822b-echo-threshold (T3, ARM-echo-threshold step 1)
Declared experiment ceiling $1.90; runner ceiling $1.55, stop-loss $1.35; actual $1.195830900.
Neither guard fired. Ceiling rose from $1.00 to $1.90 during design, before any data call, because
the pre-run critic's two rounds turned a 69-item design into a 138-item crossover plus two added
stages; the arithmetic is in design.md §9 and the reasons in critic-response.md.
| stage | seats | calls | cost | note |
|---|---|---|---|---|
| pre-run critic round 1, 2 BLOCKING / 3 MAJOR | P1 |
1 | $0.052856 | all five accepted; the design was rebuilt as a crossover |
| pre-run critic round 2, on the amended design, 2 BLOCKING / 3 MAJOR | P1 |
1 | $0.078452 | bought under note (bqp)'s second clause, and it found the orthography confound |
| stage P — orthography/phonology probe, 12 items × 3 seats | P1 P2 QR |
36 | $0.031938 | dispatched FIRST, as the price check for QR |
| stage D — prompted detection, 138 items × 3 seats | P1 P2 QR |
427 | $0.439009 | 414 cells, 0 dead, 13 re-dispatches |
| stage U — unprompted, 57 items × 1 seat | P1 |
77 | $0.308608 | 49 of 57 parsed; 8 truncated |
| stage R — naturalness ranking, 25 loci × 2 seats | P1 P2 |
90 | $0.193712 | 50 cells, 0 dead |
| stage S — fluency screen, 142 items × 1 seat | P2 |
143 | $0.091257 | screen failed its own control; see the result page §5 |
| session total | 775 bodies | $1.195830900 |
QR's billed rate was unknown at dispatch and was measured before it could matter, per note
(bqk): stage P — twelve items to all three seats — was sent first precisely so that the seat's price
would be read from its own calls rather than from config/models.md's table. Measured over the first
twelve calls each: P1 $0.001488, P2 $0.000458, QR $0.000715 per call. All three sit well
inside the estimate, so no seat was dropped and no arithmetic had to be rewritten mid-run. This is
the counterpart of S211's P3 repair — the same check, run before the money rather than after the
first stage.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-22 running total after S212: $2.263430750 of $5.00 across two sessions. $2.736569250 unspent.
S213 — 2026-08-22 — E-20260822c-persian-hands (T4, ARM-persian-hands step 1)
Declared experiment ceiling $1.30; runner stop-loss $1.05; actual $0.362263500. Neither guard
fired. The ceiling rose in two steps during design and before any data call — $0.75 → $1.10
when the round-1 critic's finding 4 turned a self-certified claim into a bought check, and
$1.10 → $1.30 when the round-2 critic's finding 3 put the Persian source into every stage-L prompt
and added eight planted controls. The arithmetic is in design.md §6 and the reasons in
critic-response.md.
| stage | seats | calls | cost | note |
|---|---|---|---|---|
| pre-run critic round 1, 1 BLOCKING / 4 MAJOR | P1 |
1 | $0.055132 | all five accepted; FORM moved to the page scans, stage V deleted |
| pre-run critic round 2, 2 BLOCKING / 1 MAJOR | P1 |
1 | $0.073746 | it struck round 1's own remedy; the per-bearer alignment replaced it |
| pre-run critic round 3, 2 BLOCKING / 2 MAJOR / 1 MINOR | P1 |
1 | $0.080701 | multi-bearer aggregation frozen; P3 relabelled; the loop stopped here, with the reason written |
| stage L — 29 items × 2 seats, source-faithfulness against the Persian | P1 P2 |
59 | $0.152685 | 58 cells, 1 truncated body re-dispatched at a doubled cap; 8 of 8 plants caught by both seats |
| session total | 62 bodies | $0.362263500 |
Three-quarters of the money went on criticism, and it bought the study. The three critic rounds cost $0.209578 against $0.152685 of data, and four BLOCKING findings came out of them — one of which (grade the words that render the source's rhyme-bearers, not the ends of the English clauses) would have turned a locus with zero carriage into one with four of four. Note (bra).
The free half of the design cost nothing and was the larger half: the page scans of all four
books, the layout XML behind layout.py's independent FORM check, and every one of the 128
rhyme gradings are $0.
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-23 — S214 — ARM-run-depth step 1 (T2), E-20260823-run-depth
Declared experiment ceiling $1.60; runner ceiling $1.35, stop-loss $1.20; actual $0.842453050.
Neither guard fired. The ceiling rose from $1.00 to $1.20 to $1.50 to $1.60 during design, before
any data call, each time on a pre-run critic finding that replaced the primary's stimulus; the
arithmetic is in design.md §8 and the reasons in critic-response.md.
| stage | seats | calls | cost | note |
|---|---|---|---|---|
| pre-run critic round 1, 1 BLOCKING / 3 MAJOR | P1 |
1 | $0.054955 | all four accepted; the couplet arm was demoted and a matched arm added |
| pre-run critic round 2, 1 BLOCKING / 2 MAJOR | P1 |
1 | $0.079273 | the matched arm was itself confounded; replaced by a four-line minimal pair |
| pre-run critic round 3, 1 BLOCKING / 2 MAJOR | P1 |
1 | $0.084429 | the four-line pair was confounded too; replaced by the six-line stimulus that ran. Round 4 declined under note (bqp) |
| stage S — rhyme-scheme labelling, 58 bodies × 3 seats | P1 P2 P3 |
175 | $0.307676 | 174 cells, 1 dead, 1 re-dispatch |
| stage F — source-faithfulness, 24 items × 2 seats | P1 P2 |
62 | $0.316120 | 48 cells; 14 P1 bodies died on hidden reasoning consuming the whole 500-token cap and were re-bought at 1600/700 |
| session total | 240 bodies | $0.842453050 |
P3's billed rate was measured before it could matter, per note (bqk): twelve stage-S cells were
dispatched first and P3 came back at ~$0.004 per call, inside the estimate, so no seat was
dropped and no arithmetic had to be rewritten mid-run.
The 14 dead P1 bodies are the session's one avoidable cost, ~$0.09. max_tokens 500 with
reasoning.max_tokens 200 satisfies note (bpv)'s inequality and still let the model spend all 500 on
reasoning and return an empty body. The inequality is not the guard it was taken for — on a task
with any deliberation in it, the output cap must leave room after the reasoning actually spent, not
after the reasoning requested. Both attempts are ledgered, per note (boe).
No reconciliation against the key-usage delta is performed or reported, per note (bof).
UTC day 2026-08-23 running total after S214: $0.842453050 of $5.00 across one session. $4.157546950 unspent.
S215 — 2026-08-23 — E-20260823b-embedding-carriage (T1, ARM-alf-layla step 8)
$0.00. No API call was made. The session's translation limb is the lead's own rendering of the
copy-text's fifth night, which is free and is never ledgered (charter §3, A4). The study limb is a
census over four published English Nights at ten loci fixed in advance: it is answerable exactly,
by reading four books, and there is nothing in it for a model to judge. The pre-run critic this
project buys for a judgment run buys nothing for a census; RS-20260814-saj-carriage (span A's
study limb, the same shape) is the precedent.
What the money would have bought and did not, recorded so the choice is on the record rather
than implied: a reader-side run asking whether the level structure is recoverable from each hand's
English. Note (brb), written the session before, says the direct form of that question is a
ceiling; the indirect form is a design and not an afterthought, and it is now ARM-alf-layla
step 9.
UTC day 2026-08-23 running total after S215: $0.842453050 of $5.00 across two sessions. $4.157546950 unspent.
S216 — 2026-08-23 — E-20260823c-member-move (T5, ARM-member-move step 1)
$0.871405250 against a declared ceiling of $3.40, built from max_tokens per note (abc):
1600/700, 316 calls, worst case $2.77 plus $0.63 of retry allowance. 316 calls, 316 parsed, 0
dead, no retries used.
| stage | calls | actual |
|---|---|---|
pre-run critic, 2 rounds (P1) |
2 | $0.132079000 |
| G — the Persian gate, 52 loci × 2 seats | 104 | $0.113275500 |
| R — the rendering pool, 106 cola × 2 seats | 212 | $0.626050750 |
| total | 318 | $0.871405250 |
Per-call rates read from a 6-call probe before the run, not from config/models.md: P1
~$0.0010, P2 ~$0.0005 on the short bodies this design sends. The declared de-scope (drop the 13
AFFIX loci if stage G exceeded $0.60) was not triggered — stage G came in at $0.113.
The translation limb is free and is not ledgered (charter §3, A4): 27 prose blocks and 42 bayts rendered by the lead. The contamination measurement also cost $0 — Eastwick 1852 was reachable as Internet Archive OCR, which is the first time a published English Gulistan has been available to this project without paying seats to transcribe page images (S213 paid for exactly that).
UTC day 2026-08-23 running total after S216: $1.713858300 of $5.00 across three sessions. $3.286141700 unspent.
UTC day 2026-08-22 running total after S213: $2.625694250 of $5.00 across three sessions. $2.374305750 unspent.
S217 — 2026-08-24 — E-20260824-eye-or-ear (T3, ARM-echo-threshold step 2)
$2.621514975 against a declared ceiling of $3.40, built from max_tokens per note (abc):
1600/700. Key-usage figure moved 137.501387742 → 140.122902717, a delta of 2.621514975 —
recorded as provenance, not as a cross-check (note (bof)).
| stage | calls billed | actual |
|---|---|---|
pre-run critic, 2 rounds (P1) |
2 | $0.139775500 |
| D — detection, 130 items × 3 seats | 1,271 | $2.213162225 |
| N — neutral framing, 56 items × 2 seats | 112 | $0.170304250 |
| R — naturalness ranking, 17 loci × 2 seats | 34 | $0.098273000 |
| total | 1,419 | $2.621514975 |
Stage D was designed as 390 calls and billed 1,271, and the reason is an operator error, not the
design. A background dispatcher was still alive when a second was launched; each computed its
finished-cell set once at start, so 368 of 390 cells were called two to five times. The run stayed
inside its declared ceiling and the analysis takes one call per cell (RS-20260824 §11a), but
roughly $1.60 of this session's spend bought nothing that was asked for. The remedy is note
(brf): a dispatcher that resumes from a file must hold a lock, or the launcher must confirm no
prior process is alive.
The de-scope declared in design.md §10 was not triggered: the 6-call probe put the run at
about $0.75 and the ceiling was never approached by the design itself.
The translation limb is free and is not ledgered (charter §3, A4): 44 blocks — 17 prose, 27 bayts — rendered by the lead. The contamination measurement against Eastwick 1852 also cost $0 (Internet Archive OCR).
UTC day 2026-08-24 running total after S217: $2.621514975 of $5.00 across one session. $2.378485025 unspent.
2026-08-24 (S218) — E-20260824b-hariri-hands, the second-author census
$0.483099750 against a declared ceiling of $0.90 (raised from $0.60 by the round-1 critique,
which took stage L from 82 calls to 98 and added a critic round that had not been budgeted).
max_tokens 700 on stage L, 2500 on the nine re-dispatched cells, 7000 on the critic.
| stage | calls billed | actual |
|---|---|---|
| pre-run critic, attempt 1 — empty replies, bought nothing | 2 | $0.041604000 |
pre-run critic, round 1 (P1, P2) |
2 | $0.076446500 |
| L — addition check, 49 items × 2 seats | 98 | $0.310099250 |
| L — re-dispatch of 9 cells that hit the token cap | 9 | $0.054950000 |
| total | 111 | $0.483099750 |
$0.0416 of this bought nothing and the cause is note (bnk), already on the books: max_tokens
2000 caps the answer plus the hidden reasoning, and P1 spent all of it thinking. The same fault
recurred on 9 of the 98 stage-L calls at max_tokens 700; those were re-dispatched once, in a
single foreground process, per note (brf) — no second dispatcher was launched and the billed
call count (111) matches the designed count plus the nine repeats.
The translation limb is free and is not ledgered (charter §3, A4): the whole of al-Ḥarīrī's first maqāma — 139 prose cola and 9 verse lines — rendered by the lead. The contamination measurement, the census, the chance rate and the verifier all cost $0 (Internet Archive OCR and stdlib Python).
Key-usage figure moved 140.122902717 → 140.606002467, a delta of 0.483099750 — recorded as provenance, not as a cross-check (note (bof)).
UTC day 2026-08-24 running total after S218: $3.104614725 of $5.00 across two sessions. $1.895385275 unspent.
2026-08-26 (S225) — E-20260826c-register-cost, and ARM-alf-layla closed
Third session of the same UTC day. Opening snapshot 145.871654426, +0.266959000 above S224's closing figure — non-project drift, thirty times the ordinary figure and recorded as such.
Pre-flight: two stages, worst case from max_tokens (note (abc)). Critic — two seats on a
~5,500-word design plus the exact item texts, cap 9000 from the start per note (brr), worst
case $0.15. Run — 204 single-passage calls, three seats, caps 2500 / 2500 / 3000, worst case
$2.64. Declared ceiling $2.80, against a day headroom of $3.943045025.
| stage | calls | billed |
|---|---|---|
C pre-run critic, P1 + P2 — NEEDS REDESIGN, 18 findings, 5 BLOCKING |
2 | $0.118089000 |
| run, 34 conditions × 3 seats × 2 replicates | 204 | $1.851452175 |
| total | 206 | $1.969541175 |
Came in at 70% of the ceiling, and 30 calls of 204 died — all of them QR, 29 with an empty
body at finish_reason: length on the cap note (brr) certified after 56 clean calls at S224. The
money for those calls was billed and bought nothing; that is note (brt), and the design's own
10% failure criterion fired and is reported on the result page rather than hidden.
The money bought a bound, not an answer. The registered question could not be answered, and what the run established instead is what the instrument can and cannot see: 22 of 22 blatant seeded defects, 0 of 10 subtle ones. The critic's five BLOCKING findings, at $0.118, caught a stale item file, an item that did not test what it claimed, an unscoreable locus and a control whose window did not contain its own defect — all before a scored call was dispatched.
The translation limb cost $0 and is not ledgered (charter §3, A4): span J of «ألف ليلة وليلة» rendered whole — 1,411 Arabic words into 2,754 English, two poems, sixteen logged decisions — plus the source assembly, the collation against a second Arabic witness, the dependence check against Lane and Burton, the craft report closing the arm, and the 94-check verifier, all stdlib Python.
Key usage 145.871654426 → 148.342551051, delta 2.470896625 against a billed total of 1.969541175 — the delta is larger by $0.501355450. Per-request costs are ledgered.
UTC day 2026-08-26 running total after S225: $3.026496150 of $5.00 across three sessions. $1.973503850 unspent.
2026-08-26 (S224) — E-20260826b-radif, the radif taken apart
Second session of the same UTC day. Opening snapshot 144.764653501, identical to S223's closing figure.
Pre-flight: three stages, worst case from max_tokens (note (abc)). Critic — two seats on a
~5,000-word design plus the exact materials, cap 9000 from the start per note (brr),
worst case $0.13. Control C1 — 34 items, P2, cap 3000, $0.39. Judging — 168 forced
choices, three seats, cap 3000, $2.39. Re-dispatch allowance 10%, $0.29. Declared ceiling
$3.20, staged so it could be stopped after the control.
| stage | calls | billed |
|---|---|---|
C pre-run critic, P1 + P2 — NEEDS REDESIGN, 15 findings, 7 BLOCKING |
2 | $0.108660750 |
control C1, P2, 28 real + 6 planted — gate PASS 6 of 6 |
34 | $0.125767500 |
| judging, 7 loci × 4 pairs × 2 orders × 3 seats | 168 | $0.605613675 |
| total | 204 | $0.840041925 |
Came in at 26% of the ceiling, and the reason is that nothing had to be re-dispatched: 0 dead, 0
retries, 0 unparsed across 204 calls, because every seat was budgeted at the cap note (brr)
prescribes rather than at the cap the visible answer needs. The one truncation was P2 at
max_tokens 9000 on the critic prompt — the fourth form of (brr) — and it cost nothing extra
because no re-dispatch was attempted; three findings are all this session has from that seat.
The money bought the design, not the answer. The critic's seven BLOCKING findings changed the manipulation itself twice, and the 168 judging calls that followed returned a withheld result. That is the correct order for the spend to come out in.
Key usage 144.764653501 → 145.604695426, delta 0.840041925 against a billed total of 0.840041925 — exact to 1e-11.
The translation limb cost $0 and is not ledgered (charter §3, A4): six Sa'di ghazals rendered whole — 39 bayts, 45 rhyming positions — plus the source freeze, the mechanical rhyme grading, the contamination measurement against a 68,000-word comparator, the whole analysis and the 218-check verifier, all stdlib Python.
UTC day 2026-08-26 running total after S224: $1.056954975 of $5.00 across two sessions. $3.943045025 unspent.
2026-08-26 (S223) — E-20260826-balanced-period, the operationalisation gate and the critic
New UTC day; the ledger resets to $5.00. Opening key snapshot 144.547740451, +0.009646800 above S222's closing figure — ordinary between-session non-project drift, note (abf).
Pre-flight: two stages only, both of them talk-about-the-design rather than data. Gate — three
seats, one short prompt, worst case from max_tokens 1500 at list prices $0.03. Critic — two
seats on a ~3,000-word design, worst case at max_tokens 3500 $0.12. Declared ceiling
$0.40. Everything else in this session is arithmetic and costs nothing.
| stage | calls | billed |
|---|---|---|
gate operationalize, three seats — the metric fixed before any measurement |
3 | $0.034493650 |
P1 returned an EMPTY body and P2 a truncated one at max_tokens 1500; both re-dispatched at 3500 |
2 | $0.029599250 |
C pre-run critic, two seats — NEEDS-REDESIGN, six BLOCKING |
2 | $0.064606500 |
both critic seats truncated at 3500; both re-dispatched at 9000, and P2 truncated again |
2 | $0.088213650 |
| total | 9 | $0.216913050 |
Four of nine calls were re-dispatches after a truncated or empty body, and they are 54% of the
spend. Note (bph) in a fifth form and its third seat: P1 (openai/gpt-5.6-terra) spends the
whole cap on hidden reasoning and returns nothing at 1500 on a 250-word prompt, and P2
(google/gemini-3.6-flash) truncated at 3500 and again at 9000 on the critic prompt, so two of
its findings are all this session has from it. Budget any P1 or P2 stage that wants prose at
3500 from the start, and any long-context critic at 9000. Unlike S222's QR episode the money
did buy something — P1's complete critic response is the reason six BLOCKING findings were acted
on rather than two.
Key usage 144.547740451 → 144.764653501, delta 0.216913050 against a billed total of 0.216913050 — exact to 1e-9, the shape a session gets when its calls are few and settle before the closing snapshot.
The translation limb cost $0 and is not ledgered (charter §3, A4): al-Ḥarīrī's second
Assembly «المقامة الحلوانية» rendered whole under R50 — 140 prose cola, 17 verse lines — plus the
contamination measurement, the recovery of two published hands from OCR, all three segmentations,
the 10,000-draw shuffle nulls and the 421-check verifier, all stdlib Python.
UTC day 2026-08-26 running total after S223: $0.216913050 of $5.00, one session. $4.783086950 unspent.
2026-08-25 (S222) — E-20260825c-worth-paying, the three-policy forced choice
Pre-flight, written into the frozen design after the critic corrected its arithmetic: worst case
from max_tokens at list prices, $0.00467 per call over 483 data calls = $2.26, expectation
≈ $1.20 from E-20260825b's realised $0.002472; declared ceiling $2.60 with hard per-stage
stops (> $0.40 after the controls → abort · > $1.60 after H → drop V · > $2.10 after B → drop
V · > $2.45 after V → drop R). No stop fired. Opening key snapshot 142.742144401.
| stage | calls | billed |
|---|---|---|
C pre-run critic, two seats, no data — both NEEDS-REDESIGN, 15 findings accepted |
2 | $0.097545250 |
F floor control (intact prose vs the same words deranged) — passed 24 of 24 |
24 | $0.045218375 |
F6 near-duplicate control (one text against itself, one word changed) |
24 | $0.050643175 |
T pilot, registered as proceed-or-abort only |
9 | $0.031694450 |
H rhyme-forward vs restrained, 12 segments × 3 conditions × 2 orders × 3 seats |
216 | $0.614050050 |
B balanced vs restrained, 8 × 2 × 2 × 3 |
96 | $0.299411325 |
V rhyme-forward vs balanced, 8 × 2 × 2 × 3 |
96 | $0.285239700 |
R repeatability |
18 | $0.059276650 |
40 QR calls returned EMPTY at finish_reason: length and were re-dispatched once — bought nothing |
40 | $0.304777775 |
| total | 485 | $1.787856750 |
Seventeen per cent of this run's spend bought nothing, and it is one seat's behaviour. QR
(qwen/qwen3.7-max) spends its whole max_tokens on hidden reasoning and returns an empty body;
40 of 483 data calls (8.3%) came back that way even after the cap was raised at the F gate from
1200 to 1600. All 40 re-dispatched at 3000 and all 40 parsed, so 0 cells were dropped — the
money bought the completeness, not the answers. Note (bnk)'s family, in a fourth form.
The token caps were also raised once at the F gate before any main-stage call, from
600/600/1200 to 800/1400/1600, because P2 truncated 2 of its 8 floor calls; recorded on the design
rather than done quietly.
Key usage 142.742144401 → 144.538093651, delta 1.795949250 against a billed total of 1.787856750 — provenance, not a cross-check (note (bof)); $0.0081 above, the non-project direction. The opening snapshot also read $0.087059725 above S221's closing figure.
The translation limb cost $0 and is not ledgered (charter §3, A4): al-Ḥarīrī's first Assembly
rendered whole a third time under the new R50, 139 prose cola, plus the contamination
measurement and every chime and balance grading, all stdlib Python.
UTC day 2026-08-25 running total after S222: $2.998953875 of $5.00 across three sessions. $2.001046125 unspent.
2026-08-25 (S220) — E-20260825-level-recovery, the reference-recovery run
$0.595071375 against a declared ceiling of $1.40 (raised from v1's $1.10 by the pre-run critic,
which added a fourth arm — 18 more calls — that had not been budgeted; the same mechanism as S218).
max_tokens 2000 with reasoning.max_tokens 900 on stage P, 7000/4000 on the critic.
| stage | calls billed | actual |
|---|---|---|
pre-run critic (P1, P2) — bought no data and produced a fourth arm |
2 | $0.120788250 |
| P — 4 arms × 3 seats × 6 bodies | 72 | $0.412440800 |
P — the five finish_reason: length originals, re-dispatched once |
5 | $0.061842325 |
| total | 79 | $0.595071375 |
Five calls came back at the token cap, all on QR, which wrote its full working into the answer
body; three recovered on the single permitted repeat and two were dropped and counted under the
design's own failure criterion. Note (bnk) again, in its visible-output form rather than its
hidden-reasoning form.
The translation limb is free and is not ledgered (charter §3, A4): the copy-text's sixth night
whole, 880 Arabic words → 1,619 English. The collation against the second Arabic witness and the
dependence_check against Lane and Burton both cost $0 (stored Project Gutenberg texts and
stdlib Python).
Key-usage figure moved 141.368450912 → 141.963522287, a delta of 0.595071375 — recorded as provenance, not as a cross-check (note (bof)).
UTC day 2026-08-25 running total after S220: $0.595071375 of $5.00 across one session. $4.404928625 unspent.
2026-08-25 (S221) — E-20260825b-flippancy, the declared-reason run
$0.616025750 against a declared ceiling of $2.60 (raised from v1's $2.20 by the pre-run critic,
which tripled stage A, doubled stage G and added a pilot — the same mechanism as S218 and S220).
max_tokens 1200 on P1/P2, 2000 on QR, reasoning.max_tokens 600 on the scored stage;
7000/4000 on the critic.
| stage | calls billed | actual |
|---|---|---|
C pre-run critic (P1, P2) — bought no data and rewrote the primary item |
2 | $0.098553000 |
pilot — G4 × RHY/DRH × 3 seats, run before the 135 were committed |
6 | $0.014791300 |
P — 15 segments × 3 arms × 3 seats |
135 | $0.333775925 |
A — audibility, 15 × RHY/DRH × 3 seats |
90 | $0.124388225 |
G — gravity, 15 × RHY/DRH × P2 |
30 | $0.018693000 |
R — repeatability, 5 × 3 arms × P2 |
15 | $0.022036500 |
the one finish_reason: length cell, re-dispatched once and dropped |
1 | $0.003787800 |
| total | 279 | $0.616025750 |
Note (bnk) fired in a third form. QR appended a duplicate JSON object and a </think> block
after a valid answer on five audibility calls — recovered by tightening the extraction regex to the
first flat object, at no cost — and on one call wrote its whole working into the body and hit the
cap twice, which is a genuine dropped cell.
The translation limb cost $0 and is not ledgered (charter §3, A4): al-Ḥarīrī's first Assembly
rendered whole a second time under R48, 139 prose cola, plus the mechanical R49 control. The
contamination measurement and every chime grading cost $0 (stdlib Python and stored comparators).
Key-usage figure moved 142.042846726 → 142.655084676, a delta of 0.612237950 — recorded as provenance, not as a cross-check (note (bof)); it sits $0.00378 below the billed total, the cost of the truncated re-dispatch. The opening snapshot also read $0.079324439 above the figure S220 recorded at its close, which is non-project spend on the same key and is why the delta is provenance.
UTC day 2026-08-25 running total after S221: $1.211097125 of $5.00 across two sessions. $3.788902875 unspent.
2026-08-27 (S226) — E-20260827-declared-play, and ARM-declared-function closes
Declared ceiling $2.20 (design v2 §9.6, revised down from v1's $3.00 when the pre-run critic cut
the run from 384 calls to 240). Seats P1 openai/gpt-5.6-terra, P2 google/gemini-3.6-flash,
third seat QR qwen/qwen3.7-max — GL z-ai/glm-5.2 was the frozen third seat and returned an
empty body at finish_reason: length on 1 of 6 pilot calls, which is the exact condition the
design's frozen fallback rule names. QR's cap fixed at 4000 by a 2-call mechanical probe
(2/2 parsable, ~2,600 completion tokens each) — note (brt).
| stage | calls billed | actual |
|---|---|---|
C pre-run critic (P1, P2) — both BLOCKING on the same two defects, no data |
2 | $0.115140750 |
pilot on the frozen third seat GL |
6 | $0.028242181 |
cap probe on the fallback seat QR |
2 | $0.024302100 |
P — 20 units × 4 arms × 3 seats, plus 17 re-dispatches at a repaired cap |
257 | $1.373982900 |
| per-request sum | 267 | $1.541667931 |
2 of 240 cells dropped (0.83%), both QR on the two longest passages, against an F3 void
threshold of 10%. Providers: OpenAI 80, Alibaba 80, Google 40, Google AI Studio 40. Per-seat:
P1 $0.300810, P2 $0.476293, QR $0.596879.
The ledgered figure is the key delta, $1.747847221, not the per-request sum — see the snapshot row and note (brw). This is the first session in this project able to account for a divergence between the two rather than merely record it.
Free and not ledgered (charter §3, A4): the whole «گلستان» باب دوم ۱–۱۰ rendered under R52
(68 blocks), the machine enumeration and adjudication of all 62 locus candidates, the alignment key,
the dependence check on all six pairs, the notes census, and the 306-check verifier.
UTC day 2026-08-27 running total after S226: $1.747847221 of $5.00, one session. $3.252152779 unspent.
2026-08-27 (S227) — E-20260827b-worth-paying, and ARM-worth-paying closes
Pre-flight. Declared ceiling $2.20 against a UTC-day headroom of $3.252152779. Note (abc):
the worst case is built from the cap the request permits. A 6-call cap probe was run first, on
this task shape, per note (brt) — measured completion tokens P1 32–35, P2 601–726, QR
1058–1264 — and the caps were set at twice the largest measured figure, floored at 1200: P1 1200,
P2 1500, QR 2600. The cap-worst-case then came to $3.83, which exceeds both the ceiling and the
day's headroom; that is recorded in raw/cap-probe.json rather than hidden, together with the
statement that what protects the ceiling is the stage stop, not the cap. Stages were dispatched
M, D, F, X, R with the cumulative spend checked after each.
| stage | calls | actual |
|---|---|---|
C pre-run critic (P1, P2) — both NEEDS-REDESIGN, and the design was rebuilt |
2 | $0.063360250 |
T cap probe, 3 seats × 2 conditions |
6 | $0.022633475 |
M main, RHY vs ORD, 9 × 4 × 2 × 3 |
216 | $0.828626550 |
D placebo, DCH vs ORD, 9 × 2 × 2 × 3 |
108 | $0.400120450 |
D placebo at I2 — UNREGISTERED extension, dispatched after the primaries were seen |
54 | $0.190531950 |
F operational floor |
12 | $0.030021600 |
X near-duplicate scale |
12 | $0.028672775 |
R repeatability |
18 | $0.067804375 |
| per-request sum, and the ledgered figure | 428 cells, 449 attempts | $1.631771425 |
8 of 378 preference cells void (2.1%) against an F4 threshold of 10%, all eight the same
openai/gpt-5.6-terra unterminated-string defect that a temperature-0 re-dispatch reproduces exactly
— note (brx). Per-seat and per-provider figures are in analysis.json.
The ledgered figure is the per-request sum, $1.631771425, which here is the larger of the two figures; see the snapshot row.
Free and not ledgered (charter §3, A4): «المقامة الحلوانية» rendered whole twice more under R43
and R48 (280 prose cola of new English), the R49 de-chime built and machine-checked, the 58 saj'
loci enumerated and adjudicated by script, the dependence measured on eleven pairs, and the 461-check
verifier with its three mutation tests.
UTC day 2026-08-27 running total after S227: $3.379618646 of $5.00, two sessions. $1.620381354 unspent.
S228 — 2026-08-27, E-20260827c-declared-page, ARM-balanced-period step 2 (T4)
Declared ceiling before spending: $0.25, for a pre-run critic pass on two seats and nothing
else. Everything the session measures is arithmetic over stored text and cost nothing; the
translation limb — «المقامة الدمياطية» rendered whole under the new R53, 182 prose cola — is the
lead's own and is not ledgered (charter §3, A4).
| stage | calls | actual |
|---|---|---|
C pre-run critic, first dispatch — both seats truncated at max_tokens 3500, note (brt) |
2 | $0.06540875 |
C pre-run critic, re-dispatch at 9000 — the findings that were acted on |
2 | $0.07994815 |
| per-request sum, and the ledgered figure | 4 | $0.14535690 |
The truncated first dispatch is ledgered and not written off, per note (brw): a cap that was verified for step 1's shorter design was carried over to a design half again as long, both replies were cut off mid-finding, and the second seat returned one finding of the four it had. The cost of finding that out is part of the cost of the critic pass.
Key usage 152.139205997 → 152.284562897, delta 0.145356900, against a per-request sum of 0.145356900 — exact to 1e-9. The opening snapshot read +0.063764175 above S227's close, which is non-project spend on the same key.
UTC day 2026-08-27 running total after S228: $3.524975546 of $5.00, three sessions. $1.475024454 unspent.
S229 — 2026-08-28, E-20260828-forced-half, ARM-radif step 2 (T2)
Declared ceiling before spending: $2.60, raised from $2.20 by the pre-run critic's BLOCKING 3,
which bought a fifth information condition (IL, the length- and form-matched Persian decoy) and
54 more calls. UTC day 2026-08-28 had no prior rows: the whole $5.00 was available. The six whole
renderings — «غزلیات» ۹۷, ۹۴, ۱۲۶ each rendered twice under the new R54, 21 bayts and 24 rhyming
positions — are the lead's own and are not ledgered (charter §3, A4), as are the mechanical
grading and both contamination measurements.
| stage | calls | actual |
|---|---|---|
C pre-run critic, P1 + P2, cap 12000 — both NEEDS-REDESIGN, 16 findings, neither truncated |
2 | $0.071679500 |
T cap probe on this task shape, I0/IS/IL × 3 seats — no truncation; caps set 600 / 1200 / 2500 |
9 | $0.035806700 |
C1 sense-equivalence gate, P2, 15 items — 6 of 6 planted caught, 3 real pairs flagged |
15 | $0.045930750 |
F2 operational floor — 12 of 12 |
12 | $0.035803850 |
M main, 9 windows × 5 conditions × 2 orders × 3 seats — 0 void, 0 re-dispatches |
270 | $0.981363500 |
F4 repeatability, 18 stratified cells — 14 of 18 |
18 | $0.060042650 |
| per-request sum, and the ledgered figure | 326 | $1.230626950 |
Note (brt) did not fire, for the first time in five sessions, because the cap probe ran on this task shape before the main run and the critic cap was raised to 12000 on the note's own record. $1.369373050 of the declared ceiling unspent.
Key usage 152.551501597 → 153.782128547, delta 1.230626950 against a per-request sum of 1.230626950 — exact to 1e-9. The opening snapshot read +0.266938700 above S228's close, non-project spend on the same key.
UTC day 2026-08-28 running total after S229: $1.230626950 of $5.00, one session. $3.769373050 unspent.
S230 — 2026-08-28, E-20260828-purchased-figure, ARM-gulistan step 1 (T1)
Declared ceiling before spending: $1.60, raised to $2.40 by the pre-run critic, whose two
BLOCKING findings required a third translation arm (an unrhymed verse control) and forty more
judging items. UTC day 2026-08-28 had one prior row, S229's $1.230626950, so $3.769373050 was
available. The whole translation limb — «گلستان» باب دوم حکایات ۱۱–۲۰, 87 blocks, 1,964 English
words — the chapter collation, the mechanical rhyme coding and both contamination measurements are
the lead's own and are not ledgered (charter §3, A4).
| stage | calls | actual |
|---|---|---|
C pre-run critic, C1 + C2, cap 4000 — both NEEDS-REDESIGN, 12 findings; C2 truncated and was re-dispatched once at cap 12000, note (brr) |
3 | $0.093240500 |
UNR the control arm — 39 bayts as unrhymed English verse, x-ai/grok-4.5, 0 void |
39 | $0.150823200 |
T cap probe on this task shape, the six longest prompts × 3 seats — 18 of 18 parsed, none truncated; the registered acceptance rule passed |
18 | $0.078744400 |
M main, 140 items × 3 seats — 0 lost; 9 P1 bodies re-parsed rather than re-dispatched, note (brx) |
420 | $1.209492625 |
| per-request sum, and the ledgered figure | 480 | $1.532300725 |
Note (brt) did not fire again, for the second session running: the cap probe went first and the
main run's 3000-token cap was the one it verified. Note (brr) fired on the critic stage exactly
as it says — P2 truncated a long critic prompt at 4000 — and cost one re-dispatch, $0.029186250.
$0.867699275 of the declared ceiling unspent.
Key usage 153.782128547 → 155.314429272, delta 1.532300725 against a per-request sum of 1.532300725 — exact to 1e-9. The opening snapshot was identical to S229's close: no non-project spend on the key between the two sessions, the first time this month.
UTC day 2026-08-28 running total after S230: $2.762927675 of $5.00, two sessions. $2.237072325 unspent.
S234 — 2026-08-30, E-20260830b-rhyme-family, ARM-rhyme-family step 1 (T2)
UTC day 2026-08-30, second session. After S233's $1.420460500 the day stood at $3.579539500.
Declared ceiling $2.60, set in design v2 §10 and not raised; spent $2.190483225,
$0.409516775 unspent. The translation limb T-hafez-amad-R57-v1 (Hafez غزل ۱۷۶ whole), the
Ganjoor fetch, all extraction, the CMUdict maximisation, the contamination measurement against 136
Payne odes and the verifier are the lead's own and are not ledgered (charter §3, A4).
| stage | calls | actual |
|---|---|---|
C pre-run critic, P1 + P3, cap 14000 — both NEEDS-REDESIGN, 27 findings; note (brr) honoured, P2 kept off the long prompt |
2 | $0.112091900 |
PR instrument probe — four seats × two caps × two reasoning settings, on both task shapes |
16 | $0.310233425 |
F2 the rhyme-search stage, abandoned after the probe and replaced by G (amendment v3) |
5 | $0.152718000 |
F1 contextual senses, P1, cap 2500 |
42 | $0.371738400 |
F1 re-buys at cap 12000 after finish_reason: length |
2 | $0.058684000 |
G word lists, P1, cap 2500 |
38 | $0.287182000 |
G re-buys at cap 12000 |
2 | $0.016878000 |
S realised coverage, P2, cap 6000 |
38 | $0.741221250 |
S re-buys at cap 12000 |
6 | $0.139736250 |
| per-request sum, and the ledgered figure | 151 | $2.190483225 |
$0.462951425 of that — 21% — went on instrument work rather than on the question: the probe, and
the five abandoned F2 calls. It bought the amendment that made the predictor mechanical and six
times cheaper, so it is not waste, but it is the largest instrument share this project has recorded
and note (bsi) exists so that the next design reaches the split without paying for it.
Note (bsf) fired a fourth session running, and on four seats at once: at cap 2500 on a nine-line
answer, P2 and qwen/qwen3.7-max both truncated on the very first item probed, and
z-ai/glm-5.2 returned an empty body after 6000 completion tokens. At 6000 the panel seats came
back clean and six of thirty-eight stage-S bodies truncated anyway. Note (bsh)'s remedy —
price from a probe, not from the cap — is what kept this run inside its ceiling; note (bsh)'s other
half held too, since reasoning: {"effort": "low"} cut x-ai/grok-4.5 from $0.0162 to $0.0027 on
one task shape and did nothing at all for P1 on another.
Key usage 159.995500773 → 162.059280998, delta 2.063780225 against a per-request sum of 2.190483225 — NOT exact, and short by $0.126703000, which is the opposite sign to S233's discrepancy. The conservative figure is ledgered. The session also opened $0.212799500 above S233's close, the fourth consecutive inter-session gap.
UTC day 2026-08-30 running total after S234: $3.610943725 of $5.00, two sessions. $1.389056275 unspent.
S233 — 2026-08-30, E-20260830-leaf-contract, ARM-leaf-contract step 1 (T4)
UTC day 2026-08-30 had no prior row, so the whole $5.00 was available. Declared ceiling
$1.80 — raised from the design's first draft of $0.60 before dispatch, with the reason written
into design v2 §7: the disputed rhyme set came out at 91 unique pairs, and v1's response to an
overrun was to truncate it, which both pre-run critic seats showed would bias the measure against
exactly the licences the design says it respects. Ledgered $1.420460500, $0.379539500 of the
ceiling unspent. The translation limb (T-hafez-bekonad-R57-v1, Hafez غزل ۱۸۷ rendered whole twice
under R57), all extraction, all coding and scansion of 444 printed lines, the census and match
reuse, and every contamination measurement are the lead's own and are not ledgered (charter §3,
A4).
| stage | calls | actual |
|---|---|---|
C pre-run critic, P1 + P3, cap 12000 — both NEEDS-REDESIGN, 28 findings; neither truncated |
2 | $0.087631000 |
A blind rhyme adjudication, 115 items × 3 seats, cap 900 |
345 | $1.151420300 |
A2 re-buy of the 10 bodies that returned no parsable bit, cap 3000 — all 10 clean |
10 | $0.159115200 |
| two truncated bodies discarded before the run restarted — billed, bodies deleted; note (brw) | 2 | $0.022294000 |
| ledgered figure | 359 | $1.420460500 |
Note (bsf) fired for the third session running, in a form that earns its own note (bsh). At
max_tokens 900 on a task whose answer is two lines, 8 of 115 P1 bodies returned
finish_reason: length with content: null — 900 completion tokens, every one of them hidden
reasoning, and no answer at all. Note (brx) could not fire: there was nothing to re-parse. A
reasoning: {"effort": "low"} passthrough was added and the provider billed 900 reasoning tokens
anyway; the parameter had no effect on this model and provider. The re-buy at cap 3000 returned
all ten clean, and its cost is in the table rather than netted out (note (brw)).
Key usage 158.280830373 → 159.782701273, delta 1.501870900 against a per-request sum of 1.398166500 — NOT exact, and this is the first non-exact reconciliation this project has recorded. $0.022294000 of the $0.103704400 gap is the two discarded truncations, which are ledgered above, bringing the honest project figure to $1.420460500. $0.081410400 remains unexplained. The three preceding session-starts each opened with an inter-session gap on this key ($0.353300441, $0.872971860, $0.333671850), so non-project spend on the key is the standing explanation and a within-session gap is consistent with it — but nothing here establishes that, and the residue is recorded rather than attributed.
UTC day 2026-08-30 running total after S233: $1.420460500 of $5.00, one session. $3.579539500 unspent.
S232 — 2026-08-29, E-20260829b-inversion-price, ARM-inversion-price step 1 (T3)
UTC day 2026-08-29, second session. After S231's $0.197706300 the day stood at $4.802293700
available. Declared ceiling $2.20; spent $1.208750650, $0.991249350 of the ceiling
unspent. The translation limb (T-hafez-darad-R56-v1, two Hafez ghazals in four machine-checked
arms, 72 English couplets), the census selection, the archaism map and all contamination
measurements are the lead's own and are not ledgered (charter §3, A4).
| stage | calls | actual |
|---|---|---|
C pre-run critic, P1 + P3, cap 12000 — both NEEDS-REDESIGN, 30 findings, 9 BLOCKING; neither truncated |
2 | $0.126422400 |
S within-idiom well-formedness rating, 90 bodies × 3 seats, cap 300 |
270 | $0.677516100 |
S re-dispatch of 6 bodies at cap 1200 — note (bsf), below |
6 | $0.016076000 |
F within-lexis forced choice, 36 × 3, cap 1200 |
108 | $0.388736150 |
| per-request sum, and the ledgered figure | 386 | $1.208750650 |
Note (bsf) fired for the second time in two sessions, and it fired as written. At max_tokens
300, 92 of 270 rating bodies returned finish_reason: length — overwhelmingly P2, whose
hidden reasoning is variable. 86 of them had already delivered the SCORE: line and were
re-parsed, not re-bought (note (brx), saving 86 calls at roughly $0.22). Six lost the score
line entirely, were re-dispatched at cap 1200, and all six returned clean for $0.016076000.
The stage F cap was raised to 1200 before that stage ran. Note (brw) applies and the
re-dispatch cost is in the table, not netted out.
Key usage 156.738407873 → 157.947158523, delta 1.208750650 against a per-request sum of 1.208750650 — exact. The opening snapshot was $0.872971860 above S231's close — the second inter-session gap in a row, and larger than the first ($0.353300441). It is non-project spend on the key and is recorded in the snapshot table.
UTC day 2026-08-29 running total after S232: $1.406456950 of $5.00, two sessions. $3.593543050 unspent.
S231 — 2026-08-29, E-20260829-radif-hands, ARM-radif-hands step 1 (T5)
UTC day 2026-08-29 had no prior row, so the whole $5.00 was available. Declared ceiling
$1.20; spent $0.197706300, $1.002293700 unspent. The translation limb
(T-hafez-ghazals-radif-R55-v1, two Hafez ghazals whole), the 495-ghazal radif census, all
matching, all carriage coding across three published hands and all three contamination measurements
are the lead's own and are not ledgered (charter §3, A4).
| stage | calls | actual |
|---|---|---|
C pre-run critic, P1 + P3, cap 12000 — both NEEDS-REDESIGN, 25 findings, 9 BLOCKING; neither truncated, no re-dispatch |
2 | $0.092818000 |
G blind grammatical classification, 9 radifs × 3 seats, cap 1200 — 29 of 29 usable |
27 | $0.104888300 |
| per-request sum, and the ledgered figure | 29 | $0.197706300 |
Note (brr) did not fire: P2 was kept off the long critic prompt by the design, which is what
the note asks for. Note (brt) fired in a new form and is written up as note (bsf): the
max_tokens cap was probed on the first three classification calls and passed, and P2 then
truncated on a later item of the same shape, because its hidden reasoning is variable — 1,123 of
1,196 completion tokens on that call. Note (brx) fired: the truncated body had every checklist
bit and only lost its trailing gloss, so it was re-parsed, not re-bought, at $0.
Key usage 155.667729713 → 155.865436013, delta 0.197706300 against a per-request sum of 0.197706300 — exact. The opening snapshot was $0.353300441 above S230's close; that is non-project spend on the key between sessions and is recorded in the snapshot table.
UTC day 2026-08-29 running total after S231: $0.197706300 of $5.00, one session. $4.802293700 unspent.
UTC day 2026-09-04 — S243, ARM-french-in-russian step 1 (T4), E-20260904-french-in-russian
New UTC day; S243 is its first session, so the whole $5.00 was available. Pre-flight ceiling
declared in design-v2.md §7: $0.50, built from the max_tokens cap and from prices read
from the API in this session (note (bsw)): openai/gpt-5.6-terra $2.00/$12.00,
x-ai/grok-4.5 $2.00/$6.00, google/gemini-3.6-flash $0.75/$3.75 and the reserve
qwen/qwen3.7-max $1.48/$4.42 — all four unchanged from the S242 re-read, including the one
that moved upward then, so no ceiling was mis-built and no revisit trigger fires.
Spent $0.208631650, $4.791368350 of the day unspent. The 38 loci, the collation against a
second Russian witness, the whole translation limb (2,780 Russian words), all 96 lead codings, the
ABBYY italic recovery and every contamination measurement are the lead's own and are not
ledgered (charter §3, A4).
| stage | calls | actual |
|---|---|---|
C pre-run critic, two seats, cap 12000 — both NEEDS-REDESIGN, 31 findings, 11 BLOCKING; neither truncated |
2 | $0.0896509 |
S1 blind second coder, 96 cells in batches of 8, cap 2000 — 6 of 12 batches truncated, 81 of 96 cells returned |
12 | $0.0894150 |
S1 refill, the 15 missing cells in batches of 4, cap 6000 — 4 of 4 clean |
4 | $0.0295658 |
| per-request sum, and the ledgered figure | 18 | $0.208631650 |
Note (bsf) fired again and harder than it has before. google/gemini-3.6-flash truncated half
the second-coder batches at a 2,000-token cap on a task whose visible output is eight short lines of
the form LI.11/GAR KEPT; completion tokens pinned at exactly 1,996 on every truncated call. The
same items at four per batch and a 6,000-token cap finished every time. The overrun is the hidden
reasoning, not the answer, and it scales with the number of items in the batch — which is the
sharpest form the note has taken. Nothing was lost: the truncated bodies carried no parsable lines,
so the missing cells were re-bought, not re-parsed, and the re-buy is in the table.
Note (bsw) was honoured: every dispatched seat's price was read from /api/v1/models in this
session before the ceiling was set, and one price new to the project's table was recorded.
Key usage 251.723042776 → 251.931674426, delta 0.208631650 against a per-request sum of 0.208631650 — exact. The opening snapshot was $0.503849600 above S242's close; per note (bso) that is recorded as an observation and nothing is claimed about how much of it is endpoint lag and how much is non-project spend on the key.
UTC day 2026-09-04 running total after S243: $0.208631650 of $5.00, one session. $4.791368350 unspent.