Repository path: workshop/experiments/E-20260902-rhyme-slot/amendment-v2-1-probe.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260902-rhyme-slot-v2-1 |
| status | frozen |
| created | 2026-09-02 |
| updated | 2026-09-02 |
| links | workshop/experiments/E-20260902-rhyme-slot/amendment-v2.md, config/models.md, config/budget.md, wiki/method-notes.md |
Amendment v2.1 — the seat probe, and what it cost the design
Run after v2 was frozen and before any item of the run was bought, on the single ghazal
sh35 (the lead's own poem, so no published hand's item is spent on instrument work). Five probe
calls, $0.094771750. This is note (bsf) firing for the fifth time — a cap is a per-seat,
per-task-shape measurement, not a number carried across a session — and it fires in a form the note
does not yet cover.
What the probe found
| seat | shape | cap | result | cost |
|---|---|---|---|---|
P1 openai/gpt-5.6-terra |
gloss (stage G) | 2500 | clean, 14 of 14 parsed, 1,635 completion tokens | $0.020452 |
P2 google/gemini-3.6-flash |
locate (stage L) | 2500 | parsed; 8 of 14 senses located, 2,486 completion tokens — at the cap | $0.009703 |
P2 google/gemini-3.6-flash |
locate, same item | 6000 | parsed; 4 of 14 located, 5,833 tokens | $0.022254 |
P1 openai/gpt-5.6-terra |
locate | 2500 | finish_reason: length, empty body |
$0.030978 |
P3 x-ai/grok-4.5, reasoning: {"effort":"low"} |
locate | 2500 | clean, 12 of 14 located, 1,659 tokens | $0.011330 |
The finding that matters, and it is new
P2 returned fewer locations at the larger cap than at the smaller one, on the same item, same
prompt, same temperature — 4 of 14 against 8 of 14, losing three it had already found. This is not
truncation. Note (bsf) so far covers a cap that is too small for a seat's hidden reasoning; here the
seat had more room, used it (5,833 completion tokens against 2,486), and its answer got worse. A
cap is therefore not only a pre-flight measurement but a parameter of the answer on this seat and
this shape, and a run that fixes a cap by probing one item at one setting is not thereby safe.
Recorded as method note (bst).
What changed in the design
- The locating seat is
P3x-ai/grok-4.5atreasoning: {"effort": "low"}, cap 2500.P2is out on this shape for the reason above.P1truncates and is the gloss seat. Stage G remainsP1openai/gpt-5.6-terra, cap 2500. The two stages remain at disjoint labs (OpenAI / xAI), which is the property both of step 1's critics blocked on. Both slugs also served as pre-run critics this session. That is not the rater-sharing defect the critics named — the predictor and the outcome still come from different seats, the critic calls saw the design and no data, and no call carries state into another — but it is recorded here rather than left to be noticed. contentlessgloss screen. The stage-G probe glossed the marked unit ofsh354a (مستغنیست) as the bare word "is". A gloss with no content word cannot be located by anyone, so such glosses are dropped from the item list asUNGLOSSABLE. This is a screen, not a repair: the gloss is not rewritten. Declared here, before the run, and applied bystages.py::contentless.- The declared ceiling is revised from $2.20 to $2.80, on the measured per-call figures rather
than an assumption: 54 gloss × $0.0205 + 58 locate × $0.0113 + 16 control × ~$0.015 ≈ $2.01,
plus $0.104133 already spent on critics and $0.094772 on this probe. The UTC day 2026-09-02 has
no prior row and $5.00 available; the revision is recorded in
config/budget.mdwith this page as its basis, per note (abc) — build the worst case from the cap the request actually permits.