Repository path: workshop/experiments/E-20260812-unlicensed-typography/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260812-critic |
| status | frozen |
| created | 2026-08-12 |
| updated | 2026-08-12 |
| internal-judgment-only | true |
| provisional | true |
| links | workshop/experiments/E-20260812-unlicensed-typography/design.md |
Pre-run critic pass — NEEDS-REDESIGN, 7 findings, 7 BLOCKING, 7 accepted
Seat openai/gpt-5.6-terra, one call, max_tokens 3,000, temperature 0, $0.0235655, 50.8s.
Whole body persisted at runs/critic-openai_gpt-5.6-terra.json (note (bco)). The critic was given
design.md including §4a and manifest.json.
One defect on the lead's side, first. The body came back finish_reason: "length" — the cap cut
it inside finding 7's amendment. Seven complete findings arrived and the verdict line arrived first,
so this is not a dead body and note (bhf) rule (iii) does not license a re-dispatch: the seat
answered. It was not re-dispatched, and findings 8+ are unknown. The cap was mine and it was too
small; note (abc) prices the worst case from max_tokens and says nothing about whether max_tokens
holds an answer, which is the same gap RS-20260811c §6 fell into from the other side.
Every finding is accepted. Nothing is overruled. The design as frozen claimed a confirmatory structure it did not have, and its primary measure did not measure what its question asked. The run below is what survives.
The findings and what was done
1 (BLOCKING) — the design is not pre-run for P1–P4. manifest.json, printed by
build_materials.py under gate G1, contains the word counts and every primary mark total for
every text. Those are sufficient to compute all four predictions before census.py exists. "Calling
the subsequent arithmetic 'unseen' does not preserve confirmatory status."
ACCEPTED IN FULL. P1, P2 and P4 are retired as confirmatory predictions and the Corpus A
census is reported as an explicitly exploratory, descriptive count. No P-value, no pass/fail
verdict, no "as predicted" anywhere in the result.
2 (BLOCKING) — A1 is an outcome-informed change to P3 with an incoherent criterion, and
absolute-count agreement is not a meaningful correspondence measure between texts of different
lengths. ACCEPTED. A1's pass/fail machinery is struck. What survives is the critic's own
amendment, which the run can satisfy exactly: "preserve and publish a timestamped/hash-identified
copy of the original D11 … evidence that it predates access to source and translation totals."
D11 is frozen in git commit 447e362, which predates build_materials.py by construction — the
script reads the file that commit created. So D11's content is verifiably pre-registered and its
scoring convention is not. Both scorings are reported side by side and no verdict is declared on
either.
3 (BLOCKING) — A2's zero rule is post-hoc, and "max ≥ 5 in the larger hand" is ambiguous
(larger rate? longer text? larger count?). ACCEPTED. A2 is struck entirely. Corpus A reports
raw counts and rates with no ratio statistic, so no zero-denominator rule is needed.
4 (BLOCKING), and it is the finding that redesigns the run — aggregate rates cannot answer
"licensed by the source". "A translation and source may have identical total exclamation, dash, or
ellipsis rates while every individual occurrence is moved, deleted, or newly introduced." The census
as frozen measures marginal typographic frequency and the question asks about licensing.
ACCEPTED IN FULL, and the primary measure is replaced with the one the critic specifies: aligned
units with a predeclared correspondence table — source mark retained at the aligned unit,
omitted, or target mark added with no source counterpart. It is run at paragraph grain,
not sentence grain, and only where alignment is established rather than assumed — which is
GAR_RU × LEAD, 32 paragraphs to 32 by construction. Corpus A is demoted to a secondary,
descriptive count, because its alignment is not established and this run will not assume it.
5 (BLOCKING) — P1 does not contain a source term and so cannot test its own sentence.
ACCEPTED. The sentence "the hands disagree at least as much as they disagree with the author" is
withdrawn and not replaced. The narrower claim the marginals do support is stated instead and is the
only thing Corpus A is used for: three independent hands of one source produce different amounts of
expressive typography, so the amount an English reader sees is not fixed by the source. That is an
existence argument from marginals and does not need alignment.
6 (BLOCKING) — dialogue density is not controlled, and the claim that P1 is "affected equally"
is false; also, quotation marks are excluded from the census yet used as an unvalidated dialogue
detector, and Russian direct speech is dash-led rather than quotation-delimited. ACCEPTED. G3
as written is struck. In its place, Corpus B's 32 aligned paragraphs are coded by hand for speech
status against the source's own dash-led and guillemet convention, the coding is published cell by
cell in alignment.json, and the correspondence table is reported stratified by it. For Corpus A
no dialogue claim is made at all.
7 (BLOCKING) — G2 does not detect abridgement; offsetting omissions and expansions can leave
word counts close, and rates over non-corresponding content are not comparable either.
ACCEPTED. G2's 25% rule is struck as a validity gate. Corpus A's word counts are reported as
what they are — a length statement, not a completeness audit — and the result page says in terms
that no completeness audit was performed on the three hands and that Field's possible
intermediary (G5) compounds rather than resolves it.
Findings 8+ — unknown, cap-truncated, not recovered.