Repository path: workshop/experiments/E-20260725-tierD-ladder/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260725-tierD-ladder-critic |
| status | active |
| created | 2026-07-25 |
| updated | 2026-07-25 |
| links | workshop/experiments/E-20260725-tierD-ladder/design.md |
Independent pre-run critic pass — Tier D perturbation ladder
Experiment discipline step 2. An independent adversarial critic (a separate agent instantiation, not the orchestrator that drafted the design) pressure-tested design.md v1 and the on-disk instrument before any API call. Verdict: NEEDS-REDESIGN, on twenty-one blockers.
This is the third design this project has put through an independent critic pass and the third to come back NEEDS-REDESIGN. Two of the blockers (B10, B11) are regressions: defects that a previous critic pass raised, that were accepted, and that this design reintroduced. That pattern is the most useful thing in the critique and is recorded as such at the bottom.
The critique measured the files rather than reasoning about them, and it is right about essentially everything it measured. Its single most consequential finding is D21: the sham arm — the control that makes every positive result mean anything — was measurably archaizing, and on the design's own §7.3 would probably have voided the entire run after the money was spent. Its second is E26: the pre-registered wave-2 gate arithmetically forbade wave 2 from ever running, in an experiment titled "the perturbation ladder".
design.md was rewritten as v2. A run may proceed only against v2.
Critique (verbatim)
Independent pre-run critic pass — E-20260725-tierD-ladder
I read
design.md(v1, mtime 07:25), the charter §§2–8,wiki/goodness-senses.md,config/models.md,config/budget.md,wiki/program.mdSlate A,T-berman-tendances-R04-v1, all three anchor directories, both prior critic passes, and the S014 raw run data. I also read the un-reviewed instrument that already exists on disk:materials/perturbations.md,materials/items.json,materials/build.py,materials/*.txt,tools/run_tierD.py,tools/verify_tierD.py.The materials moved under me during the review —
perturbations.md,items.json,build.pyand every variant.txtwere regenerated at 07:38:05, after I had measured them at 07:34. All numbers below are as of 07:38:05. That fact is itself finding B10.Everything numbered below marked "measured" was computed from the files, not inferred.
A. Materials
A1. Passage word counts in §3 are both wrong. LESSER. Measured: Shaw §I is 314 whitespace tokens, not 318. The lead's Genji §1‑3 is 325 words, not 380 (§1‑3 + §1‑4 together are 600). The materials are correctly cut —
A-source-ja.txt(805 chars) is exactly 蜘蛛の糸 §一 and ends at 「まっすぐにそれを御下しなさいました。」 matching Shaw's §I ending;B-source-classical-ja.txt(494 chars) is exactly 蓬生 §1‑3 and matchesB-ref-lead.txt. So this is a documentation defect, not a build defect — but it is the second time in three sessions that a design's stated file facts did not survive checking, and 380 vs 325 is a 17% error in the number that justifies passage comparability.A2. The Beowulf line range in §3 is wrong. LESSER. The design says "ll. 2015–2040". Every stored file —
beowulf-ingeld-oe.txt,gummere-1909.txt,kirtlan-1913.txt, the lead'sT-beowulf-ingeld-R04-v1— is ll. 2015–2070a (56 lines), and that is whatC-source-oe.txt(369 words) contains. Nothing is trimmed to 2040. Fix the design text; do not trim the files (a 2040 cut falls inside Kirtlan's chapter break XXIX and mid-sentence).A3. §3's justification for choosing Genji as the lead-provenance reference is false. LESSER. "Of the five lead translations available, the Genji passage is the only literary narrative prose into English whose source is stored beside it."
T-beowulf-ingeld-R04-v1is literary narrative, rendered as English prose, with its source stored in the same anchor directory — and this design uses that source in arm C.A-garnett-vankaalso supplies a stored published RU→EN reference that would have bought a third language pair for free. The choice of B may still be right; the stated reason is not a reason.A4. Provenance headers, footnote markers and chapter headings are correctly stripped. FINE — say so. Measured on the built files:
C-gummere.txtcontains zero{28a}–{28e}markers,C-kirtlan.txtcontains zero[55]–[59]markers and noXXIXheading,A-ref-shaw.txtcarries none of the=== PROVENANCE & LICENSE ===block that names Glenn W. Shaw, and no file carries a title line. This was the single largest available authorship leak and the build handles it. It is also the one thing the design document never mentions, which is why I checked it rather than assuming it.A5. Kirtlan's stored text is missing the rendering of half-line 2015a, and Gummere's is not. BLOCKER for arm C‑GK.
beowulf-ingeld-oe.txtline 2015 opensWeorod wæs on wynne. Gummere renders it ("The liegemen were lusty"); the lead renders it ("The company was in joy");kirtlan-1913.txt's own header records that the clause "sits at the end of the preceding sentence and is quoted in the anchor page rather than here" — soC-kirtlan.txtopens at 2015b ("Nor ever have I seen greater joy…"). A juror scoringaccuracyagainst the source sees the first clause of the source untranslated in one text and translated in the other. That is an extraction artefact of the anchor page presented to the jury as a translation difference.
B. Design and statistics
B6. Detection threshold 0.75 on n=12 has no stated statistical basis and cannot survive multiplicity. BLOCKER. 0.75 of 12 votes = 9/12. Exact one-sided binomial under p=0.5: P(X≥9) = 0.0730 — not significant at 0.05 even for a single cell. The firing rule is evaluated on six cells (4 operators heavy + 2 light). Under a global null with independent votes, P(at least one cell fires by chance) = 0.365. The design states no null, no CI, no correction, and no significance test anywhere. Compare 10/12 (0.833): P = 0.0193. The threshold is one vote away from a rate that would at least be nominally defensible, and the design's own §12 says "margins of ≤ 0.17 (2 votes) are reported as indistinguishable" — i.e. it already concedes the resolution is coarser than the threshold it set.
B7. The 12 votes are not 12 independent observations, and the design knows this and pools anyway. BLOCKER. The calibration-v1 critic raised exactly this (A4, "pseudo-replication from the two orderings"), it was accepted, and the S014 analyzer implements it —
analysis.jsonreportsn_units: 15, i.e. orderings averaged within juror→span before pooling. This design reverts to raw pooling of 12 (juror × ordering × passage) andverify_tierD.py'sagg()divides bylen(votes). The independent units here are 6 (juror × passage). At n=6, P(X≥5)=0.109 and P(X≥4)=0.344. Either average orderings first and state a rule on 6 units (e.g. "reference preferred in both orderings for ≥5 of 6 juror×passage units"), or justify the pooling. This is a process regression on an already-dispositioned point.B8. The specificity threshold of 0.75 scale points has zero empirical basis, and the project has never run this scoring format. BLOCKER. Every prior jury run in this repo is winner-per-sense, not numeric. I checked the S014 raw responses:
{"accuracy":{"winner":"1","confidence":"high","rationale":…}}. §10 calls S014 "an equivalent per-sense scoring task"; it is a different task. So the project has no data on how P1/P2/P5 use a 1–7 integer scale — no mean, no SD, no evidence they use more than 2–3 of the 7 points. A 0.75-point differential could be enormous or trivially attainable, and nothing in the design says which. This is the headline metric per charter §5.4. A 2-item, 2-ordering, 3-juror pilot costs ≈ $0.12 and would fix it. Not running it is the largest un-derisked element in the design.B9. Specificity's baseline includes
naturalness, so the design's own naturalness prediction mechanically inflates the margin it is supposed to be independent of. BLOCKER.verify_tierD.py:specificity()computesdrop[target] − mean(drop[s] for s ≠ target), and the "others" set includesnaturalness. The design predicts (§8.6) that naturalness rises for O2 and O3 — i.e.drop(naturalness) < 0— which lowers the baseline and inflates specificity for exactly the two operators where the design most wants a specificity result. The probe and the metric are not independent. Report specificity both with and without naturalness in the baseline, and pre-register which one fires the rule.B10. The instrument was authored after the design, is unfrozen, and changed twice during this critic pass. BLOCKER (repeat of a dispositioned defect). The anchor-verification critic's E6 — "the claims file, the actual instrument, is authored after the critic pass and is never frozen" — was accepted and fixed there by committing
claims.jsonbefore any call and recording its SHA in every run file. Hereperturbations.md,items.jsonandbuild.pyare untracked, were written after design.md, and were regenerated at 07:38:05 mid-review. At 07:34 the log declaredB-O4-heavysite 1 aselder brother → elder sister; the built file containedelder brother, and two edits present in the file (might be cut back → should be left to grow;put off → drawn/bleakness → richness) were absent from the log.verify_tierD.pydid not catch it: its check only tests that logged edits are present, never that present edits are logged. After the 07:38 rebuild the log check passes (92 edits, 0 problems), but nothing prevents the same drift recurring. Required: commit the materials in their own commit before any API call, record the items-file SHA in every run file, and add an unlogged-edit check (diff variant against reference, assert every hunk maps to a logged site).B11.
run_tierD.pyreuses cached responses keyed only on (item, order, juror). LESSER→BLOCKER given B10.if os.path.exists(stem + ".json"): spent += …; continue. There is no payload hash. Given thatitems.jsonhas already been rebuilt twice today, a re-run after a rebuild will silently score old responses against new texts. This is the anchor-verification critic's E7, accepted and fixed there, reintroduced here.B12.
parse_scoresrejects6.0and"6". LESSER, cheap. It requiresisinstance(v, int). A juror emitting6.5,6.0or"6"burns a retry at full cost, and §9 voids the run at >10% failures. Coerce integral floats and numeric strings before rejecting.
C. Confounds
C13. Length is genuinely well controlled — say so. FINE. Measured word counts: A variants 313–326 against a 314-word reference; B variants 316–322 against 325. No variant is identifiable by length. Given that O2 instantiates Berman's allongement, this was the obvious failure mode and it did not happen.
C14. O2 is identifiable without reading, by counting full stops. LESSER (intrinsic, but must be reported). Measured sentence counts: A reference 14 → A‑O2‑heavy 18 (+29%); B reference 10 → B‑O2‑heavy 18 (+80%). B‑O2 also drops all 3 semicolons and the single em-dash to zero. This is what the operator is, so it is not an artefact — but it means a pooled "worse" verdict on O2 is compatible with a jury that never read either text, and the specificity claim for
style-correspondencecannot distinguish reading from counting.C15. A‑O2‑heavy is the only A variant with 9 paragraphs; every other A text has 10. LESSER. Site 1 merges "It was morning in Paradise." into the preceding paragraph. A per-item formatting fingerprint. Restore the paragraph break or apply the same merge to the reference.
C16. O3 injects accuracy damage while the design claims "Meaning preserved". BLOCKER. Measured edits in the built O3 files: -
the golden pistils and stamens in their centers→the gold-colored centers— an omission (pistils, stamens). -as through a sterioptiscope→as if through a window— a different image. -In the eighth month of a year→In September one year— the eighth lunar month is not September, and the lead's own frozen log records rejecting "typhoon" for 野分 as inaccurate. -the herd-boys→the local kids— a different referent (総角 is a class-marked term for herd-boys).O3 and O4 are therefore not orthogonal, and the O3 cell cannot support a claim about
voicespecifically.C17.
A-O3-heavyswaps 1930 register for contemporary colloquial:"No, no, as small as this thing is, it, too, has a soul"→"Hold on now, small as it is, it's got a soul too". BLOCKER, and it interacts with C19. This is Berman's tendency 4 (vulgarisation), not 10/11. Dropped into Shaw's elevated-archaic 1930 prose, it is the sharpest possible register cue — and §6 predicts naturalness will hold flat or rise here.C18. O3 also removes all transliterated Japanese from the A variant. LESSER (intrinsic).
the Sanzu-no-Kawa and Hari-no-Yama→the river of the dead and the hill of needles. Presence/absence of romaji is a one-glance discriminator. Unavoidable given the operator, but it should be named as a limitation rather than left for a reader to find.
D. Controls
D19. The
naturalnessprobe is run without naming a register, in direct contradiction of the typology it cites. BLOCKER.wiki/goodness-senses.mdstates as a rule that "an evaluation must say period-idiomatic or contemporary-vernacular, not 'naturalness in general'".run_tierD.py'sSENSE_DEFS["naturalness"]says only "Reads as fluent, idiomatic English prose". S014 already measured what happens: naturalness was the only sense to reach 0.80 andconfig/models.mdrecords it as register-cued; the calibration-v1 critic's C1 was accepted on exactly this ground. The design calls this probe "the sharpest discrimination the design contains" and then runs it on the one sense with a documented register artefact, with the register unspecified, against a 1930 reference and a variant containing "Hold on now… it's got".D20. The naturalness prediction has no threshold and no consequence. BLOCKER (third occurrence of a dispositioned defect). §8.6 says "
naturalnessdoes not fall for O2 and O3" — what counts as falling, and what happens if it does, is nowhere stated. Calibration-v1 critic B3 ("cannot fail in any defined way") and anchor-verification critic E4 ("not a prediction, measured while deciding nothing — reintroducing it is a process regression") both flagged this shape and both were accepted. Give it an integer rule or drop it.D21. The sham is measurably directional, not quality-neutral. BLOCKER — and it is a global veto. Every A‑sham and B‑sham edit, measured by diff:
A‑sham B‑sham on→upon(×2 sites)no one→nobodyall→every oneover→onPresently→Before longrarely→seldomwhich→thatso much as→oncenoticed→caught sight ofeast and west gates→east and **the** west gateslooked→glancedaround→abouttook up X in his hand→took X up in his handso much as→evennobody→no oneFour of eight A‑sham sites push archaic-formal (
upontwice,every one) or clunkier (took the spider's thread up in his hand); one shifts sense (glanced≠looked). Three of eight B‑sham sites push archaic or clunkier (about the grounds,the east and the west gates,not even **on** some small matterafter "called on her"), and one replaces a phrase the lead's frozen log records as a deliberate register choice (never so much as occurred to him). The asymmetry is structural, not incidental: the reference is an unedited, considered text and any eight substitutions drift downward. My prediction is that the sham lands above 0.70 for the reference and voids the entire run under §7.3 — after the money is spent.D22. Even a perfectly neutral sham voids the run ~15% of the time. BLOCKER. The band [0.30, 0.70] on n=12 admits counts 4–8. Under a true null with independent votes, P(inside band) = 0.854, so P(spurious global void) = 0.146 — before adding juror clustering and the slot bias S010 measured at 0.63–0.85. A single control arm of 2 items is carrying a veto over 14 items. Either widen the band (counts 3–9), state it as an equivalence test with a stated tolerance, or give the sham more reps than any single targeted cell.
D23. The held-out arm compares 1909 alliterative verse against modern prose and calls it a same-quality control. BLOCKER. Measured on the built files:
words non-empty lines paragraphs archaic tokens C-gummere.txt403 56 (verse half-lines) 1 5 ( thou, thee, thy, canst, ere)C-kirtlan.txt489 2 2 17 ( goeth, boasteth, carrieth, dieth, lieth, walketh, cometh, escapeth, exhorteth, bringeth, shouldst, whilst…)C-lead.txt504 7 7 0 C‑GK is verse vs prose with a 3.4× archaism gap. C‑GL is 1909 verse vs 2026 plain prose with a 5-vs-0 archaism gap — and the lead's own frozen log states the plainness was chosen deliberately to be the plain pole of the study's scale (log §3a). §5 claims C‑GL "probes whether the jury systematically prefers published over lead provenance": no probe is needed, the two texts are separated by a century of English at first glance. Charter §5.3 and Slate A item 4 both specify the held-out arm as "an independent same-quality translation, chance expected"; the design silently relaxes this to "the margin should be markedly smaller", which is (a) not a pre-registered rule, (b) unlikely to hold, and (c) a change to a mandatory control that under charter §2.8/§8 should be opened as a decision page, not asserted inside the design that benefits from it. This arm costs 4 payloads ≈ $0.17 and, as built, can only produce an uninterpretable number.
D24. The dose ladder is confounded: light is a nested subset that retains the most salient edit. BLOCKER for the ladder's interpretation. §4 says only "Light = 3 edit sites. Heavy = 8 edit sites" — nothing about nesting or selection. Measured, every light set is a subset of its heavy set, and the subsets keep the loudest edits.
A-O4-lightretainsit, too, has a soul→it has no soul, producing "as small as this thing is, it has no soul: it would be rather a shame to recklessly kill it" — an internal self-contradiction detectable with the source closed, present at both doses.B-O4-lightretainsNot so much as a servant stayed on.→A few servants stayed on.. Predict: no dose effect for O4, for a reason that is a property of the site selection, not of the jury. State the selection rule and either match salience across doses or draw the light set at random from the heavy set.D25. O3's declared target contradicts the project's own frozen mapping, and the co-damaged sense is not scored. BLOCKER.
T-berman-tendances-R04-v1's mapping table — frozen before this design and cited by it as authoritative — assigns tendencies 10 and 11 tovoice,cultural-mediation. The design declares O3's target asvoicealone and scores a five-sense set that excludescultural-mediation. O3's most visible effect (removing Sanzu-no-Kawa, Hari-no-Yama, 禅師, 陸奥紙, 総角) is squarely cultural-mediation. Either score six senses or re-target O3 — as written the specificity test for O3 is measuring the wrong sense by the design's own catalogue.
E. Budget and operations
E26. The pre-registered wave-2 gate arithmetically forbids wave 2 from running, on the design's own point estimate. BLOCKER. §10: headroom $1.600774; wave 1 estimate $0.995; wave 2 estimate $0.332; gate = "wave 2 runs iff (headroom − wave 1 actual) ≥ 2 × wave 2 estimate" = $0.664. Maximum permissible wave-1 actual = 1.600774 − 0.664 = $0.936774 < $0.995. Wave 2 can run only if wave 1 comes in ≥6% under estimate — against a project record of S010 (est $1.9–3.0 → $4.49), S014 jury (est $0.90 → $1.83), S014 probe (est $0.25 → $0.40). The experiment is titled "the perturbation ladder" and its own budget rule makes the ladder unrunnable. Either restructure (run light before or interleaved with heavy), drop the held-out arm to buy headroom, or defer the dose axis explicitly rather than by arithmetic accident.
E27. The $0.04146 per-payload assumption is a mean transplanted from a smaller, different task. BLOCKER. The figures are exactly the S014 per-juror means (verified: P1 $0.01524, P2 $0.01609, P5 $0.01013). But (i) S014's task was winner-per-sense, this one is 10 integer scores + preference (B8); (ii) payloads are larger — measured source+two-translations: item A 4,307 chars, item B 3,970, item C 7,734–7,794, against S014's ≈2,000 with mean
prompt_tokens1,224; (iii) using S014's per-call maxima instead of means ($0.01738 + $0.02365 + $0.02364 = $0.06467), wave 1 alone = $1.552 — above the $1.45 abort and just under the $1.601 headroom, and all 32 payloads = $2.069. (iv) P5's measured mean cost $0.01013 is ~3× whatconfig/models.md's list prices predict from its token counts, so there is unmodeled billing in the one juror the estimate leans on for cheapness. Restate the estimate as a range with a worst case, asE-20260725-anchor-verificationdid (the only run so far to land inside its estimate).E28. Dispatch order puts the global gate last, so a budget abort destroys the run. BLOCKER, one-line fix.
run_tierD.pybuildspayloads = [(it, o, j) for it in items …]initems.jsonorder, which for wave 1 is: 8 heavy targeted → A-sham, B-sham → C‑GK, C‑GL. The pre-dispatch guardif spent + max(worst, 0.05) > cap: breaktherefore drops the sham and held-out arms first. Per §7.3, without sham data no sense may be declared calibrated whatever the other arms show — so an abort at 80% of budget yields a run worth exactly nothing. Reorder: sham → heavy → held-out. Better still, run the sham arm as a standalone $0.17 gate before committing to the rest (this also derisks D21 for two-thirds of a cent per call).E29. The abort cap ($1.45) sits above the point at which the run can complete, and below headroom. LESSER. With headroom $1.601 and a plausible wave-1 actual of $1.2–1.55, the cap will fire mid-wave-1 rather than cleanly between waves. Either set the cap at a value that guarantees a complete wave (and shrink the wave to fit) or make the guard arm-aware so it never truncates an arm.
E30.
--max-tokens 9000is undocumented in the design and is the S010 failure surface. LESSER. S014 used 10,000 and P5 averaged 3,079 completion tokens on a smaller payload. Truncation → parse failure → retry → double billing, and §9 voids the run at >10% failures. State the value in the design, state the retry budget, and state what happens to a truncated call's cost in the ledger.
F. Charter compliance
F31. Relaxing a charter-mandated control without a decision page. BLOCKER. Charter §5 step 3 and Slate A item 4 both make the held-out arm "an independent same-quality translation, chance expected". §5 of the design overrides this ("Not 'chance expected'") inside the experiment page. Charter §2.8: "Value-laden choices are opened as decision pages and ratified cross-session… A session never ratifies a decision it opened." Weakening the gate experiment's mandatory control is exactly such a choice.
F32. The
config/models.mdrevisit trigger is declared discharged in the wrong file, and on the wrong evidence. BLOCKER.config/models.mdcurrently reads "FIRING as of 2026-07-25 … Re-probe withtools/panel_probe.pybefore the next jury-heavy run — cheap, and Tier D depends on the instrument being current." The design declares the trigger's premise false and discharges it in §6 of an experiment page, leaving the config page saying FIRING; CLAUDE.md rule 6 makesconfig/models.mdthe sole configuration locus. Separately, the substitute evidence offered — that four of five were "exercised at length" in S015 — is a factual-adjudication task thatconfig/models.mditself records as one that "does not bear on calibration". Either editconfig/models.mdwith the reasoning and the evidence, or run the ~$0.05 probe.F33. The order-flip failure threshold is set just above the one juror already known to be unstable. BLOCKER. §9 down-weights any juror with flip rate > 0.50. Measured in
analysis.json: P1 0.194, P2 0.167, P3 0.167, P4 0.139, P5 0.444 — andconfig/models.mdrecords P5 as "down-weight until reps increase". P5 is one of three jurors here, i.e. a third of every cell, and the pre-registered rule cannot fire on it. Also, with 16 items each juror yields at most 16 order pairs, so the flip-rate estimate itself has a 95% interval roughly ±0.25 — a hard threshold on it is not measurable at this n. Fix: pre-commit a P5-excluded sensitivity analysis (the anchor-verification design's accepted disposition for the identical objection, C2), not a threshold P5 cannot cross.F34. Charter §5.4 makes cross-sense specificity the headline metric, and the design predicts its own headline will fail with no fallback decision. LESSER but worth stating. Prediction 3 says specificity fails on ≥2 of 4 senses, and S014's floor contrast already showed the "worse across the board" pattern (record-fit 0.98 / 1.00 / 0.87 / 0.77 / 0.80 / 0.92 on a blatantly worse translation). §7 names the outcome ("detects damage, not sense-calibrated") but never says what Tier D's gate status becomes if every sense lands there — which is the whole point of running it. Pre-commit that.
F35. Compliant where I checked. FINE. Order swap mandatory and implemented; dispatch strictly sequential (§6, charter §5); raw JSON preserved including
.attemptN.jsondiscards (the S015 fix); pre-dispatch worst-case cost guard (the S015 fix, anchor-verification G4); slugs resolved fromconfig/models.mdat run time, tools referring to roles P1/P2/P5; Berman read in French and correctly ledgered inwiki/base/consulted.md; freeze ordering verified (commit0111d38predatesdesign.md, whose mtime is 07:25); the "none of the twelve tendances targetsaccuracy" finding is real and correctly used to license O4 under Slate A item 2; the kimi exclusion figures are accurate (measured P4 mean $0.09089, i.e. 5.96× P1 and 8.97× P5); noprivate-texts/content is reproduced in the public materials; the absence of asenses:field on a page that evaluates the jury rather than a translation is defensible.
The single most important weakness
Three of the four things this run is meant to buy are unbuyable as designed. Detection is nearly a foregone conclusion — S014's floor contrast already showed 0.77–1.00 record-fit against a genuinely worse translation — so the detection leg mostly re-purchases something the repo has. Specificity, the charter's headline, is measured on a 1–7 scale that has never been piloted with these models, against a 0.75-point threshold with no empirical grounding, using a baseline that the design's own naturalness prediction mechanically inflates. The dose ladder cannot run, because the pre-registered wave-2 gate ($0.937 permissible against a $0.995 estimate) forbids it. And the one arm that would make any positive result mean something — the sham — is a measurably archaizing, clunkier edit set that is more likely to fail its own band than to sit in it, carries a ~15% spurious-void rate even if perfect, and is dispatched after everything it gates.
The cheap path forward, in order: (1) commit the materials and add an unlogged-edit check; (2) run the sham arm alone, 12 calls ≈ $0.17, as a gate — it simultaneously pilots the 1–7 scale (B8) and tests D21; (3) only if it lands in band, rebuild the firing rule on 6 independent units with a stated null, fix the naturalness prompt and rule, re-target or re-score O3, drop or replace the held-out arm, and re-plan the budget so the ladder can actually run.
VERDICT: NEEDS-REDESIGN
Disposition of every point
design.md v2 implements the critique's own recommended sequence. Twenty-one blockers and eight lesser points; 26 implemented, 3 accepted-and-recorded-as-limitations, 0 rejected.
| # | severity | disposition |
|---|---|---|
| A1 | LESSER | Fixed. v2 §3 states the measured counts: A 314 words, B 325 words. The critic is right that this is the second such defect in three sessions; the standing method note is extended in NEXT.md. |
| A2 | LESSER | Fixed. v2 records ll. 2015–2070a. Files untrimmed, per the critic's advice. |
| A3 | LESSER | Fixed. The false reason is struck. v2 gives the actual reason: the Genji reference is the only lead translation whose source is modern-readable prose narrative in a language the panel has been competence-screened on and whose passage is not also used by another arm. The Beowulf lead translation is excluded from the reference role because its source is used elsewhere and because the lead's own log records its diction as deliberately at one pole of a scale. |
| A4 | FINE | Recorded. v2 §3 now states the stripping explicitly and names what was stripped, instead of leaving it implicit in the build. |
| A5 | BLOCKER | Moot — arm C is dropped entirely (see D23). The half-line defect is recorded on A-beowulf-ingeld as a caveat on the stored Kirtlan file, where it belongs. |
| B6 | BLOCKER | Fixed. v2 §7 replaces the 0.75-of-12 threshold with a sign test on 6 independent units (juror × passage), unit = "prefers the reference in both orderings". Under the null a unit is +1 with p = 0.25, so ≥5 of 6 gives exact P = 0.00464 and family-wise P ≈ 0.018 across the four operator cells. The null, the exact P, and the multiplicity are all stated in the design. |
| B7 | BLOCKER | Fixed, and the regression is recorded. Orderings are averaged within (juror × passage) before anything is pooled; verify_tierD.py aggregates units, not votes. That this had already been raised, accepted and implemented in S014 and was then reverted is the most important single lesson of this critic pass. |
| B8 | BLOCKER | Fixed, by adopting the critic's recommended sequence. Stage 1 of the run is the sham arm alone, which is the scale pilot: it measures the range, mean and SD the three jurors actually use on a 1–7 scale before any threshold is applied to targeted arms. Stage 2 is dispatched only if stage 1's scale usage is wide enough to support a 0.75-point criterion, and the design pre-registers what "wide enough" means. |
| B9 | BLOCKER | Fixed. naturalness is excluded from the specificity baseline, pre-registered. Both versions are computed and reported; the rule fires on the excluding version only. |
| B10 | BLOCKER (regression) | Fixed. Materials are committed in their own commit before any API call; the SHA-256 of items.json is recorded in every raw response file by the runner; verify_tierD.py gains a diff-based unlogged-edit check that reconstructs the variant from the reference plus the logged edits and asserts byte equality — which catches edits present but unlogged, the direction the old check could not see. |
| B11 | BLOCKER | Fixed. The cache key includes the SHA-256 of the rendered payload; a cached file whose payload hash differs is ignored and re-dispatched. |
| B12 | LESSER | Fixed. parse_scores coerces integral floats and numeric strings before rejecting. |
| C13 | FINE | Recorded in v2 §12 with the measured numbers. |
| C14 | LESSER | Accepted as a limitation and stated. O2 is repunctuation; a jury could in principle score it without reading. v2 §12 states that the O2 cell cannot distinguish reading from counting, and the result page must repeat it. |
| C15 | LESSER | Fixed. The paragraph break is preserved; the O2 site now joins the sentences within the paragraph rather than deleting the break. |
| C16 | BLOCKER | Fixed. O3 is rebuilt so that every edit preserves the referent. The four meaning-changing edits the critic measured are gone: 八月 stays "the eighth month", 総角 becomes "the cowherds" (referent kept, class-marking lost), the pistils/stamens omission is dropped, the sterioptiscope becomes "a viewing-glass" rather than a window. |
| C17 | BLOCKER | Fixed. The vulgarising edit ("Hold on now… it's got a soul too") is removed. O3 is now strictly de-marking, not register-crashing. |
| C18 | LESSER | Accepted and stated. Removing transliterated Japanese is what tendency 10 is. v2 §12 names romaji presence/absence as a one-glance discriminator in the O3 cell. |
| D19 | BLOCKER | Fixed. The naturalness definition put to the jury now names its frame explicitly and instructs that a text is not to be penalised for being of an older or newer idiom, only for being unidiomatic within its own evident register. Verbatim wording is in v2 §6 and in tools/run_tierD.py. |
| D20 | BLOCKER (third occurrence) | Fixed. The naturalness prediction gets an integer rule and a consequence: for O2 and O3, drop(naturalness) ≤ 0.25 scale points; if it exceeds that, the pre-registered reading is "the jury lowers every sense together", and no specificity claim is made for that operator regardless of its margin. |
| D21 | BLOCKER | Fixed as far as it can be, and the residue is stated rather than denied. The archaising edits (upon ×2, every one, about the grounds, the east and the west gates) and the sense-shifting one (glanced) are removed; the sham is rebuilt from free variants that do not move register. But the critic's structural point is conceded: a perfectly neutral sham is not constructible against a considered text. v2 therefore reframes the sham as measuring an upper bound on the false-alarm rate, runs it first, and reports its direction as a result rather than assuming it away. |
| D22 | BLOCKER | Fixed. The sham no longer holds a global veto with a 15% spurious-void rate. v2 states the sham outcome as a three-way interpretation rule on 6 units, not a pass/void gate, and gives the sham the same number of units as any targeted cell. |
| D23 | BLOCKER | Accepted in full: arm C is dropped. The critic is right that verse-vs-prose with a 3.4× archaism gap cannot be a same-quality control. The project has no stored materials from which a genuine held-out arm can be built. Rather than relax a charter-mandated control (see F31), v2 declares the control not satisfiable from stored materials and, in consequence, that this run cannot pass Tier D at all — it is a pilot, and config/models.md stays NOT CALIBRATED whatever the arms show. Finding a genuine same-quality pair goes to NEXT.md as the named blocker on a real Tier D pass. |
| D24 | BLOCKER | Moot — the dose ladder is dropped (see E26), and the reason is stated rather than left to arithmetic. The nested-subset defect is recorded so a future ladder does not repeat it: draw the light set at random from the heavy set, or match salience. |
| D25 | BLOCKER | Fixed. Six senses are scored, cultural-mediation included, matching the frozen mapping. O3's target is declared as voice + cultural-mediation and specificity for O3 is computed against that pair. |
| E26 | BLOCKER | Fixed by dropping the dose axis outright, explicitly and in NEXT.md, rather than by an unrunnable gate. v2 has two stages, not two waves, and the stage-2 decision is made on stage-1's actual spend against a stated worst case. |
| E27 | BLOCKER | Fixed. v2 §10 gives a range with a worst case built from S014 maxima ($0.06467/payload), not means, and sizes the run so that the worst case fits inside headroom. |
| E28 | BLOCKER | Fixed. The sham is stage 1 and is dispatched first, standalone. A budget abort can now only truncate targeted arms, never the control. |
| E29 | LESSER | Fixed. The run is sized so the worst case completes under the cap; the cap is stated per stage. |
| E30 | LESSER | Fixed. max_tokens (10,000, matching S014), the one-retry budget, and the ledger treatment of truncated calls are all stated in v2 §10. |
| F31 | BLOCKER | Fixed by not relaxing the control. v2 does not weaken the held-out requirement; it declares it unsatisfiable, drops the arm, and declines to claim a Tier D pass. No decision page is needed to decline a claim. |
| F32 | BLOCKER | Fixed. config/models.md is edited directly with the reasoning and the evidence, and the critic's second point is honoured: the S015 factual-adjudication evidence is not offered as competence evidence, because that file itself says it does not bear on calibration. What is claimed is only what was checked — slug liveness, pricing, and release recency from the OpenRouter model list. |
| F33 | BLOCKER | Fixed. The unmeasurable flip-rate threshold is replaced by a pre-committed P5-excluded sensitivity analysis, the disposition already accepted for the identical objection in E-20260725-anchor-verification (C2). Flip rates are reported, not thresholded. |
| F34 | LESSER | Fixed. v2 §7 pre-commits the gate status for every outcome, including the all-cells-fail-specificity case. |
| F35 | FINE | Recorded. Note that the commit hash cited there, 0111d38, changed to 0aeeb93 when the branch was rebased to reset commit authorship; the freeze ordering the critic verified is unaffected (the tree is identical and the freeze commit still precedes design.md). |
The pattern worth keeping
Three of the blockers — B7, B10, B20/D20 — are defects that an earlier critic pass in this project raised, that were accepted, and that this design reintroduced. The project's critic passes are working; its memory of them is not. Neither workshop/experiments/README.md nor continue-prompt.md carries forward the specific dispositions that previous passes established, so each new design re-derives them from scratch and sometimes gets them wrong. That is a process defect, and the fix is a standing checklist rather than a better designer. It goes to NEXT.md.