Repository path: workshop/experiments/E-20260806-carrier-or-knowledge/critic.md · rendered 2026-09-09
Page metadata (front matter)
| type | note |
|---|---|
| id | E-20260806-critic |
| status | frozen |
| created | 2026-08-06 |
| updated | 2026-08-06 |
| links | workshop/experiments/E-20260806-carrier-or-knowledge/design.md, workshop/translations/zapiski-oct3/R19-v4/translation.md |
Stage 0 — the independent pre-run critic pass, and the six amendments it forced
Seat: nvidia/nemotron-3-ultra-550b-a55b (non-panel; the panel's five seats are all used
downstream and a critic that has seen the hypothesis may hold no rating role). Provider and full raw
body at runs/critic_nemotron-3-ultra-550b-a55b_try1.*. 186.5s, 6,663 chars, $0.0477954.
VERDICT: NEEDS-AMENDMENT — six findings, four BLOCKING. All six accepted; five in the form the
critic specified, one in part with the departure recorded. Every amendment below was applied
before any call after stage 0 was dispatched.
F1 (BLOCKING) — the primary does not control the 17.7× extent confound
Accepted. The design named the confound in §3 and left the statistic uncorrected. The critic's own
fix — a paired contrast (r(B) − r(N)) against (r(A) − r(N)) — is algebraically identical to
r(B) − r(A) and buys nothing, so it is not the amendment taken.
Amendment A1. A new failure criterion:
F7— the loud arm must beat the floor. Ifmean r(A) ≤ mean r(N),P1pmay not be read as support for anything.Awould then not be a manipulation that reached readers at all, only a rewrite, andBoutrunning it would beBoutrunning noise.
And P2p is formally demoted, in writing, before the data: a P2p firing may not be reported as
support for the entry's carrier list, because extent alone predicts it. It may be reported only as
not distinguishable from extent.
F2 (BLOCKING) — the parity screen was built blind to the only arm it needed to screen
Accepted in full, and this is the finding that changed the run most. The frozen stage-3 prompt
excluded "first-person statements about how the narrator regards or handles his own material" — which
is precisely B's three added propositions. A screen instructed not to look at the manipulation is
not a screen, and B was guaranteed to pass by construction.
Amendment A2. The stage-3 prompt now counts every propositional difference, with no exclusion category named to the seat, and sorts each into two neutral labels the seat is given without explanation of why:
EVENT— a difference in what the text says happened, or in what it states about any person other than the narrator.SELF— a difference in what the text states about the narrator's own thoughts, habits or handling of his material.
F2 now fires on EVENT > 2 for an arm. SELF is reported, not gated — and it is now a
check on the lead's own logs rather than a formality: B's log declares exactly three added
self-descriptive propositions and A/A′'s logs declare none. If the screen finds more, the
logs are wrong, and that goes in the result.
F3 (BLOCKING) — leakage into the held carriers, and too few seats to detect it
Accepted, one half in part. The critic read both logs correctly: A's log itself records that
obtaining for wheedling may have carried K6 with the register, and B's log records that its
site-7 fragments strain K2.
Amendment A3. The manipulation check goes from two seats to three, and F1 gets the
numerical thresholds the critic required, fixed here before the data:
A isolates iff mean|Δ_A(K1..K5)| - |Δ_A(K6)| >= 0.50
A' isolates iff mean|Δ_A'(K1..K5)| - |Δ_A'(K6)| >= 0.50
B isolates iff |Δ_B(K6)| - mean|Δ_B(K1..K5)| >= 0.50
where Δ_X(K) = profile_X(K) - profile_C(K), averaged over the manipulation-check seats. F1
fires — primary WITHHELD — if any of the three fails.
The part not accepted, with the reason. The critic required the manipulation check to pass
before stage 4 is dispatched. It is run before stage 4 and its result is known before the primary is
read, but stage 4 is dispatched regardless of the outcome. A run whose gate fires and which then
has no ratings at all cannot say what the gate's firing means; RS-20260805d §8's standing is that a
withheld primary whose value is hidden is worse than one whose value is printed. If F1 fires, every
stage-4 figure is descriptive and the primary is withheld — which is what withholding means here.
F4 (NON-BLOCKING) — six paired cells cannot support the claim
Accepted, both halves.
Amendment A4. Primary seats go from three to four — P1 openai/gpt-5.6-terra, P2
google/gemini-3.6-flash, P3 x-ai/grok-4.5, P5 deepseek/deepseek-v4-pro — giving 8 paired
cells and an exact sign-flip permutation over 2⁸ = 256. And the design now states in §7, before
the data: this run supports a pilot-level signal on one passage and licenses no generalisable claim
about voice.
None of the four primary seats holds any other role in this run. The critic, the paraphraser, the three manipulation-check seats and the two parity seats are all drawn from outside them.
F5 (NON-BLOCKING) — the screen prompt leaked the hypothesis by naming its exclusion
Accepted. Amendment A2 fixes F5 by the same edit: no exclusion category is named to the seat, and the two labels are given flat, with no indication which one the design cares about.
F6 (BLOCKING) — the two manipulations are not parallel in kind
Accepted — the deepest finding of the six, and the one that added a rendering. The critic is right
that A (70.1% of tokens rewritten, five distributed properties moved by wholesale substitution) and
B (4.0% rewritten, one property moved by three marking clauses) are not two instances of one
operation, and that the primary as frozen therefore compared heavy rewrite with light rewrite plus
explicit self-reflection.
Of the critic's three fixes: (a) — re-specify B without propositional addition — is refused, and
B's log gives the argument at length (self-knowledge has no substitutional exponent in English).
(c) — a normalised contrast — is adopted as a secondary only, because dividing by a 0.0396
denominator is not a statistic anyone should lean on.
(b) is adopted and executed: A′, T-zapiski-oct3-R19-v4, the five named carriers moved under a
hard budget of ten edited sites, drafted after this critic pass and before any further dispatch.
Amendment A5. The run becomes a four-arm extent ladder, and stage 4 gains the pair C–A′:
| arm | what moved | change fraction from C |
|---|---|---|
B |
K6 alone |
0.0396 |
A′ |
K1–K5, ten sites |
0.1767 |
A |
K1–K5, unbudgeted |
0.7012 |
N |
nothing, aimlessly | measured at stage 1 |
P1q, a second registered contrast: d′(s,o) = r(B,s,o) − r(A′,s,o) over the same 8 cells, same
bars as P1p. A′ is 4.5× B's extent rather than 17.7×, so P1q is the fairer of the two and is
reported beside P1p rather than instead of it.
Amendment A6 — what the primary is now licensed to say. The critic's underlying point is that this design cannot isolate a latent property; it can only compare operations a translator can actually perform. So the licensed sentence, fixed here before the data, is about the operations:
A translator who changed N% of a passage at the site of the narrator's self-relation moved source-blind readers' judgment of who was speaking further/less far than one who changed M% of it along all five properties the entry names.
Not the five carriers do not carry voice. A′'s log already establishes half of why the two
operations cannot be made parallel, from the translator's side and before any reader was asked:
K4 and K5 are located properties that ten edits move, K1 moves partially, and K2 and K3 are
distributed over every sentence and every concrete noun in the paragraph and cannot be moved by a
small edit at all.
Cost of the pass
$0.0477954 — 1.7% of the run's revised declared worst case, and it bought a fourth rendering, a rebuilt screen, two more primary seats, a new failure criterion, and a narrowed conclusion.