Repository path: workshop/experiments/E-20260728j-classb-marking/design.md · rendered 2026-09-09
Page metadata (front matter)
| type | experiment |
|---|---|
| id | E-20260728j-classb-marking |
| status | frozen |
| created | 2026-07-28 |
| updated | 2026-07-28 |
| senses | accuracy, voice, style-correspondence, naturalness, consistency |
| internal-judgment-only | true |
| provisional | true |
| links | wiki/arms/ARM-longwork.md, wiki/findings/results/RS-20260727-jeli-fid.md, workshop/translations/jeli-il-pastore/R05-v1/span2-fid-prereg.md, workshop/translations/jeli-il-pastore/R05-v1/translation.md, workshop/regimes/R05-serial-long-work.md, wiki/base/consulted.md |
E-20260728j — is "English has no marked conditional to reach for" a fact about the pair, or about the translator?
Frozen before any forced rendering exists and before Dole 1896 is opened. This page is committed in its own commit; the translation commit follows it; the comparator read follows that. That ordering is the design.
ARM-longwork step 6, the arm's last step. The pair-level question this answers is the one
RS-20260727-jeli-fid §5 wrote down and deferred: "the pair-level question — what does a published
English «Jeli» do at these nine sites? — is answerable, because Dole 1896 is stored. It cannot be asked
until the work is finished, since reading the comparator would prime spans 3–5 and destroy the arm."
The work is finished (S047). The question can now be asked.
1. The claim under test
Log decision D22, span 2, frozen at 215434c, and repeated as the headline of
RS-20260727-jeli-fid §2:
The other three are conditionals —
saprebbe curarla,si sarebbe guarita,avrebbe saputo aiutarsi— and here the English would is the ordinary unmarked form, so the Italian's out-of-sequence marking simply does not appear: nothing is lost that a reader can see, and nothing is carried. I did not notice this until the sentences were written, and I have not repaired it, because there is nothing to repair with — English has no marked conditional to reach for. Six of nine carried; three of nine flattened with no available alternative.
Two clauses are doing different work and the page ran them together. "English has no marked conditional" is a claim about English grammar. "No available alternative" is a claim about what a translator can do at these three sites, which is a much larger claim and does not follow from the first: a marked conditional is not the only way to mark a clause as standing outside the narrative sequence.
RS-20260727-jeli-fid §5 already says the result "says nothing about C1" and is one translator's
choices. This design attacks it from two independent directions, neither of which needs a jury.
2. Materials — the nine Class B sites, frozen at S039
Enumerated in span2-fid-prereg.md §2 before span 2 was translated and unchanged since. Every
Italian string was machine-verified present in source-it-full.txt at S039 (note (g)) and is
re-verified by verify.py in this run.
| id | ¶ | Italian | mark | frozen English (S039, 215434c) |
|---|---|---|---|---|
| B1 | 28 | come fa il bottaio sui cerchi delle botti | gnomic pres. fa |
the way the cooper does on the hoops of a barrel |
| B2 | 41 | tremava come fanno le foglie in novembre | gnomic pres. fanno |
shook the way the leaves do in November |
| B3 | 41 | non rispondeva altro che con un guaito come fa un cagnuolo di latte | gnomic pres. fa |
answered nothing but a whimper, the way a suckling pup does |
| B4 | 42 | che anche un ragazzo saprebbe curarla | pres. conditional | that even a boy would know how to treat |
| B5 | 42 | se la febbre non era di quelle che ammazzano ad ogni modo | gnomic pres. ammazzano |
if the fever was not of the kind that kills you whatever you do |
| B6 | 42 | col solfato si sarebbe guarita subito | past conditional | with the sulphate it would have been cured at once |
| B7 | 42 | pensando che il ragazzo avrebbe saputo aiutarsi, quando fosse rimasto solo | cond. + impf. subj., tagged FID | thinking the boy would know how to fend for himself when he was left alone |
| B8 | 50 | colla curiosità inquieta che destano le cose spaventose | gnomic pres. destano |
with the uneasy curiosity that frightening things stir up |
| B9 | 51 | ed anche Mara doveva esser cresciuta, pensava egli sovente | modal, tagged FID | and Mara too must have grown up, he often thought |
Comparator: Nathan Haskell Dole, "Jeli, the Shepherd", in Under the Shadow of Etna (1896), Project Gutenberg #37979 — public domain, freely reachable, and the comparator every contamination gate on this work has used since S036.
3. What counts as MARKED — the definition, frozen before any rendering exists
This is the load-bearing part of the design, because the lead grades its own prose (§6). The definition is therefore mechanical: the presence of a listed form, not a judgment that a clause "feels" out-of-sequence.
An English clause rendering a Class B site is MARKED iff it contains at least one of:
- m1 — grammatical, tense. A finite verb in the present tense inside past narration, or a were-subjunctive or inversion (had the fever been…), where the propositional content does not require it.
- m2 — grammatical, modal. A modal other than plain
would/couldcarrying prospective-in-past or reported-reasoning force:was to,shouldin its older non-deontic use,must have. - m3 — lexical. An added evaluative or epistemic particle or adverb attributable to a speaker rather than to a neutral narrator, with no lexical counterpart in the Italian: right enough, no doubt, sure enough, of course, mind you, after all.
Otherwise UNMARKED. m1/m2 are recorded as grammatical markings, m3 as lexical.
The frozen S039 renderings are graded by the same definition, so the comparison is like with like.
P1 (§5 rules) still binds on every forced rendering: no added attributive frame — no as they said, no people thought, no quotation marks, no italics. A rendering that marks by adding a frame is recorded as FRAME, not MARKED, and counts as a failure of the site.
4. Procedure — three stages, each committed before the next begins
Stage 1 — the forced re-translation (the translation limb). From the Italian alone, without opening
Dole, I attempt at each of the nine sites a rendering that is MARKED under §3 and does not violate P1.
Each attempt is recorded whether it succeeds or fails, with the alternatives considered. Then the whole
of ¶42 is re-translated as continuous prose under the same brief — ¶42 holds four of the nine sites
(B4, B5, B6, B7) and is the only place where the accumulated effect of the policy on a continuous stretch
can be seen. Committed as forced/renderings.md, frozen.
Stage 2 — the comparator read. fetch_dole.py fetches PG #37979 and locates Dole's rendering of each
of the nine sites. Each is graded MARKED/UNMARKED/FRAME by §3, or NO-COUNTERPART where Dole omits or
paraphrases so freely that no clause corresponds. Committed as dole/sites.md.
Note (abm) — the standing rule that no comparator prose is displayed except a measured shared run — is
lifted for this stage, and the reason is on the record: that rule exists to protect unfrozen prose
from priming, and there is none. The work is translated whole and every span is frozen in git. Dole 1896
is public domain; §7 hygiene applies (brief attributed excerpts, logged in wiki/base/consulted.md).
Stage 3 — analysis. Three-way table (frozen lead / forced lead / Dole) at nine sites, predictions
scored, verify.py recomputes every count from the stored files.
Cost. $0 for all three stages: lead translation, a public-domain fetch, and arithmetic. One independent pre-run critic call is budgeted separately below.
5. Predictions, with failure conditions
P1 — pre-seen, and declared as such so the design's knowledge state is auditable. At least one of B4/B6/B7 admits a MARKED rendering, and at least one such rendering is grammatical (m1/m2), not lexical (m3) alone. Falsified if every forced attempt at the three conditional sites is UNMARKED or FRAME, or if the only successful markings are m3.
Pre-seen: while writing this page I noticed that B4 admits "that even a boy knows how to cure it" — a gnomic present, which is m1. Confirmation of P1 therefore carries almost nothing, and it is written down because a design that discovers what it already knew and does not say so is the failure mode this project keeps finding in its own pages. The weight is on P2–P6.
P2 — the real availability question. All three conditional sites admit a MARKED rendering.
Falsified if any of B4, B6, B7 admits none. B6 (si sarebbe guarita, a past conditional stating a
counterfactual outcome) and B7 (a future-in-past inside an explicit pensando che tag) are the two I
cannot call.
P3 — the cost, and it is a cost RS-20260727-jeli-fid did not name. Every grammatical marking
achieved at B4, B6 or B7 changes the proposition, not only the mood: it converts an episodic
statement (about this fever, this boy) into a generic one (about fevers, about boys). Falsified if
a MARKED grammatical rendering is produced at any of the three whose clause keeps the same referents as
the frozen rendering's and remains episodic rather than generic.
P4 — the pair-level test, and the one that can most embarrass the published page. Dole 1896 renders at least 4 of the 5 gnomic-present sites (B1, B2, B3, B5, B8) as MARKED, and marks none of the three conditionals. Falsified if Dole leaves 2 or more gnomic-present sites UNMARKED, or if Dole marks 1 or more conditional. If Dole flattens the gnomic presents, then "English admits a gnomic present" described a capability the frozen prose used and a published translator declined — which makes the six-carried result a policy effect, not a pair fact, and that is the finding.
P5 — convergence. At the three conditional sites, the forced renderings and Dole's do not coincide: no site where both are MARKED by the same class (m1/m2/m3) and the same construction. Falsified if they coincide at 1 or more sites. A coincidence would mean the repair is a resource of the pair rather than the lead's invention, which is the strongest possible answer to §5's question.
P6 — what the policy does to continuous prose. The forced re-translation of ¶42 requires 2 or more added lexical items with no counterpart in the Italian (m3 material, counted as tokens added), and they cluster at the conditional sites rather than the gnomic ones. Falsified if the forced ¶42 needs 1 or 0.
6. Failure criteria — what makes this run unreportable
- NO-COUNTERPART. If 4 or more of the nine sites have no corresponding clause in Dole, P4 and P5 are not reportable and are recorded as not run. Fewer than four: the affected sites are excluded and the exclusion count is stated with every figure.
- The lead grades its own prose, and that is this design's principal weakness. It is stated here
rather than discovered in the results. Two mitigations, both partial: the §3 definition is mechanical,
and every rendering — successful or failed — is quoted in full, so a later reader can regrade every
cell without re-running anything. No quality claim is made about any forced rendering. They are
probes of availability, not proposed improvements, and the frozen span-2 prose is not revised
(
R05is append-only; the translation page changes only by erratum). - What this cannot establish. Nothing here is evidence about C1 or about any language pair in general. Two translators at nine sites in one story is a case, and the honest statement of what a confirmed P4 would license is: one published English «Jeli» does X at these nine sites. Charter §5: the lead does not judge its own translation, and no judgment of quality is asked for or made.
- If
verify.pycannot reproduce a reported count from the stored files, the figure is withdrawn, not corrected.
7. What the arm takes from this either way
ARM-longwork step 6 is the arm's closing report, per goodness sense, of what the long form pressured
that short units left idle. This run feeds one line of it and the line reads differently under each
outcome, which is why it is worth running:
- P4 confirmed — the pair genuinely lacks the resource at the conditionals, D22's second clause was
right for a reason it did not give, and
style-correspondencehas a documented site class where a source marking is unrecoverable in this pair regardless of policy. - P4 falsified via the gnomic presents — D22's "six of nine carried" is a fact about the lead's policy and not about English, and every per-sense claim in the closing report that rests on "what English admits" is downgraded to "what this policy chose".
- P4 falsified via a marked conditional in Dole — D22 is wrong as written and the translation page owes an erratum.
8. Concurrency
The id E-20260728j-classb-marking was minted without any way to check whether another session had
claimed it (NEXT.md, standing hazard since S048). Same for method-note ids added at hand-off.
9. Budget
One independent pre-run critic call, panel role P1 (openai/gpt-5.6-terra), reserve
qwen/qwen3.7-max on an empty or length return (note (b): fall through, never retry the same slug).
No panel model is a subject in this design, so any of P1/P2/P3 could critique it; P1 is chosen for the
list-price record (note (x)) and terse output. Worst case built from max_tokens = 8000 at list
out-price plus the prompt at list in-price (note (abc)). Everything else in this design costs $0.
AMENDMENT, 2026-07-28, after the independent pre-run critic pass — verdict NEEDS-REDESIGN
Nothing above is edited. The design as frozen at 7968abe stands on the page; what follows amends
it, and where an amendment contradicts a sentence above, the amendment governs. Critic: P1
openai/gpt-5.6-terra, provider OpenAI, finish_reason: stop, $0.024405625, one call, no
fall-through. Ten findings. Nine accepted in substance, one accepted with a reasoned substitution, and
one clause declined in writing. The critic's response is stored verbatim at runs/critic.response.md.
Four of the amendments withdraw or narrow a claim this design had made, including the sentence its closing use rested on.
A1 — accepted. The MARKED definition was not mechanical, and §3 is replaced.
The critic quoted m1's "where the propositional content does not require it" and m3's "attributable to a speaker rather than to a neutral narrator" and is right: both are attribution judgments made by the lead after seeing its own rendering. §3 is replaced by the following, which is form-detection only.
m1 — grammatical, tense. A finite verb in the present tense, or a were-subjunctive, or a
conditional inversion (had the fever been…), occurring inside a clause that is not dash-set direct
speech, in a paragraph whose matrix narration is past. (The exclusion clause is gone. Direct speech is
identified by Verga's em-dashes and by V8's quotation marks, both mechanical.)
m2 — grammatical, modal. One of the exact strings was to, were to, should (non-deontic), or
must have, as the finite verb of the site's clause.
m3 — lexical, and the list is now closed. One or more of exactly: right enough · well enough · no doubt · sure enough · of course · mind you · after all · to be sure · naturally · indeed. Presence of a list member in the site's clause is m3; nothing else is.
FRAME is unchanged and is also mechanical: the presence of an attributive expression (as they said, people thought, so they reckoned) or of quotation marks or italics round the site.
A2 — accepted. P3 could not fail, and it is replaced by a textual test.
The critic: "the lead selects the forced wording and then decides whether it is 'episodic rather than generic'; it can make P3 true by choosing generic wording." Correct. P3 is restated on definiteness, which is on the page:
P3′. At every site among B4/B6/B7 where a grammatical (m1/m2) marking is achieved, the marked clause contains at least one noun phrase that is bare-plural or generic-indefinite where the frozen S039 rendering's corresponding phrase is definite or pronominal. Falsified if at any of the three a grammatical marking is achieved with every noun phrase's definiteness unchanged from the frozen rendering.
Definiteness is read off the determiner. Both clauses are quoted side by side for every site so a later reader regrades without re-running anything.
A3 — accepted, and it is the strongest finding. P4 conflated two things.
The critic: "Dole's present in 'leaves do' may be ordinary English simile idiom, not a deliberate FID/gnomic carry. The design has no 'source-mark correspondence' field or rule." Right, and the same objection lands on the frozen S039 result it is testing. A second, mechanical field is added and applied to all nine sites before Dole is opened:
ELECTIVE / FORCED. A site is ELECTIVE if the past-tense counterpart of the clause is grammatical English ("the way the leaves did in November"), FORCED if it is not. Only at an ELECTIVE site is a present tense a choice, and only there can it carry a marked/unmarked contrast; at a FORCED site the present is English idiom and carries nothing about the source's marking.
P4 is restated on the elective subset, and the denominators are recomputed and published before Dole is read.
A4 — accepted, and §7's first bullet is withdrawn.
The critic: "No prediction outcome is allowed to threaten the intended closing use … §7's 'the pair genuinely lacks the resource' and 'regardless of policy' generalize from Dole's choices to English–Italian pair capacity." Both points accepted.
- §7's first bullet is withdrawn as written. A confirmed P4 licenses exactly this and nothing wider: one published English «Jeli» declined to mark at the three conditionals, and so did the lead's policy. Two translators at three sites in one story is a case. No outcome of this design licenses any claim about English–Italian as a pair, and §6.3 is the limit for every outcome, not only for P4.
- A pre-committed retraction, so that at least one outcome costs this session something. If P4 is
falsified via the gnomic presents — Dole leaving two or more ELECTIVE gnomic sites unmarked — then
RS-20260727-jeli-fid§2's inference that the six carried because English admits the form is retracted, andARM-longwork's closing report may not use "six of nine carried" as evidence about anything but the lead's policy. That is the arm's own headline sentence on this feature, named in advance, with the condition that would remove it.
B1 — accepted. Freezing the prose was not enough.
The critic: "Dole can contaminate the lead's grading rules, forced-rendering alternatives and rationales,
… NO-COUNTERPART decisions, and the closing report." Stage 1's commit must therefore contain, frozen,
every lead judgment this design uses: the nine forced renderings and the ¶42 continuous rendering; the
MARKED/UNMARKED/FRAME grading of the forced renderings and of the frozen S039 renderings; the
ELECTIVE/FORCED classification of all nine sites; the P3′ definiteness table; and the P6′ count. Only then
is fetch_dole.py written and run.
B2 — accepted with a reasoned substitution, and one clause declined.
Accepted: discretionary NO-COUNTERPART let the lead rescue P4 by declaring an inconvenient marked conditional non-corresponding. It is replaced by an anchored procedure: for each site, a lexical anchor is named from the Italian before Dole is opened (e.g. B2's anchor is leaves + November); Dole's enclosing sentence is quoted in full; a site is AMBIGUOUS rather than excluded whenever the anchor is present but the clause structure does not correspond, and AMBIGUOUS counts against the lead's prediction — as not-MARKED for P4's "at least 4 of 5", and as possibly-marked for its "none of the three". Denominators are therefore fixed at 5 and 3 and no exclusion can move them, which answers C's "unstable denominator" and E's "unenforceable exclusion rule" together.
Declined, with the reason: the critic asked that uncertainty "count against P4/P5 reportability" — i.e. sink the run. It is declined because it creates exactly the incentive it is meant to remove: a rule under which ambiguity destroys the result gives the lead a motive to resolve ambiguities in whichever direction keeps the run alive. Counting AMBIGUOUS as evidence against the lead's own prediction removes that motive completely and is stricter in the only direction that matters. §6.1's four-site rule is withdrawn as superseded.
C — accepted. P6 is restated on a countable quantity.
The critic: "'added lexical items with no counterpart in the Italian' requires a frozen token/alignment policy", and "a result with two additions at gnomic sites satisfies the numeric condition while falsifying 'cluster at the conditional sites'." Both right.
P6′. The forced continuous rendering of ¶42 contains N ≥ 2 occurrences of members of the closed m3 list (§A1), and at least one of them falls inside a Class B conditional clause (B4, B6 or B7). Falsified if N ≤ 1, or if no m3 item occurs inside a conditional clause.
Counting a closed list of fixed strings needs no alignment.
D — accepted as a stated limitation.
"Independent arithmetic is not independent semantic verification." verify.py checks string presence,
counts, and that every reported cell is reproducible from the stored files; it cannot check that a
MARKED grading is correct. The mitigation is §6.2's and is unchanged: every cell, successful or failed,
is quoted in full. This design has no independent grader and does not claim one.
E — accepted. Costs and the reserve condition, stated.
- Worst case, now with the arithmetic. P1 at list $2.50 in / $15.00 out: prompt capped at 8,000 tokens
= $0.020, output capped at
max_tokens8,000 = $0.120, $0.140. Reserveqwen/qwen3.7-maxat its observed routed cost (S044–S046: $0.029–0.060), $0.060. Total cap $0.200. Actual: $0.024405625, 17% of the P1 worst case, no fall-through. - "Empty" is defined as
choices[0].message.contentstripping to zero length, orfinish_reason == "length"with such a content. The raw body is written toruns/critic__<slug>.rawbefore any parse, so a failed call is preserved; the runner already does this. - Stage ordering is procedural, not technical, and this design does not pretend otherwise. The
enforcement is the git history: this page frozen at
7968abe, Stage 1 at a hash recorded on the result page,fetch_dole.pywritten and run only after that. A reader who does not trust the ordering can check the commit dates and the file contents at each hash.