Repository path: workshop/canon/njala/manifest.md · rendered 2026-09-09
Page metadata (front matter)
| type | manifest |
|---|---|
| id | canon-njala |
| status | frozen |
| created | 2026-08-01 |
| updated | 2026-08-01 |
| links | workshop/canon/README.md, workshop/translations/njala/, wiki/base/consulted.md |
Brennu-Njáls saga, chapters 1–2 — source manifest
Anonymous, Brennu-Njáls saga, c. 1280. The longest and most widely read of the Icelandic family sagas.
- Language: Old Icelandic. The project's fourteenth source language and its first medieval
source in a Scandinavian language. Distinct from the Swedish of
canon-pelsenand the Norwegian ofworkshop/translations/karen/both in period and in morphology. - Extent: chapters 1 and 2 entire. 1,151 words, 6,348 characters, 43 paragraphs (ch. 1: 321 words / 11 paragraphs; ch. 2: 830 / 32).
- Orthography: MODERNISED, and this is a fact about the edition and not about the saga. The snerpa.is text prints maður, sonur, Mörður, eg where the medieval manuscripts and the normalised Old Norse editions have maðr, sonr, Mǫrðr, ek. Nothing in chapters 1–2 turns on the difference for a translator working into English, but a study of period marking on the source side could not use this text.
Provenance
https://www.snerpa.is/net/isl/njala.htm, fetched programmatically by fetch.py (this directory)
on 2026-08-01, session S080. HTML SHA-256
0baa39d2328ac0b57706e289dd05bb5de52c69f5005dfa87f829a41a023ac1d5; extracted text SHA-256
a8e205197dfa1efb6cf84b9bcd10ec1f6639ee2ac4e167b99ffc1b863ed44007
(source.meta.json carries both).
Public domain. The saga is thirteenth-century and the edition is an unannotated reading text.
No consulted.md entry is required for the source — nothing copyrighted was opened. The comparator
used for the contamination measurement (Dasent 1861) is separately out of copyright and is logged
with the experiment that opened it.
Why this passage
Proper-name density. Chapter 1 is very largely genealogy and chapter 2 is a betrothal
negotiation at the Althing: between them the two chapters name twenty-three people and fifteen
places in 1,151 words. That is what E-20260801d needs — tools/ngram_overlap.py::name_tokens
is a heuristic for deciding which tokens are proper names, it has never been checked against a
ground truth, and a text where a fifth of the content words are names is where the check has power.
The passage was chosen for that property before any translation was made and before any comparator was opened, and the choice is recorded here rather than in the experiment's design so that the ordering is visible.
What the fetcher does
fetch.py splits the single-file edition on its <b>N. kafli</b> chapter headings, strips tags,
unescapes entities, joins the edition's hard-wrapped lines inside paragraphs while keeping the
blank-line paragraph breaks, and asserts that no <, & or non-breaking space survives into the
frozen text. It takes --offline <cached.html> so the extraction can be re-run without the network.