aboutsummaryrefslogtreecommitdiffhomepage
path: root/src/sisudoc
Commit message (Collapse)AuthorAgeFilesLines
* html & epub header fixes, miscellanyRalph Amissah8 days1-132/+85
| | | | | | | | - adjust metadata, cleanup fixes - no <h0> realign title on h1 along with bespoke css - show group bullets (assisted by Claude-Code)
* abstraction 2.1: a poem states its object rangeRalph Amissah9 days9-23/+106
| | | | | | | | | | | | | | | - A poem is a container of verse, and stores its range of verse ocn, - its verse are the citable units with ocn. - A note in the last verse of a poem previously was not gathered into the endnotes section, this now is fixed Format 2.0 -> 2.1: the property is an addition, the poem is not a citable object (but contans a range of objects), its verse are (individual citable objects), and every reader checks the major version. A 2.0 database reads as having no ranges. (assisted by Claude-Code)
* search database: show all included notesRalph Amissah9 days1-3/+9
| | | | | | | | | | | | | | | | search database: show notes other than regular numbered (foot)notes. make sure all types of notes for all object types set the objects has.inline_notes_reg flag required for them to be rendered search results' html (doc_objects.body) show notes only where an object's has.inline_notes_reg flag is set, which occurred for numbered notes only - only numbered notes, regular footnotes were included - and notes in some objects were omitted (quotes and verse) (assisted by Claude-Code)
* notes: editor's notes, ~[* ]~ & ~[+ ]~, as seriesRalph Amissah9 days6-16/+84
| | | | | | | | | | The markup documents ~[* note ]~ and ~[+ note ]~ as editor's notes, each a separately numbered series, and a bare ~[ note ]~ which sisu put in the asterisk series. They are included as notes. Each series is numbered through the document, *1, *2 ... and +1, +2 ..., apart from the author's notes. (assisted by Claude-Code)
* unused function removedRalph Amissah9 days2-27/+11
| | | | | | | | - meta_processing_general() in spine.d replaced by per-stage predicates. - doc_matters.generated_time() commented out, unused since the odt and epub clock stamps were removed. (assisted by Claude-Code)
* --version, as an optionRalph Amissah9 days1-0/+13
| | | | | | --version prints the: version, compiler, and platform, and exits. (assisted by Claude-Code)
* options: request absent flag, warn; --strict failsRalph Amissah9 days1-0/+13
| | | | | | | | | | unrecognised options are named on stderr before any work is done, and --strict makes such fatal (exit 1). (prior to this commit a mistyped or retired flag was ignored without acknowledgement). This applies to options only, an argument that is not a source, such as a README beside pods being processed, is still passed over quietly. (assisted by Claude-Code)
* markdown output, --markdown (--md)Ralph Amissah11 days7-6/+693
| | | | | | | | | | | | | | | | | | | | | A plain-text output that keeps the object numbers, written as CommonMark with GFM pipe tables, one .md per document-language in <lang>/markdown/, (linked from the metadata page) with images into that shared directory itself, skipping any already there, so --markdown alone still gives a complete tree. provided as an alternative to --text rather than a replacement --markdown produces a document that renders, (both provide objext numbering survives both. Every object carries an anchor and a superscript number linking to itself, so a citation by object number is a link in any markdown renderer. --text may still be of value to read in a terminal A note is written where it is referenced, with a link back to the object it came from, rather than gathered into an endnotes section. (assisted by Claude-Code)
* typst output --typist & --pdf that makes the pdfRalph Amissah11 days6-5/+947
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | typst provides a new pdf build path replacing latex as the primary pdf builder. --typst (--typ) writes the .typ; --pdf writes the .typ and compiles it by spawning `typst compile`. (--pdf is redefined, being typst to pdf rather than latex). --latex is unchanged and writes the .tex; One .typ per document-language (not one per paper size). The paper and the orientation are read from sys.inputs with a default, so the same file compiles to every paper --set-papersize asks for, the pdfs keep their original names. Spine's paper names (derived from what suited latex) and typst's are different vocabularies and are mapped; a paper typst does not have is named and skipped. With the flag --pdf spine invokes a typesetter to generate the pdf output directly, which is new (the latex path writes a .tex and leaves xelatex to the caller). If typst is not available on the machine it is reported absent, naming the binary and printing the command, and the run continues, the .typ is written. Several useful features of an object-centric pdf are easily met, usually being a few lines each. - The object number goes in the margin as one `place` at a coordinate inside the object's own block, carrying the label every citation points at. - A note is a real footnote at the point of reference, with the document's own number and one link back to the object. - The table of contents locates by object number rather than by page, so it holds true across every paper and every setting. That typst compiles silently across the sample set means every internal link in every document resolves (--pdf over the collection with two paper sizes gives 144 pdfs, none skipped and none failed, each the page size its name claims and each tagged). paths: .typ, pdf named in same way as .tex based built pdfs so site linkage works either way flake.nix: added a typst dev shell, for .typ pdf typesetter - dsh-typst-pdf providing a pdf typsetter for spine's .typ. - dsh-latex-pdf provides xelatex for .tex which can still be written. (assisted by Claude-Code)
* flags: --ocda-db and matching --ocda-db-round-tripRalph Amissah11 days1-6/+15
| | | | | | | | | | | | | | | | flags names tidied so that each names the artefact it acts on. --ocda-db and --abstraction-db do the same thing. --ocda-db is the official flag (--abstraction-db remains but undocumented) --db-round-trip becomes --ocda-db-round-trip, so that it agrees with --ssp-round-trip: each names the artefact it reads rather than one naming a container and the other a format. --abstraction-source is untouched named after he stage as it takes .sst, a pod, .ssp or .ocda.db, rather than any one artefact. (assisted by Claude-Code)
* ocda: warn if heading claims reserved segment nameRalph Amissah2026-09-234-10/+281
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | spine automatically builds some segments, including: (toc, endnotes, glossary, bibliography, bookindex, blurb, _the_title) these are now identified as reserved names and a user is now warned if any of these names have been manually assigned to a heading by markup. A document still builds but to disambiguate the ocn of the heading is attached to the markup (reserved) name and this is seeded before a document not after. It is read off the finished abstraction rather than reported by the parser, and beside the ocn alignment check for the same reason: it is a statement about a document rather than a step in building one. WARNING reserved segment name: the_autonomous_contract... [en] heading 1 at ocn 135 asks for "endnotes", which spine gives its own generated section it is named "endnotes-135" instead; ... A warning: the document is correct and complete and the name it ends up with works. --strict makes it a failure for the run, as it does for ocn alignment, and by the same reasoning: the outputs are written and can be looked at, and the exit status is taken at the end. Two of the thirty-six sample documents have reserved segment names, "1~endnotes" heading. Output is unchanged: nothing here touches the abstraction. (assisted by Claude-Code)
* uid and paths separator "~" in place of ":"Ralph Amissah2026-09-237-37/+44
| | | | | | | | | | | | | | | | | | | need a character to split filenames on in certain circumstances. there problems with use of a colon in filenames, which is legal on posix but not on Windows."~" fits the bill better being legal on every filesystem of interest and is unreserved in rfc 3986, needing no escaping in a url. The reference abstraction is renamed, its content unchanged. Over the sample collection two filenames move and nothing else does: the abstraction and database digests are identical, the archive's members and their sizes are unchanged, and only the member order shifts, "~" collating after letters where ":" sorted before them. The document's epub dc:identifier is a v5 uuid derived from the uid, so it changes for those filenames. Every document's uid changes in the search database (spine.search.db). (assisted by Claude-Code)
* help: what spine takes, what an artefact carriesRalph Amissah2026-09-231-7/+63
| | | | | | | | | | | | | | | | --help now names the five source forms before the options, and after them states the contract: what a .ssp and a .ocda.db each carry and do not, that both record the digest of the markup they were built from, which five actions need that markup and what happens when they cannot have it, and the three ways to verify. And the version, 0.24.1 to 0.25.0, for abstraction format 2.0 and what it made possible: (one database per document holding every language of it, carrying the markup, images, configuration, manifest and translation catalogues it was built from, with the implications that carries) (assisted by Claude-Code)
* read: db names outputs from manifest carriedRalph Amissah2026-09-221-5/+39
| | | | | | | | | | | | | ensure that db uses the names it carries in manifest which produce deterministic output (and fix divergence in db rendering of output names, (which previously also looked for variable input from the environment)) Output built from a database with --config naming the site configuration is now byte identical to output built from the pod, on the render route as it already was on the materialise route. (assisted by Claude-Code)
* cleanup: misc. (previously deferred)Ralph Amissah2026-09-225-5/+24
| | | | | | | | | | | | | | | | | | | | | dr_document_make in three identifier spellings becomes document_make, which is what the file has been called since the rename. The mixin import audit, a template declares its imports at template scope and they are then visible in every scope it is mixed into (which is how a template-level split once hijacked a UFCS lookup in code that had not been touched). Disabling htmlSnippet's imports and rebuilding names the consumers that were relying on them rather than on their own: across twelve mixin sites, exactly one. metadata.d now imports the `to` it uses. The templates keep their imports, their own functions needing them, and narrowing those is a change of its own rather than a cleanup. And the standing FIX in source_pod.d: the insert digest line for a non pod source recorded the insert's path, where every other line in digests.txt is a filename and a digest naming a file that is not in the pod under that name cannot be checked against it. (assisted by Claude-Code)
* verify: this spine still produces this abstraction?Ralph Amissah2026-09-222-2/+296
| | | | | | | | | | | | | | | | | | | | | | | does this spine still produce this abstraction? A database states an abstraction and carries the markup it was built from, so re-parsing the one and comparing against the other says whether the two still agree. Checked as follows: - automatically and unskippably when a document is built from a database, at the moment it is parsed and before anything has been written from it. - on --ocda-verify=<file>, an action of its own that materialises, parses every language, compares, reports and writes nothing, the exit status being the answer. - on --no-verify as the escape, warning per document, for a newer spine reading an older artefact where the abstraction is expected to differ and the output is wanted anyway. A build ends at the first failure rather than skipping the language. A document rebuilt from a database now has its whole abstraction built. (assisted by Claude-Code)
* dispatch: pod materialised if source requiredRalph Amissah2026-09-221-24/+22
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | actions requiring source first materialise the pod One route per argument, decided by what was asked for. --source, --pod, --pod2, --show-abstraction and --ocda-db need original markup, so an .ocda.db given with any of them is written back out as a pod and parsed. Asked only to render, it is read as an abstraction as before. To ensure consistency, where a pod is materialized by an argument, that pod is used for the whole run, including for html and epub that could have been built from the loaded abstraction instead. --show-abstraction and --ocda-db join the markup side deliberately. They could be satisfied by re-serialising the abstraction already loaded, and were. But an artefact is a statement about the source: written from a loaded abstraction it says only that the loader is self consistent, where written from the markup it says what this spine (whatever current version) makes of that document today. The chain is checkable against itself: --ocda-db from a database now goes database, pod, parse, database, and returns the same 9,908,224 bytes it started from. --show-abstraction likewise re-emits the .ssp files byte for byte. A .ssp carries no markup by construction and a database written before the format carried it has none, so both are refused for those actions, once and by name, and the rest of the run goes on. (assisted by Claude-Code)
* dispatch: pod materialised if source requiredRalph Amissah2026-09-221-12/+57
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | actions requiring source first materialise the pod One route per argument, decided by what was asked for. --source, --pod, --pod2, --show-abstraction and --ocda-db need original markup, so an .ocda.db given with any of them is written back out as a pod and parsed. Asked only to render, it is read as an abstraction as before. To ensure consistency, where a pod is materialized by an argument, that pod is used for the whole run, including for html and epub that could have been built from the loaded abstraction instead. --show-abstraction and --ocda-db join the markup side deliberately. They could be satisfied by re-serialising the abstraction already loaded, and were. But an artefact is a statement about the source: written from a loaded abstraction it says only that the loader is self consistent, where written from the markup it says what this spine (whatever current version) makes of that document today. The chain is checkable against itself: --ocda-db from a database now goes database, pod, parse, database, and returns the same 9,908,224 bytes it started from. --show-abstraction likewise re-emits the .ssp files byte for byte. A .ssp carries no markup by construction and a database written before the format carried it has none, so both are refused for those actions, once and by name, and the rest of the run goes on. (assisted by Claude-Code)
* ocda db: carry catalogues & text blobs compressedRalph Amissah2026-09-223-22/+132
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The pod carries tools/po4a and the database (until now) did not, so a pod written back out of a database came back a lossy copy, without its translation catalogues. the database now carries them, walked and name-checked exactly as the pod writer walks and checks them, once on the last language. They are compressed to save space, the test sample carrying 6.25 MB of catalogue as read would have made that one document's database two thirds larger. Files gains a compression column: NULL or 'none' for the bytes as read, 'zstd' for a frame. What to compress is decided by role, not by size (source, conf, manifest and tools are text and are compressed; images being compressed already are not). The same kind of file is always stored the same way. Document objects are untouched and stay raw: objects_fts is external content over objects.text and reads that column directly. bytes and sha256 remain those of the original file, so every digest check, digests.txt line and source.digest rebuilt from these rows works unchanged, and a reader that does not decompress can still say what it is looking at. dbReadFiles decompresses, so no caller learns how a blob is stored. It asks pragma_table_info whether the column is there at all: a database written before it reads as raw. A row that claims zstd and will not decompress is named and left out, rather than handed back as a frame where markup should be. live-manual: 9,490,432 bytes without the catalogues, 9,908,224 with them. Output from the materialised pod is identical to output from the original pod, 679 files each side, no differences. (assisted by Claude-Code)
* ocda db: materialise a pod from carried sourceRalph Amissah2026-09-222-0/+284
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A database that carries its source can recreate the pod. --source and --pod2 given a .ocda.db now write the pod to a directory of the run's making and carry on with the pod, so everything after that point is handling a pod like any other. Nothing renders from the database. A materialiser that also rendered would be a second path to every output format and the two would drift; one that only writes files means the document is built by the same code over the same bytes as the original, and identical output is a consequence rather than an aspiration. Held against the original pod, site configuration constant: every output file identical across ten languages. The pod's name comes from the database's filename, which inverts the naming rule exactly: <doc>.ocda.db is named by doc_uid_out_no_lang, the pod name and the document's filename joined by ":" when they differ and the one name when they do not. So the materialised pod recomputes the uid it was named by and every output file lands on the name it had. Names are checked before anything is created, and one bad name refuses the artefact rather than skipping a file, as the zip reader does with a zip. Markup that does not match the digest stored with it is refused outright, where a mismatched image is written with a warning: a wrong image makes a document that looks wrong, a wrong markup file makes one that is wrong, in its text, with nothing downstream to notice. (assisted by Claude-Code)
* ocda db: now carry markup source, conf & manifestRalph Amissah2026-09-222-1/+86
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The database already carried the images, because without them no output can be produced. It now carries the markup as well, and with it the document's configuration and the manifest as the author wrote it. That is the difference between a serialised abstraction and a document source. An abstraction can be rendered but not re-parsed, and a reader who wants to correct a sentence needs the sentence as written. With these a pod can be written back out of the sqlite-file. Three new values in the existing role column, no schema change: the format stays 2.0. Source rows are named by their path within the pod, media/text/<lang>/<file>, and not by bare filename as images are. Every language of a document has a file of the same name, so bare names would collide under UNIQUE(role, name) and nine of ten would be dropped without a word. The path is also what a pod materialised from this database has to be told. Bytes stored as read with the digest over them, so that each row is checkable against the line source.digest was built from. Also build.spine_version, a file level row saying which spine wrote the file: not a property of the document, and the one thing a file cannot be asked for afterwards. The reader skips the build. prefix as it skips schema. and translation., or it would come back inside the document header and the round trip would differ. (assisted by Claude-Code)
* check: warn when document's languages ocn divergeRalph Amissah2026-09-223-1/+151
| | | | | | | | | | | | | | | | | | | The check runs once every language of a document has been abstracted. Providing a warning by default. --strict makes divergence a failure, taken at the end of the run so that outputs are complete and can be examined. (the name --strict is general, so that later checks can be added without a second flag). Under --parallel the profiles are appended under synchronized, and the comparison sorts a document's languages by name rather than taking them in the order the threads finished, so two runs print the same lines in the same order. The outcome is noted per language in the database as translation.ocn_aligned, for documents that have more than one language. (assisted by Claude-Code)
* check: languages of a document, against each otherRalph Amissah2026-09-221-0/+282
| | | | | | | | | | | | | | | | | | | | | | | | | | | Each language leaves a profile as it is abstracted: how many numbered objects it has, and what kind of object each ocn is. The comparison is against the language the manifest names first. A count difference is reported alone. Past the first dropped or added object every ocn names something else, and a kind comparison after it would print hundreds of lines that are all the one fault. The languages of a document share their object numbering, this being the basis of ocn citation in a multi-language document: an ocn names the same object in every language. Nothing enforces it. It holds where translations follow the source object for object, and stops holding otherwise (e.g. the moment a translator drops a paragraph, merges two, or turns a heading into a sentence, at which point the numbering no longer matches). An ocn is not unique: a poem and its first verse share one, by design. So the walk is in document order, the section order the .ssp is written in rather than the abstraction's key order, which an associative array does not promise, and the first object at an ocn is the one recorded. (assisted by Claude-Code)
* digests: db digest taken once on file completionRalph Amissah2026-09-222-34/+28
| | | | | | | | | | One file for the document (including all languages) with digest of the document. digests.txt lists it once, beside the shared images, taken when the last language has been written. digests.txt carries all digests. (assisted by Claude-Code)
* ocda.db: read a document language of a databaseRalph Amissah2026-09-224-34/+139
| | | | | | | | | | | | | | | | | | | | | | | | | A database now holds every language of its document, so reading one means naming a language. Three small pieces: - spineDbLanguages says what languages a file holds. Its own template rather than part of the reader (asking a file what it holds does not instantiate the object setter and the markup regexes with it). - abstractionLoad and spineDocFromArtefact take a language and pass it to dbReadFile, which resolves it to a doc_id, takes that language document where there is one, and lists the languages and stops where there is a choice. - spineArtefactLanguages turns an artefact into the documents to build: one for a .ssp, and for a database the languages it holds filtered by --lang. The artefact loop then has a language loop inside it, so a single .ocda.db argument builds every language, as the pod it came from would. --db-round-trip takes --lang for the same reason, and ocda-db looses --parallel (would cause a race). (assisted by Claude-Code)
* ocda db: one <doc>.ocda.db holding every languageRalph Amissah2026-09-221-10/+29
| | | | | | | | | | | | | | | | | | | The file is named for the document, not for one of its languages (and is written a language at a time: ten calls fill one file for a ten language document). The removal on entry was right when a call wrote the whole file. It now happens on the first language of the document, taken from the manifest rather than from whatever order the loop ran in, so nine languages are no longer written and thrown away. Also a partial unique index for the file level metadata rows. In SQL no null equals any other null, so UNIQUE(doc_id, key) does not constrain the rows written with a null doc_id, and schema.name and schema.version were inserted once per language: ten rows each, with nothing for INSERT OR REPLACE to replace. (assisted by Claude-Code)
* paths: doc_uid_out_no_lang, doc name sans languageRalph Amissah2026-09-221-0/+17
| | | | | | | | | | doc_uid_out ends in ".{lng}", which is right for an artefact written per language and wrong for one that holds every language of a document. Rather than strip the suffix at each place that needs the bare name, build it by the same branches without the language, so the two names cannot drift apart. (assisted by Claude-Code)
* ocda db: the schema gains doc_idRalph Amissah2026-09-222-40/+146
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | each file still holds one language, here groundwork for one database per document. The shape changes here and nothing merges yet: a database is still written per language, so every doc_id is 1. Real and testable on its own, where the writer and the reader together are not. documents, a row per language, is what lets one file hold a document's whole set. Everything that tells one language's rows from another's keys on documents.id. metadata is keyed on (doc_id, key) (no longer on key alone). Every language has a title and a creator, and a key-only primary key refuses the second one. schema.name and schema.version describe the file rather than a document in it, so they are written with a null doc_id and are the only rows that are. objects gain doc_id and its uniqueness widens from (section, seq) to (doc_id, section, seq). ('body', 0) exists once per language, so the narrow constraint was the thing that would have refused a second language outright. idx_objects_section leads with doc_id, or reading one language scans them all. objects.id stays a global INTEGER PRIMARY KEY, so object_images, object_links, object_anchors and object_subtoc keep their schema and their keys, and objects_fts keeps content_rowid='id'. The reader's four sweeps filter through objects rather than gaining a column of their own. outline and citable name the language and order by it first. A view over a file that may hold several languages and does not say which reads as one document and is several. The DDL is IF NOT EXISTS throughout, since a second language will open a file that already has its schema. dbReadFile takes an optional language and means "the only document in it" without one. Given none where there are several it reports the languages and stops, rather than returning the first: the round trip compares byte for byte, and a quietly wrong answer there would read as a spine fault. (assisted by Claude-Code)
* abstraction: format 2.0, source.digestsRalph Amissah2026-09-223-4/+65
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | format 2.0, source.digest covers every markup file, the sha256 over one line per markup file, "<sha256> <filename>", sorted by filename. source.digest was the sha256 of the master file as read, before any insert. This is right for a .sst, which is the whole document. For a .ssm with inserts .ssi (containing the substantive part of a documents text) this is close to meaningless. Considered change of meaning (rather than an addition) and given the major part of the format version with it: 1.1 becomes 2.0, in the .ssp header line and in the database's schema.version row together. A 1.x reader refuses a 2.0 artefact, which is what that check is for. - Sorted, so the value does not depend on the order a filesystem hands back a directory. - By filename (rather than by path), so it does not depend on where the pod sits, which is the property the pod rebuild comparison rests on: a pod unzipped elsewhere must still give the same abstraction. Within one language every markup file lives in one directory, so a filename identifies it. Each line is checkable on its own against a digests.txt line or a files row, which the single value was not. The whole is reproducible with sha256sum and sort, and was verified that way for a one file document and for a twenty file one. An unreadable file is named in the digest rather than skipped. The parse has failed elsewhere by then, and a digest that quietly left a file out would claim the document is something it is not. The reference abstractions are regenerated. Two lines change in each and no others: the version, and the digest. (assisted by Claude-Code)
* pod: generate the po4a configuration, --po4a-cfgRalph Amissah2026-09-223-7/+214
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A document's po4a configuration is derived rather than kept. Spine already knows both halves without looking: the manifest says which languages, and the markup says which files, through the << lines the parser follows. Derived, it cannot disagree with the document, where a hand-kept file eventually does: a language added to the manifest and forgotten in the Makefile is a translation nobody updates. It prints, rather than writing. Spine does not write into a source tree, and the file belongs beside the catalogues in the pod it describes: spine --po4a-cfg <pod> > <pod>/tools/po4a/po4a.cfg Three things the Makefile it replaces knew, which had to be found by running po4a rather than by reading: --keep 0 is the difference between regenerating a document and destroying it. po4a will not write a translated file below a completeness threshold, and that threshold defaults to 80%. A partly translated document is the normal state of one being worked on, and sisu covers it: an untranslated string falls back to the source text. Measured at the default, 181 of live-manual's 189 translated files would be discarded, and nine languages would collapse to eight surviving files, reading as the document reverting to its source language rather than as a failure. index.html.in is not markup and no << line names it, so a configuration built from the manifest and the insert list alone drops it, and with it index.html.in.pot and its nine .po files at the next regeneration. It is found by asking whether the source language has one. neverwrap is not set, and the measurement is recorded beside the code: setting it turns 889 of 1126 strings fuzzy in four languages, and po4a does not use a fuzzy translation, so Catalan's coverage would fall from about 46% to about 12%. Tidy formatting is not worth that. The run-complete banner is suppressed for this action. The configuration goes to stdout and the banner would otherwise be part of the file, which is what the round trips already avoid by leaving early; this one prints from inside output processing and cannot. (assisted by Claude-Code)
* pod: carry tools/po4a, the translation working setRalph Amissah2026-09-221-0/+41
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A pod now carries its catalogues, so a translation can be updated rather than only retyped: providing the .pot and .po (the means of maintaining it). A directory walk, where every other thing the pod writer carries is named. A catalogue set is whatever the translator has, and enumerating .pot and .po files would mean the writer deriving the document's languages a second time, from a second place, to say what it already reads from the manifest. Done once, on the last language, as the markup blocks are. The tree is not language specific, so a pass for each language would re-read all of it, 5.5 MB ten times over for live-manual, into an archive the next pass overwrites. Every name is checked before it is carried, by the rule the pod reader applies to an archive it is handed. The writer is reading a directory someone else may have filled, which is the reason the reader checks. The cost is small, and it is the whole argument for compressing the archive as a stream rather than per member: uncompressed content 6972418 -> 13167489 +89% the .sisupod 692942 -> 891388 +28.6% 5.5 MB of catalogues cost 198 KB, because the .po files largely repeat markup already in the archive and whole-stream compression sees it. Per-member deflate could not. (assisted by Claude-Code)
* pod: set entry limit for the writer and readerRalph Amissah2026-09-223-1/+77
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Both now take the number from one constant now set at 10,000, the writer checks it before the archive is written. Previously the number set for the pod reader was less than 500 members. The writer had no limit, (so spine could write a pod it would then refuse to read, reporting too many entries: a good file that looks corrupt, and only for a document with enough parts). 500 was set when a pod held markup and images. A pod carrying translation catalogues has a different arithmetic, and the count follows from how many files a document is made of times how many languages it has, not from how large the document is: the_wealth_of_networks 213,405 words, 1 file per language 12 languages -> 49 entries live-manual 24,724 words, 20 files per language 12 languages -> 506 entries The big book is not what runs into this; the modular manual is. live-manual stood two languages from an unreadable artefact. 10,000 covers a hundred-insert manual in thirty languages, about 6,200 entries, with room. It is deliberately well clear of any real document rather than snug above the largest one known: the two errors are not comparable, since too low refuses a legitimate document with a message that reads as corruption, while too high defers to a size cap a moment later. The count is the weakest of the three guards and is not what bounds resource use; the per entry and total size caps do that, both before a byte is written. Refusing also removes any archive an earlier language left. The writer runs once per language and only the last pass holds every language, so it is the last that goes over, and the passes before it wrote smaller archives that passed. Without this, refusal left a pod missing a language: an artefact that reads perfectly well and is wrong. A missing file is an error someone notices. The arithmetic is recorded beside the constant so the next person can re-derive the number rather than guess at it. (assisted by Claude-Code)
* pod: write <doc>.sisupod (single zstd frame)Ralph Amissah2026-09-224-9/+55
| | | | | | | | | | | | | | | | | | | | | | | | | | | | The pod archive spine writes is now named .sisupod and carries the same archive inside one zstd frame. The reader has accepted both since the previous commit, so every pod already published stays readable and this changes only what is written. free_culture 2863504 -> 1100180 live-manual 6972418 -> 692942 10.06x The wrap is at the pod call site and not in createZipFile, which also writes every epub and every odt. Best compression from whole-stream (rather than per-member compression). (for live-manual per-member deflate manages about 3x compared to 10x for whole pod content compression). The suffix is named as pod_archive_suffix. --pod-compression sets the level, 19 by default: a pod is written once and fetched many times (692942 bytes at 19, 837703 at 9, 2054365 at 1). (A value that is not a number warns and the default is used). Output built from a written .sisupod is byte identical to output built from the markup it came from. (assisted by Claude-Code)
* pod: read when zstd-wrappedRalph Amissah2026-09-226-10/+230
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | sisupod using zstd compression recognised by first bytes. <doc>.sisupod: spine reads one before it writes one, so when the name changes (from .zip to .ssiupod) every pod already published stays readable. A .sisupod is one zstd frame wrapping the archive spine already builds. The reader unwraps the bytes and hands the same archive to the same parser, so every guard downstream is untouched: entry names, per-entry and total size, path depth, escape and symlinks. Which container it is comes from the first four bytes (rather than the name). A pod published as a plain .zip reads as it always did, a .sisupod reads, and either one renamed reads too. The suffix is still recognised, both spellings, for the argument and for a url. libzstd is declared rather than bound: provides the whole surface of fifteen extern C prototypes (there is nothing to generate and no upstream tree to track, which is the arrangement sqlite3 already has). dub links it with "libs": [ "zstd" ]; nix needs zstd.out rather than zstd, whose default output is the binaries and carries no library at all. A frame declares its uncompressed size in its own header, and for a pod fetched over https that number is attacker controlled. The declared size is checked against a ceiling before a buffer is asked for, a frame that will not declare one is refused, and what comes out is checked against what was promised. The ceiling is the extraction limit the archive reader already applies, so the two bounds agree. Measured: output built from a .sisupod is byte identical to output built from the same pod's .zip, 359 files over text, html, epub, odt and .ssp for three documents, live-manual's ten languages included. A truncated frame is refused and the document skipped. free_culture 2863504 -> 1099736 2.60x the_wealth_of_networks 4294022 -> 1172649 3.66x live-manual 6972418 -> 675970 10.31x (assisted by Claude-Code)
* odt: office:meta from the document, not the runRalph Amissah2026-09-223-72/+88
| | | | | | | | | | | | | | | | | | | | | | | | | Three faults in four lines of the odt metadata. meta:creation-date and dc:date were taken from generated_time, the run's own clock, so two runs over unchanged markup produced different odt bytes. They now carry the document's own instant, so an odt is a function of its source. That was the last clock stamp in any writer: every output format is now reproducible. generated_time is a human-readable run stamp, of the form "2026-9-17 [38/3] 18:39:35", with unpadded numbers and an iso week in brackets. It is not a valid xsd:dateTime, which is what ODF asks for in both of those fields, so what was written there was malformed as well as unstable. dc:language now carries the document's own language tag. The instant is doc_matters.modified_utc, moved there from epub3.d in this commit: epub3 needs the same value for its dcterms:modified, and one definition cannot drift from the other. generated_time now has no caller and is left in place. (assisted by Claude-Code)
* metadata: use document's languageRalph Amissah2026-09-221-2/+6
| | | | | | | Every metadata page body element carried lang="en" xml:lang="en" as a literal. It now carries the document's own language tag. (assisted by Claude-Code)
* epub: dcterms:modified timestamp, reproducibilityRalph Amissah2026-09-221-3/+62
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | timestamp dcterms:modified from the document, not the clock EPUB3 requires exactly one dcterms:modified on the package It is now read in order: SOURCE_DATE_EPOCH where the environment sets it, which is the reproducible-builds convention and lets a build pin every artefact of one run; then the document's own date.modified; then date.published. The clock remains the last resort, for a document that says nothing about when it is from. A document date is loose markup. "1991", "2006-03" and "2021-12-00" all occur in the sample set, and a month or day of 00 is not a date. A missing or zero part becomes 01, and a day out of range for its month falls back to the first, so what is written is always a valid CCYY-MM-DDThh:mm:ssZ. Over the sample set two consecutive runs now produce all 36 epubs byte identical, and epubcheck reports no change: 0 fatal, 0 errors. (previously taken from the clock at build time. Two runs over unchanged markup therefore produced different epub bytes, so no comparison of epub output could tell a real change from the seconds between two builds. It also meant an epub was a function of when it was built rather than of what it was built from.) (assisted by Claude-Code)
* conf: read conf/document_make, drop the dr_ prefixRalph Amissah2026-09-223-8/+8
| | | | | | | Spine now looks for conf/document_make, and bundles that name. (the dr_ prefix was a sisu-era name to distinguish it) (assisted by Claude-Code)
* xml: markup language attributes take a bcp 47 tagRalph Amissah2026-09-223-10/+25
| | | | | | | | | | | | | | | | Spine names a language as the directory under media/text/ is named, which for a region is the posix form with an underscore, pt_BR. A lang, xml:lang or dc:language value is a BCP 47 (RFC 3066) tag and wants a hyphen, pt-BR. Spine put its own code there unchanged, so a document in a region language produced an invalid attribute in every one of its xhtml files: 138 epubcheck errors for a document of ten segments and a navigation file. src.language_tag is the same language as a tag. The attributes and the dc metadata now read it. Output paths keep src.language, since a path is named after the directory, not after the tag. (assisted by Claude-Code)
* pod digests: multilingual document digests sortedRalph Amissah2026-09-221-7/+16
| | | | | | | | | | | | | | Fixes for three faults in the insert (.ssi) digests. The key was the language captured from whichever path resolved the insert. It did not change across the loop over languages, so every language's inserts were filed under one language. The key is now the language being digested, as the digest of the primary file above it already did. - digest is taken after the check that the file exists. - digest is now taken inside the guard. - non-pod path now correctly keyed on the document's own language. (assisted by Claude-Code)
* downloads: honour --allow-downloadsRalph Amissah2026-09-221-1/+11
| | | | | | | | | The flag has existed since a url argument could be fetched and nothing ever read it, so every url was fetched whether or not the flag was given. A refused url is now dropped from the arguments, which is what a failed download already did. (assisted by Claude-Code)
* ocda db: before writing check carried file namesRalph Amissah2026-09-222-1/+230
| | | | | | | | | | | | | | | | | | | | bugfix: content image names are now checked in a pass of their own, before anything is created or written, and one bad name refuses every image in the artefact rather than skipping the one, which is what the zip reader does with a zip. The write then re-checks containment. (prior to this an image name read out of a .ocda.db was written without checks, and such a database can be fetched over https. A name could climb with "..", and being handed to chainPath an absolute name dropped everything before it, so "/x" did not land under the image directory at all. The zip reader has guarded against this since it was written; the database reader did not). carried_names.d holds both rules: a bare filename, as an image is carried, and a relative path, for the markup and conf a database is to carry. (assisted by Claude-Code)
* markup: keep a utf-8 bom out of the yaml headerRalph Amissah2026-09-211-1/+14
| | | | | | | | | | | | | | | | | | | A file beginning with a utf-8 byte order mark had the mark read as part of its first yaml key, so "title:" arrived as "title:" and every value under that key was dropped silently: the document kept its text and lost its title. Downstream that showed as an empty dc:title and an empty <title> element in the epub, which epubcheck reports as two errors. The mark is now kept out of the header and body split, and so out of the parse. It is not stripped when the file is read, because that text is what source.digest is taken over and the digest names the file as it sits on disk. One document in the sample set begins with a mark. It is left as it is: it is the case this guards against. (assisted by Claude-Code)
* images alt text, and epub says what it isRalph Amissah2026-09-147-31/+313
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Four things, not visible in an error count: - Images had no alt text, even where markup carried the text, and the field alt text is present, incorrectly used in code; fixed as the alt. { sm_tux.png 64x80 "Gnu/Linux - a better way" }image an image with nothing to say for itself gets alt="" - Images with no dimensions were incorrectly written width="0" height="0", fixed - Accessibility metadata, which EPUB Accessibility 1.1 requires: schema:accessMode, accessibilityFeature and accessibilityHazard, plus accessModeSufficient and accessibilitySummary, which epubcheck does not test for; used by - readers looking for a book it can use, - anyone distributing into the EU. Every value is derived from the document, not asserted: a document with no images says so, one whose images all carry alt text claims alternativeText and textual sufficiency, one where an image says nothing claims neither. Nothing claims WCAG conformance, which would be a claim about an evaluation that has not happened. - DPUB-ARIA roles beside the structure spine already knew about. doc-toc, doc-endnotes, doc-bibliography, doc-glossary and doc-index on the sections, doc-noteref on a note reference and doc-footnote on the note it points at. "doc_endnotes" as a class name means something to a stylesheet and nothing to a reading system. Not doc-biblioentry: DPUB- ARIA 1.1 deprecated it. These are plain ARIA, so the html gets them too. Also dealt with here: - Spine makes certain sections such as endnotes, bookindex, ... and a heading that has a name for itself cannot share the same name, if it does the name is now disambiguated, disambiguated. (disambiguation similar to that used for a repeated anchor tag) - Endnote anchors fixed to use "id" (being one unique tag in the document that is target of every reference to it)instead of obsolete form (they were <a name="note_1">. "name" on an <a>). Book index markers keep "name": the same marker sits on every object an entry names, 296 times in one document, and those are markers of where a term occurs rather than distinct targets. (assisted by Claude-Code)
* epub and html: no errors left over the sample setRalph Amissah2026-09-147-73/+318
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | epubcheck 5.3.0 over the 35 sample epubs went from 1302 remaining errors to none. The six causes for those remaining errors fixed here. most of the count was one of them: - Tables (1123 errors, 86%) where the writer emitted obsolete attributes, removed in 2014; table also now in div instead of a paragraph. also fixed: - Navigation (115). Every level 4 heading's nav entry ended "#0", the ocn of an object that has no ocn and so no id. - Duplicate ids (37). (from three unrelated sources). - The publication identifier (27 warnings). dc:identifier was a hardcoded hex string, not a UUID, and the same one in all 35 files, so two documents in one library collided. It is now UUID v5 over the document's own uid: a real RFC 4122 identifier, stable across builds and distinct per document. - Image manifest entries (15). media-type was "image/" plus the file extension, giving "image/jpg", which is not a media type; epubcheck reads it as a foreign resource and wants a fallback. There is now a mapping, which also covers svg and webp. Item ids came from the file basename, and "2bits_02_01-100.png" gave an id starting with a digit, which is not an XML name; ids are now prefixed and sanitised. - An ocn resolving to no segment (9). A poem block and its first verse share one ocn; the verses are written into a segment and the block itself is not. Also dropped from the package document: an xmlns:xsi that nothing used, and a prefix declaration for the rendition vocabulary, which is a reserved prefix and needed no declaring. The package now carries xml:lang. test-epub-validity.sh now allows no fatal and no error. Two NAV-011 warnings remain and are described in it. Three reference .ssp files change, all of them the anchor tag disambiguation above. (assisted by Claude-Code)
* regex: confine url within an endnote (delimiter)Ralph Amissah2026-09-141-4/+4
| | | | | | | in error regex permitted a url to go past an endnote close delimiter, fixed (assisted by Claude-Code)
* epub, html: a named anchor, and subtitle fixesRalph Amissah2026-09-142-16/+15
| | | | | | | | fixes related to: - named anchor id - subtitles, emit only where they exist (assisted by Claude-Code)
* epub, html: correctly identify code block as blockRalph Amissah2026-09-142-10/+11
| | | | | | | | - wrapper set to <div>, with class unchanged, and the six stylesheets select div.code where they selected p.code. - invented element <codeline> removed (assisted by Claude-Code)
* epub, html: reference issues fixesRalph Amissah2026-09-142-18/+10
| | | | | | | | - _part_eof.xhtml was declared in the manifest and never written. - image_sys/bullet_09.png was referenced as a background image, a throwback to the Ruby version of sisu, never used here. (assisted by Claude-Code)
* epub, html: document heads and ids, fixesRalph Amissah2026-09-142-6/+13
| | | | | | | One document head per file, and ids must be unique. No remaining fatal errors in the epub in the marup sample set. (assisted by Claude-Code)