aboutsummaryrefslogtreecommitdiffhomepage
path: root/src/sisudoc/ocda
Commit message (Collapse)AuthorAgeFilesLines
* abstraction 2.1: a poem states its object rangeRalph Amissah9 days8-21/+100
| | | | | | | | | | | | | | | - A poem is a container of verse, and stores its range of verse ocn, - its verse are the citable units with ocn. - A note in the last verse of a poem previously was not gathered into the endnotes section, this now is fixed Format 2.0 -> 2.1: the property is an addition, the poem is not a citable object (but contans a range of objects), its verse are (individual citable objects), and every reader checks the major version. A 2.0 database reads as having no ranges. (assisted by Claude-Code)
* notes: editor's notes, ~[* ]~ & ~[+ ]~, as seriesRalph Amissah9 days2-9/+41
| | | | | | | | | | The markup documents ~[* note ]~ and ~[+ note ]~ as editor's notes, each a separately numbered series, and a bare ~[ note ]~ which sisu put in the asterisk series. They are included as notes. Each series is numbered through the document, *1, *2 ... and +1, +2 ..., apart from the author's notes. (assisted by Claude-Code)
* unused function removedRalph Amissah9 days1-11/+11
| | | | | | | | - meta_processing_general() in spine.d replaced by per-stage predicates. - doc_matters.generated_time() commented out, unused since the odt and epub clock stamps were removed. (assisted by Claude-Code)
* markdown output, --markdown (--md)Ralph Amissah11 days1-0/+5
| | | | | | | | | | | | | | | | | | | | | A plain-text output that keeps the object numbers, written as CommonMark with GFM pipe tables, one .md per document-language in <lang>/markdown/, (linked from the metadata page) with images into that shared directory itself, skipping any already there, so --markdown alone still gives a complete tree. provided as an alternative to --text rather than a replacement --markdown produces a document that renders, (both provide objext numbering survives both. Every object carries an anchor and a superscript number linking to itself, so a citation by object number is a link in any markdown renderer. --text may still be of value to read in a terminal A note is written where it is referenced, with a link back to the object it came from, rather than gathered into an endnotes section. (assisted by Claude-Code)
* typst output --typist & --pdf that makes the pdfRalph Amissah11 days1-0/+9
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | typst provides a new pdf build path replacing latex as the primary pdf builder. --typst (--typ) writes the .typ; --pdf writes the .typ and compiles it by spawning `typst compile`. (--pdf is redefined, being typst to pdf rather than latex). --latex is unchanged and writes the .tex; One .typ per document-language (not one per paper size). The paper and the orientation are read from sys.inputs with a default, so the same file compiles to every paper --set-papersize asks for, the pdfs keep their original names. Spine's paper names (derived from what suited latex) and typst's are different vocabularies and are mapped; a paper typst does not have is named and skipped. With the flag --pdf spine invokes a typesetter to generate the pdf output directly, which is new (the latex path writes a .tex and leaves xelatex to the caller). If typst is not available on the machine it is reported absent, naming the binary and printing the command, and the run continues, the .typ is written. Several useful features of an object-centric pdf are easily met, usually being a few lines each. - The object number goes in the margin as one `place` at a coordinate inside the object's own block, carrying the label every citation points at. - A note is a real footnote at the point of reference, with the document's own number and one link back to the object. - The table of contents locates by object number rather than by page, so it holds true across every paper and every setting. That typst compiles silently across the sample set means every internal link in every document resolves (--pdf over the collection with two paper sizes gives 144 pdfs, none skipped and none failed, each the page size its name claims and each tagged). paths: .typ, pdf named in same way as .tex based built pdfs so site linkage works either way flake.nix: added a typst dev shell, for .typ pdf typesetter - dsh-typst-pdf providing a pdf typsetter for spine's .typ. - dsh-latex-pdf provides xelatex for .tex which can still be written. (assisted by Claude-Code)
* ocda: warn if heading claims reserved segment nameRalph Amissah2026-09-233-8/+225
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | spine automatically builds some segments, including: (toc, endnotes, glossary, bibliography, bookindex, blurb, _the_title) these are now identified as reserved names and a user is now warned if any of these names have been manually assigned to a heading by markup. A document still builds but to disambiguate the ocn of the heading is attached to the markup (reserved) name and this is seeded before a document not after. It is read off the finished abstraction rather than reported by the parser, and beside the ocn alignment check for the same reason: it is a statement about a document rather than a step in building one. WARNING reserved segment name: the_autonomous_contract... [en] heading 1 at ocn 135 asks for "endnotes", which spine gives its own generated section it is named "endnotes-135" instead; ... A warning: the document is correct and complete and the name it ends up with works. --strict makes it a failure for the run, as it does for ocn alignment, and by the same reasoning: the outputs are written and can be looked at, and the exit status is taken at the end. Two of the thirty-six sample documents have reserved segment names, "1~endnotes" heading. Output is unchanged: nothing here touches the abstraction. (assisted by Claude-Code)
* uid and paths separator "~" in place of ":"Ralph Amissah2026-09-235-26/+33
| | | | | | | | | | | | | | | | | | | need a character to split filenames on in certain circumstances. there problems with use of a colon in filenames, which is legal on posix but not on Windows."~" fits the bill better being legal on every filesystem of interest and is unreserved in rfc 3986, needing no escaping in a url. The reference abstraction is renamed, its content unchanged. Over the sample collection two filenames move and nothing else does: the abstraction and database digests are identical, the archive's members and their sizes are unchanged, and only the member order shifts, "~" collating after letters where ":" sorted before them. The document's epub dc:identifier is a v5 uuid derived from the uid, so it changes for those filenames. Every document's uid changes in the search database (spine.search.db). (assisted by Claude-Code)
* read: db names outputs from manifest carriedRalph Amissah2026-09-221-5/+39
| | | | | | | | | | | | | ensure that db uses the names it carries in manifest which produce deterministic output (and fix divergence in db rendering of output names, (which previously also looked for variable input from the environment)) Output built from a database with --config naming the site configuration is now byte identical to output built from the pod, on the render route as it already was on the materialise route. (assisted by Claude-Code)
* cleanup: misc. (previously deferred)Ralph Amissah2026-09-222-2/+2
| | | | | | | | | | | | | | | | | | | | | dr_document_make in three identifier spellings becomes document_make, which is what the file has been called since the rename. The mixin import audit, a template declares its imports at template scope and they are then visible in every scope it is mixed into (which is how a template-level split once hijacked a UFCS lookup in code that had not been touched). Disabling htmlSnippet's imports and rebuilding names the consumers that were relying on them rather than on their own: across twelve mixin sites, exactly one. metadata.d now imports the `to` it uses. The templates keep their imports, their own functions needing them, and narrowing those is a change of its own rather than a cleanup. And the standing FIX in source_pod.d: the insert digest line for a non pod source recorded the insert's path, where every other line in digests.txt is a filename and a digest naming a file that is not in the pod under that name cannot be checked against it. (assisted by Claude-Code)
* verify: this spine still produces this abstraction?Ralph Amissah2026-09-221-0/+28
| | | | | | | | | | | | | | | | | | | | | | | does this spine still produce this abstraction? A database states an abstraction and carries the markup it was built from, so re-parsing the one and comparing against the other says whether the two still agree. Checked as follows: - automatically and unskippably when a document is built from a database, at the moment it is parsed and before anything has been written from it. - on --ocda-verify=<file>, an action of its own that materialises, parses every language, compares, reports and writes nothing, the exit status being the answer. - on --no-verify as the escape, warning per document, for a newer spine reading an older artefact where the abstraction is expected to differ and the output is wanted anyway. A build ends at the first failure rather than skipping the language. A document rebuilt from a database now has its whole abstraction built. (assisted by Claude-Code)
* ocda db: carry catalogues & text blobs compressedRalph Amissah2026-09-222-5/+44
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The pod carries tools/po4a and the database (until now) did not, so a pod written back out of a database came back a lossy copy, without its translation catalogues. the database now carries them, walked and name-checked exactly as the pod writer walks and checks them, once on the last language. They are compressed to save space, the test sample carrying 6.25 MB of catalogue as read would have made that one document's database two thirds larger. Files gains a compression column: NULL or 'none' for the bytes as read, 'zstd' for a frame. What to compress is decided by role, not by size (source, conf, manifest and tools are text and are compressed; images being compressed already are not). The same kind of file is always stored the same way. Document objects are untouched and stay raw: objects_fts is external content over objects.text and reads that column directly. bytes and sha256 remain those of the original file, so every digest check, digests.txt line and source.digest rebuilt from these rows works unchanged, and a reader that does not decompress can still say what it is looking at. dbReadFiles decompresses, so no caller learns how a blob is stored. It asks pragma_table_info whether the column is there at all: a database written before it reads as raw. A row that claims zstd and will not decompress is named and left out, rather than handed back as a frame where markup should be. live-manual: 9,490,432 bytes without the catalogues, 9,908,224 with them. Output from the materialised pod is identical to output from the original pod, 679 files each side, no differences. (assisted by Claude-Code)
* ocda db: materialise a pod from carried sourceRalph Amissah2026-09-221-0/+231
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A database that carries its source can recreate the pod. --source and --pod2 given a .ocda.db now write the pod to a directory of the run's making and carry on with the pod, so everything after that point is handling a pod like any other. Nothing renders from the database. A materialiser that also rendered would be a second path to every output format and the two would drift; one that only writes files means the document is built by the same code over the same bytes as the original, and identical output is a consequence rather than an aspiration. Held against the original pod, site configuration constant: every output file identical across ten languages. The pod's name comes from the database's filename, which inverts the naming rule exactly: <doc>.ocda.db is named by doc_uid_out_no_lang, the pod name and the document's filename joined by ":" when they differ and the one name when they do not. So the materialised pod recomputes the uid it was named by and every output file lands on the name it had. Names are checked before anything is created, and one bad name refuses the artefact rather than skipping a file, as the zip reader does with a zip. Markup that does not match the digest stored with it is refused outright, where a mismatched image is written with a warning: a wrong image makes a document that looks wrong, a wrong markup file makes one that is wrong, in its text, with nothing downstream to notice. (assisted by Claude-Code)
* ocda db: now carry markup source, conf & manifestRalph Amissah2026-09-221-1/+4
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The database already carried the images, because without them no output can be produced. It now carries the markup as well, and with it the document's configuration and the manifest as the author wrote it. That is the difference between a serialised abstraction and a document source. An abstraction can be rendered but not re-parsed, and a reader who wants to correct a sentence needs the sentence as written. With these a pod can be written back out of the sqlite-file. Three new values in the existing role column, no schema change: the format stays 2.0. Source rows are named by their path within the pod, media/text/<lang>/<file>, and not by bare filename as images are. Every language of a document has a file of the same name, so bare names would collide under UNIQUE(role, name) and nine of ten would be dropped without a word. The path is also what a pod materialised from this database has to be told. Bytes stored as read with the digest over them, so that each row is checkable against the line source.digest was built from. Also build.spine_version, a file level row saying which spine wrote the file: not a property of the document, and the one thing a file cannot be asked for afterwards. The reader skips the build. prefix as it skips schema. and translation., or it would come back inside the document header and the round trip would differ. (assisted by Claude-Code)
* check: warn when document's languages ocn divergeRalph Amissah2026-09-221-0/+10
| | | | | | | | | | | | | | | | | | | The check runs once every language of a document has been abstracted. Providing a warning by default. --strict makes divergence a failure, taken at the end of the run so that outputs are complete and can be examined. (the name --strict is general, so that later checks can be added without a second flag). Under --parallel the profiles are appended under synchronized, and the comparison sorts a document's languages by name rather than taking them in the order the threads finished, so two runs print the same lines in the same order. The outcome is noted per language in the database as translation.ocn_aligned, for documents that have more than one language. (assisted by Claude-Code)
* check: languages of a document, against each otherRalph Amissah2026-09-221-0/+282
| | | | | | | | | | | | | | | | | | | | | | | | | | | Each language leaves a profile as it is abstracted: how many numbered objects it has, and what kind of object each ocn is. The comparison is against the language the manifest names first. A count difference is reported alone. Past the first dropped or added object every ocn names something else, and a kind comparison after it would print hundreds of lines that are all the one fault. The languages of a document share their object numbering, this being the basis of ocn citation in a multi-language document: an ocn names the same object in every language. Nothing enforces it. It holds where translations follow the source object for object, and stops holding otherwise (e.g. the moment a translator drops a paragraph, merges two, or turns a heading into a sentence, at which point the numbering no longer matches). An ocn is not unique: a poem and its first verse share one, by design. So the walk is in document order, the section order the .ssp is written in rather than the abstraction's key order, which an associative array does not promise, and the first object at an ocn is the one recorded. (assisted by Claude-Code)
* ocda.db: read a document language of a databaseRalph Amissah2026-09-223-4/+67
| | | | | | | | | | | | | | | | | | | | | | | | | A database now holds every language of its document, so reading one means naming a language. Three small pieces: - spineDbLanguages says what languages a file holds. Its own template rather than part of the reader (asking a file what it holds does not instantiate the object setter and the markup regexes with it). - abstractionLoad and spineDocFromArtefact take a language and pass it to dbReadFile, which resolves it to a doc_id, takes that language document where there is one, and lists the languages and stops where there is a choice. - spineArtefactLanguages turns an artefact into the documents to build: one for a .ssp, and for a database the languages it holds filtered by --lang. The artefact loop then has a language loop inside it, so a single .ocda.db argument builds every language, as the pod it came from would. --db-round-trip takes --lang for the same reason, and ocda-db looses --parallel (would cause a race). (assisted by Claude-Code)
* paths: doc_uid_out_no_lang, doc name sans languageRalph Amissah2026-09-221-0/+17
| | | | | | | | | | doc_uid_out ends in ".{lng}", which is right for an artefact written per language and wrong for one that holds every language of a document. Rather than strip the suffix at each place that needs the bare name, build it by the same branches without the language, so the two names cannot drift apart. (assisted by Claude-Code)
* ocda db: the schema gains doc_idRalph Amissah2026-09-221-8/+35
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | each file still holds one language, here groundwork for one database per document. The shape changes here and nothing merges yet: a database is still written per language, so every doc_id is 1. Real and testable on its own, where the writer and the reader together are not. documents, a row per language, is what lets one file hold a document's whole set. Everything that tells one language's rows from another's keys on documents.id. metadata is keyed on (doc_id, key) (no longer on key alone). Every language has a title and a creator, and a key-only primary key refuses the second one. schema.name and schema.version describe the file rather than a document in it, so they are written with a null doc_id and are the only rows that are. objects gain doc_id and its uniqueness widens from (section, seq) to (doc_id, section, seq). ('body', 0) exists once per language, so the narrow constraint was the thing that would have refused a second language outright. idx_objects_section leads with doc_id, or reading one language scans them all. objects.id stays a global INTEGER PRIMARY KEY, so object_images, object_links, object_anchors and object_subtoc keep their schema and their keys, and objects_fts keeps content_rowid='id'. The reader's four sweeps filter through objects rather than gaining a column of their own. outline and citable name the language and order by it first. A view over a file that may hold several languages and does not say which reads as one document and is several. The DDL is IF NOT EXISTS throughout, since a second language will open a file that already has its schema. dbReadFile takes an optional language and means "the only document in it" without one. Given none where there are several it reports the languages and stops, rather than returning the first: the round trip compares byte for byte, and a quietly wrong answer there would read as a spine fault. (assisted by Claude-Code)
* abstraction: format 2.0, source.digestsRalph Amissah2026-09-223-4/+65
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | format 2.0, source.digest covers every markup file, the sha256 over one line per markup file, "<sha256> <filename>", sorted by filename. source.digest was the sha256 of the master file as read, before any insert. This is right for a .sst, which is the whole document. For a .ssm with inserts .ssi (containing the substantive part of a documents text) this is close to meaningless. Considered change of meaning (rather than an addition) and given the major part of the format version with it: 1.1 becomes 2.0, in the .ssp header line and in the database's schema.version row together. A 1.x reader refuses a 2.0 artefact, which is what that check is for. - Sorted, so the value does not depend on the order a filesystem hands back a directory. - By filename (rather than by path), so it does not depend on where the pod sits, which is the property the pod rebuild comparison rests on: a pod unzipped elsewhere must still give the same abstraction. Within one language every markup file lives in one directory, so a filename identifies it. Each line is checkable on its own against a digests.txt line or a files row, which the single value was not. The whole is reproducible with sha256sum and sort, and was verified that way for a one file document and for a twenty file one. An unreadable file is named in the digest rather than skipped. The parse has failed elsewhere by then, and a digest that quietly left a file out would claim the document is something it is not. The reference abstractions are regenerated. Two lines change in each and no others: the version, and the digest. (assisted by Claude-Code)
* pod: set entry limit for the writer and readerRalph Amissah2026-09-222-1/+39
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Both now take the number from one constant now set at 10,000, the writer checks it before the archive is written. Previously the number set for the pod reader was less than 500 members. The writer had no limit, (so spine could write a pod it would then refuse to read, reporting too many entries: a good file that looks corrupt, and only for a document with enough parts). 500 was set when a pod held markup and images. A pod carrying translation catalogues has a different arithmetic, and the count follows from how many files a document is made of times how many languages it has, not from how large the document is: the_wealth_of_networks 213,405 words, 1 file per language 12 languages -> 49 entries live-manual 24,724 words, 20 files per language 12 languages -> 506 entries The big book is not what runs into this; the modular manual is. live-manual stood two languages from an unreadable artefact. 10,000 covers a hundred-insert manual in thirty languages, about 6,200 entries, with room. It is deliberately well clear of any real document rather than snug above the largest one known: the two errors are not comparable, since too low refuses a legitimate document with a message that reads as corruption, while too high defers to a size cap a moment later. The count is the weakest of the three guards and is not what bounds resource use; the per entry and total size caps do that, both before a byte is written. Refusing also removes any archive an earlier language left. The writer runs once per language and only the last pass holds every language, so it is the last that goes over, and the passes before it wrote smaller archives that passed. Without this, refusal left a pod missing a language: an artefact that reads perfectly well and is wrong. A missing file is an error someone notices. The arithmetic is recorded beside the constant so the next person can re-derive the number rather than guess at it. (assisted by Claude-Code)
* pod: write <doc>.sisupod (single zstd frame)Ralph Amissah2026-09-221-1/+11
| | | | | | | | | | | | | | | | | | | | | | | | | | | | The pod archive spine writes is now named .sisupod and carries the same archive inside one zstd frame. The reader has accepted both since the previous commit, so every pod already published stays readable and this changes only what is written. free_culture 2863504 -> 1100180 live-manual 6972418 -> 692942 10.06x The wrap is at the pod call site and not in createZipFile, which also writes every epub and every odt. Best compression from whole-stream (rather than per-member compression). (for live-manual per-member deflate manages about 3x compared to 10x for whole pod content compression). The suffix is named as pod_archive_suffix. --pod-compression sets the level, 19 by default: a pod is written once and fetched many times (692942 bytes at 19, 837703 at 9, 2054365 at 1). (A value that is not a number warns and the default is used). Output built from a written .sisupod is byte identical to output built from the markup it came from. (assisted by Claude-Code)
* pod: read when zstd-wrappedRalph Amissah2026-09-225-8/+228
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | sisupod using zstd compression recognised by first bytes. <doc>.sisupod: spine reads one before it writes one, so when the name changes (from .zip to .ssiupod) every pod already published stays readable. A .sisupod is one zstd frame wrapping the archive spine already builds. The reader unwraps the bytes and hands the same archive to the same parser, so every guard downstream is untouched: entry names, per-entry and total size, path depth, escape and symlinks. Which container it is comes from the first four bytes (rather than the name). A pod published as a plain .zip reads as it always did, a .sisupod reads, and either one renamed reads too. The suffix is still recognised, both spellings, for the argument and for a url. libzstd is declared rather than bound: provides the whole surface of fifteen extern C prototypes (there is nothing to generate and no upstream tree to track, which is the arrangement sqlite3 already has). dub links it with "libs": [ "zstd" ]; nix needs zstd.out rather than zstd, whose default output is the binaries and carries no library at all. A frame declares its uncompressed size in its own header, and for a pod fetched over https that number is attacker controlled. The declared size is checked against a ceiling before a buffer is asked for, a frame that will not declare one is refused, and what comes out is checked against what was promised. The ceiling is the extraction limit the archive reader already applies, so the two bounds agree. Measured: output built from a .sisupod is byte identical to output built from the same pod's .zip, 359 files over text, html, epub, odt and .ssp for three documents, live-manual's ten languages included. A truncated frame is refused and the document skipped. free_culture 2863504 -> 1099736 2.60x the_wealth_of_networks 4294022 -> 1172649 3.66x live-manual 6972418 -> 675970 10.31x (assisted by Claude-Code)
* odt: office:meta from the document, not the runRalph Amissah2026-09-221-0/+72
| | | | | | | | | | | | | | | | | | | | | | | | | Three faults in four lines of the odt metadata. meta:creation-date and dc:date were taken from generated_time, the run's own clock, so two runs over unchanged markup produced different odt bytes. They now carry the document's own instant, so an odt is a function of its source. That was the last clock stamp in any writer: every output format is now reproducible. generated_time is a human-readable run stamp, of the form "2026-9-17 [38/3] 18:39:35", with unpadded numbers and an iso week in brackets. It is not a valid xsd:dateTime, which is what ODF asks for in both of those fields, so what was written there was malformed as well as unstable. dc:language now carries the document's own language tag. The instant is doc_matters.modified_utc, moved there from epub3.d in this commit: epub3 needs the same value for its dcterms:modified, and one definition cannot drift from the other. generated_time now has no caller and is left in place. (assisted by Claude-Code)
* conf: read conf/document_make, drop the dr_ prefixRalph Amissah2026-09-222-4/+4
| | | | | | | Spine now looks for conf/document_make, and bundles that name. (the dr_ prefix was a sisu-era name to distinguish it) (assisted by Claude-Code)
* xml: markup language attributes take a bcp 47 tagRalph Amissah2026-09-221-0/+15
| | | | | | | | | | | | | | | | Spine names a language as the directory under media/text/ is named, which for a region is the posix form with an underscore, pt_BR. A lang, xml:lang or dc:language value is a BCP 47 (RFC 3066) tag and wants a hyphen, pt-BR. Spine put its own code there unchanged, so a document in a region language produced an invalid attribute in every one of its xhtml files: 138 epubcheck errors for a document of ten segments and a navigation file. src.language_tag is the same language as a tag. The attributes and the dc metadata now read it. Output paths keep src.language, since a path is named after the directory, not after the tag. (assisted by Claude-Code)
* ocda db: before writing check carried file namesRalph Amissah2026-09-222-1/+230
| | | | | | | | | | | | | | | | | | | | bugfix: content image names are now checked in a pass of their own, before anything is created or written, and one bad name refuses every image in the artefact rather than skipping the one, which is what the zip reader does with a zip. The write then re-checks containment. (prior to this an image name read out of a .ocda.db was written without checks, and such a database can be fetched over https. A name could climb with "..", and being handed to chainPath an absolute name dropped everything before it, so "/x" did not land under the image directory at all. The zip reader has guarded against this since it was written; the database reader did not). carried_names.d holds both rules: a bare filename, as an image is carried, and a relative path, for the markup and conf a database is to carry. (assisted by Claude-Code)
* markup: keep a utf-8 bom out of the yaml headerRalph Amissah2026-09-211-1/+14
| | | | | | | | | | | | | | | | | | | A file beginning with a utf-8 byte order mark had the mark read as part of its first yaml key, so "title:" arrived as "title:" and every value under that key was dropped silently: the document kept its text and lost its title. Downstream that showed as an empty dc:title and an empty <title> element in the epub, which epubcheck reports as two errors. The mark is now kept out of the header and body split, and so out of the parse. It is not stripped when the file is read, because that text is what source.digest is taken over and the digest names the file as it sits on disk. One document in the sample set begins with a mark. It is left as it is: it is the case this guards against. (assisted by Claude-Code)
* images alt text, and epub says what it isRalph Amissah2026-09-143-2/+27
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Four things, not visible in an error count: - Images had no alt text, even where markup carried the text, and the field alt text is present, incorrectly used in code; fixed as the alt. { sm_tux.png 64x80 "Gnu/Linux - a better way" }image an image with nothing to say for itself gets alt="" - Images with no dimensions were incorrectly written width="0" height="0", fixed - Accessibility metadata, which EPUB Accessibility 1.1 requires: schema:accessMode, accessibilityFeature and accessibilityHazard, plus accessModeSufficient and accessibilitySummary, which epubcheck does not test for; used by - readers looking for a book it can use, - anyone distributing into the EU. Every value is derived from the document, not asserted: a document with no images says so, one whose images all carry alt text claims alternativeText and textual sufficiency, one where an image says nothing claims neither. Nothing claims WCAG conformance, which would be a claim about an evaluation that has not happened. - DPUB-ARIA roles beside the structure spine already knew about. doc-toc, doc-endnotes, doc-bibliography, doc-glossary and doc-index on the sections, doc-noteref on a note reference and doc-footnote on the note it points at. "doc_endnotes" as a class name means something to a stylesheet and nothing to a reading system. Not doc-biblioentry: DPUB- ARIA 1.1 deprecated it. These are plain ARIA, so the html gets them too. Also dealt with here: - Spine makes certain sections such as endnotes, bookindex, ... and a heading that has a name for itself cannot share the same name, if it does the name is now disambiguated, disambiguated. (disambiguation similar to that used for a repeated anchor tag) - Endnote anchors fixed to use "id" (being one unique tag in the document that is target of every reference to it)instead of obsolete form (they were <a name="note_1">. "name" on an <a>). Book index markers keep "name": the same marker sits on every object an entry names, 296 times in one document, and those are markers of where a term occurs rather than distinct targets. (assisted by Claude-Code)
* epub and html: no errors left over the sample setRalph Amissah2026-09-143-4/+39
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | epubcheck 5.3.0 over the 35 sample epubs went from 1302 remaining errors to none. The six causes for those remaining errors fixed here. most of the count was one of them: - Tables (1123 errors, 86%) where the writer emitted obsolete attributes, removed in 2014; table also now in div instead of a paragraph. also fixed: - Navigation (115). Every level 4 heading's nav entry ended "#0", the ocn of an object that has no ocn and so no id. - Duplicate ids (37). (from three unrelated sources). - The publication identifier (27 warnings). dc:identifier was a hardcoded hex string, not a UUID, and the same one in all 35 files, so two documents in one library collided. It is now UUID v5 over the document's own uid: a real RFC 4122 identifier, stable across builds and distinct per document. - Image manifest entries (15). media-type was "image/" plus the file extension, giving "image/jpg", which is not a media type; epubcheck reads it as a foreign resource and wants a fallback. There is now a mapping, which also covers svg and webp. Item ids came from the file basename, and "2bits_02_01-100.png" gave an id starting with a digit, which is not an XML name; ids are now prefixed and sanitised. - An ocn resolving to no segment (9). A poem block and its first verse share one ocn; the verses are written into a segment and the block itself is not. Also dropped from the package document: an xmlns:xsi that nothing used, and a prefix declaration for the rendition vocabulary, which is a reserved prefix and needed no declaring. The package now carries xml:lang. test-epub-validity.sh now allows no fatal and no error. Two NAV-011 warnings remain and are described in it. Three reference .ssp files change, all of them the anchor tag disambiguation above. (assisted by Claude-Code)
* regex: confine url within an endnote (delimiter)Ralph Amissah2026-09-141-4/+4
| | | | | | | in error regex permitted a url to go past an endnote close delimiter, fixed (assisted by Claude-Code)
* spine: a .ocda.db can be fetched, as a pod zip canRalph Amissah2026-09-122-26/+23
| | | | | | | | | | | | | | | | | A URL ending in .ocda.db is now downloaded and processed in place, through the path that already did it for .zip: the same curl call, the same size and timeout limits, the same refusal of local and private addresses, the same --allow-downloads guard, the same temp file cleanup. Only the pattern had to widen. rgx_url_zip ^https?://...[.]zip$ rgx_url_source ^https?://...([.]zip|[.]ocda[.]db)$ downloadZipUrl is no downloadSourceUrl and the download temp directory is spine-download rather than spine-zip-pod; (the extraction directory keeps its name, being still only for zips). (assisted by Claude-Code)
* ocda: the images an artefact carries, and the ones it only describesRalph Amissah2026-09-122-9/+149
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The document markup sample collection builds byte identically from either artefact: 1457 files, no differences, from markup, from the 35 .ssp files, and from the 35 .ocda.db files. A .ocda.db carries its images as blobs. dbReadFiles() reads them and they are written where the output writers look for images: a directory of this run's making, removed when the run ends, not the pod beside the artefact, (which could be a working published tree). The document's source path moves with it, image_dir_path being reached from the document's own file rather than from the pod. Each blob is checked against the sha256 the database recorded beside it. The two were written together, so a mismatch means the file is damaged. (the point of having recorded the digest). A .ssp only describes its images, so those are the ones in the pod it sits in, and they are checked against the digests it recorded. That asymmetry arises from of what the two artefacts are: the ocda.db is self-sufficient and can prove it is looking at the right image, the .ssp describes a pod it sits alongside. Found while wiring it: setting the pod directory alone was not enough, since image_dir_path is derived from the source file's path (../../image from media/text/<lang>/), so the source path has to be moved with it. (assisted by Claude-Code)
* spine: process a document from .ocda.db or .sspRalph Amissah2026-09-127-47/+426
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Process a document from its .ocda.db or .ssp abstraction. A .ssp or a .ocda.db given as an argument is now a document like any other: read rather than parsed, and handed to the output writers as the same doc they already take. The 35 document sample collection built from its .ssp files is byte identical to the same collection built from markup: 1457 files, no differences. Three things had to be rebuilt rather than read, all of them a value parsed rather than a structure derived: - classify_topic_register_arr and its expanded twin. The split is not a plain one, so the rule moved to sisudoc.ocda.meta.topic_register and both the yaml reader and this one call it. - creator_author_arr, which is creator.author split on the ", " it was joined with. - title_sub, a copy of title_subtitle made where the header is read, and what epub3 puts in dc:title id="subtitle". Fixed ordering bug found by the acceptance test. A heading's own anchor can be a bare number taken from its text and this can collide with the ocn of an unrelated object. Whole output comparison found that before the fix there was one wrong link in one epub's table of contents. From a .ocda.db, 1446 of 1457 files are identical. The 11 that are not are the images: a database carries its own image blobs and nothing yet extracts them, so the five sisu_markup images are not copied and the epub that embeds them differs. That is the next step and is not a defect in this one. --source and --pod2 are refused with a warning rather than half done, no artefact carrying the markup. (assisted by Claude-Code)
* ocda: one constructor for doc_mattersRalph Amissah2026-09-122-143/+244
| | | | | | | | | | | | | | | | | | | | ST_DocumentMatters was declared inside spineAbstraction(), closing over the parser's locals, so only the parser could build it. A document loaded from a .ssp or a .ocda.db has to arrive at the same value from what the artefact carries, and had nothing to build. It is moved to sisudoc.ocda.meta.doc_matters as docMattersMake(), taking the seven things a document is described by: the run (program_info, opt_action), where it is and what it is called (manifest), its header and the site config (conf_make_meta), its counts and indexes (ST_DocHas), its source digests, and the files its markup inserted. The parser fills them from what it has just parsed; the loader will fill them from the artefact. Pure refactor: the struct's members are unchanged, only where it is declared and where its inputs come from. (assisted by Claude-Code)
* ocda: heading cross reference linkingRalph Amissah2026-09-126-23/+72
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The segment a cross reference to a heading lands on, gixed upstream in ocda, where the field is created, rather than inferred by the loader. A heading above level 4 opens no html segment of its own, so a link to it has to land on the level 4 heading that follows. The build loop cannot know that when it reads the heading, and says so in a comment of its own: "for html segname need following lv4 not yet known". It back-fills tag_assoc when the level 4 heading arrives (lv0to3_tags), and the answer then lives only in that map. tags.segment_lv4_is now holds it, resolved in a pass over the finished head and body sections: walking backwards, the last level 4 heading seen is the next one for everything above it, which is the same answer the back-fill gives. Emitted in the .ssp only where it differs from .segment_html_is, which is every heading at level 4 or below, so it is sparse: 317 lines over the 35 document reference set, none removed. Carried in the database as a column of its own, and read back by both readers. docHasFromAbstraction reads it instead of working it out. With that, and on top of the two defect fixes, a document loaded from an artefact and one parsed from markup agree: key sets identical, all 35 documents values 3 entries differ of some 30,000, all of them _the_title, where the abstraction supplies an epub segment the parser leaves unset link targets all 1,651 agree, against 89 differing before any of this work and 1 after the defect fixes alone Format stays v1.1. That version is new in this same run of work and nothing outside spine has read it, so this belongs in it rather than in a bump of its own. (assisted by Claude-Code)
* ocda: ST_DocHas from a loaded abstractionRalph Amissah2026-09-122-0/+270
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | docHasFromAbstraction() builds ST_DocHas from the objects and the @doc_has block, so a document read from a .ssp or a .ocda.db can reach the value the parser reaches. Everything it builds is an index over properties the artefact already carries, not a recomputation of something it does not: the counts come from the header, imagelist from the .image records, the segment name lists and the tag associations from .anchor, .segment, .segment_epub, .heading_lev_anchor and .segment_*_is, and section_keys_sequenced from which sections are non-empty plus the run's own flags. *The two defect fixes this now sits on did most of the closing.* Measured before them and after, over some 30,000 tag_associations entries and the 1,651 keys a document actually links to: imagelist 8 documents differed, now none. The parser was the one that was wrong, and rgx.image being anchored to image markup brought it into line with what the .image records always said. key sets four keys differed (_part_eof, "0", the empty key, and "toc" the other way about), now none. _part_eof came back the moment @tail was carried, and the rest with it. values 89 link targets resolved differently, now 1. The one left is free_culture's ocn 5, and it is the case this cannot reach: a heading above level 4 takes its segment by back-filling when the next level 4 heading arrives, so the object never holds the answer and reading the objects in order only approximates it. Fixed properly in the commit that follows, in ocda where the field is made, rather than guessed at here. (assisted by Claude-Code)
* ocda: what a document has, as plain dataRalph Amissah2026-09-122-58/+63
| | | | | | | | | | | | | | | | | | | | ST_DocHas in metadoc_object_setter, filled by the parser, in place of DocHas_, a struct of accessors nested inside docAbstraction() that closed over the parser's locals. The member names are the ones the output writers already use, so nothing downstream changes. The point is that a struct closing over parser locals can only ever be built by the parser. A document loaded from a .ssp or a .ocda.db has to arrive at the same value with only what the artefact carries, and now there is something for it to fill. Two small things fall out. imagelist is a string[] rather than the lazy uniq range it was, and images() is therefore imagelist.length rather than counting the commas in the range's string representation, which carried a "TODO not ideal rethink". Same answer unless a filename contains a comma; the reference test agrees on all 35. (assisted by Claude-Code)
* ocda: reading .ocda.db now faster than parsing markupRalph Amissah2026-09-122-74/+111
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | reading a .ocda.db was slower than parsing markup source. Taking the largest markup document sample War and peace, 12,135 objects, optimised build, before this commit and after: parse markup 0.87 s 0.87 s load .ssp 0.047 s 0.047 s load .ocda.db 1.09 s 0.128 s The database goes from being 1.25x slower than parsing the document to 6.8x faster. Two things fixed in the reader were: First, four queries per object. object_images, object_links, object_anchors and object_subtoc were queried per object as the objects were built, each statement compiled fresh from a concatenated string. For war and peace that is 48,540 statement preparations to collect 138 rows, which is all those four tables hold between them. They are now four ordered sweeps, kept by object id, so the cost is what the tables hold rather than what the document holds. Second, and even more consequentially: d2sqlite3's row["name"] is indexForName, a linear scan over the statement's columns that calls sqlite3_column_name and allocates a D string for every column it passes. At some fifty named reads per object over forty columns that is around a thousand of those per object, twelve million for the document. The column name to index map is now resolved once per statement and the reads are an integer index. Also here, since it was measured while doing this: sspReadFile no longer hashes the file it read unless asked (with_digest). Only the database writer wants that digest and it has the lines already, so every other read was paying for it. Existing outputs unaffected. (assisted by Claude-Code)
* ocda: the abstraction carries the document metadata output readsRalph Amissah2026-09-123-15/+167
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | .ssp format v0.1 -> v1.1, and the .ocda.db with it: they are two serialisations of one format and now move together. The gap this closes, measured against the writers' read set: of the 59 conf_make_meta fields outputs/io_out/ reads, 29 were in neither artefact. Five are conf.* and stay out, being site and run scoped (urls, papersize, the search db filename): the same abstraction published to two sites must take each site's. Of the rest, three are never assigned anywhere (title_short, publisher, original_publisher; read only by sqlite.d, always empty) and one is a copy of a field already carried (title_sub = title_subtitle), so 17 properties actually had to travel and now do: make breaks, footer, home_button_text meta title.edition, date.added_to_site, language.document_char, original.{title,source,language,language_char}, rights.copyright_{text,translation,illustrations, photographs,cover,audio,video} make splits by when it acts, which is worth keeping in mind: these three are read in outputs/io_out/ and must travel, while italics, bold, emphasis, substitute and headings are read at parse time and their effect is already in the objects. New @source block, and source.* rows in the database, so a reader can say which markup an abstraction came from rather than working from stale content in silence: language the document's own languages the pod's list, which is what the inter-language links in html need and neither artefact carried digest sha256 of the .sst; equals its digests.txt entry The database adds source.ssp_digest, the sha256 of the .ssp it was built from, since a file cannot hold its own hash. So the chain .sst -> .ssp -> .ocda.db is checkable end to end. The database's metadata table is now filled from the header blocks the .ssp gives back, not from a second list read off doc_matters. One list, in ssp.d: a property added there arrives in the database with nothing else changed, and one that is not in the .ssp cannot be in the database at all. That was the last place the two could drift; the objects stopped being able to on 2026-09-07. Version is checked on load. The major part must match, a newer minor is accepted (a minor bump only adds properties, and an unknown property line is ignored). A v0.1 artefact is now refused with a message saying to regenerate it, rather than loading half populated. test-abstraction-db.sh now compares the two artefacts' header blocks property by property, 36 per document on the wealth of networks, where it previously only checked that a schema.version row existed. Verified non-vacuous: dropping one property is reported. Reference regenerated, and the whole diff is this change and nothing else: 320 lines added, 35 removed over 35 files, being 35 format lines changed, 35 @source blocks (4 lines each), 35 language.document_char, 35 home_button_text (it has a default), 18 breaks, 16 footer, 4 date.added_to_site, 1 title.edition, 1 original.source. Existing outputs unaffected. (assisted by Claude-Code)
* ocda: .ssp should have all sections of abstractionRalph Amissah2026-09-121-2/+4
| | | | | | | The abstraction has nine sections and includes "tail". Both artefact writers hardcoded a list of eight, so "tail" was silently dropped. (assisted by Claude-Code)
* ocda: get image from is internal image markupRalph Amissah2026-09-123-6/+18
| | | | | | | | | | correct circumstance where regex was incorrectly able to match image shaped words as well as identified and marked up image (internal markup). take all matches. (assisted by Claude-Code)
* ssp: abstraction directory cleared onceRalph Amissah2026-09-091-13/+33
| | | | | | | | | | | | | | | | | abstraction directory cleared once, by whichever language is first The .ssp writer clears stale files out of pod/<doc>/media/abstraction/, and that one directory is shared by every language of a document. The clearing is now done once per directory per run, by whichever language reaches it first, with the lock held across it. A language that finds the directory already prepared has passed through that same lock before writing, so the clearing it skipped had completed before its own write began: no .ssp produced on a given run can be removed during it. (removes possibility of a race condition on parallelisation) (assisted by Claude-Code)
* html metadata: a link to the ocda.dbRalph Amissah2026-09-091-0/+17
| | | | | | | The metadata page gains a line between the markup source and the source digests. (assisted by Claude-Code)
* --ocda-db replaces --show-abstraction-dbRalph Amissah2026-09-091-1/+1
| | | | | | | | | ocda (object centric document abstraction) flag --ocda-db (or --abstraction-db) replaces --show-abstraction-db rename results in consequently large diff
* output: the abstraction artefacts live with the podRalph Amissah2026-09-091-8/+25
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | pod/ holds all document source representations: pod/<doc>/ source tree pod/<doc>/media/abstraction/<uid>.ssp abstraction, as text pod/<doc>.zip tree, zipped, .ssp included pod/<doc>.digests.txt sha256s of what is in them pod/<uid>.ocda.db abstraction, as sqlite db <lang>/abstraction/ is gone. pod/<doc>/media/abstraction/<uid>.ssp preferred as having the images (found within the pod tree) which .ssp needs to reproduce a document but does not carry on its own. The .ocda.db sits carries the images as well and (like the pod.zip) can be used to reproduce a document directly. digests.txt now covers the database as well as the zip, the source and the .ssp; (as does the metadata html page). Two ordering issues addressed: - the pod builder clean-slates pod/<doc>/ before regenerating it, and the .ssp is now written before that runs. The clean slate now leaves media/abstraction/ alone, and the .ssp writer clears that directory itself on the first language of a run, so a .ssp for a language the document no longer is removed and cannot be bundled. - for a multi-language document the .ssp files accumulate one language at a time and are bundled on the last, which is why the directory cannot simply be emptied by whichever (language) gets there first. (assisted by Claude-Code)
* ocda loader: one way in whatever the sourceRalph Amissah2026-09-092-0/+206
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | sisudoc.ocda.abstraction.load names the five things a document can be read from, tells them apart, and loads the two that are self-describing artefacts: .sst / .ssm + images the markup source pod (dir) + images the same, bundled pod .zip the same, zipped .ssp + images the abstraction, as text .ocda.db the abstraction, sqlite, images inside abstractionSourceOf(path) is the detection, by name and for a directory by whether it holds pod.manifest. abstractionLoad(path) returns a LoadedAbstraction: the source kind, whether it was loaded, why not when it was not, and the document itself. The three source forms are deliberately not loaded here. Reading them is the parser's job (sisudoc.ocda.meta.metadoc spineAbstraction) and it needs the manifest, environment and configuration that spine.d assembles, none of which belongs in a loader. What this gives that case is the dispatch and a plain statement of where it is handled, rather than a silent empty result. spine --abstraction-source=<path> says what a path is and, for an artefact, loads it and reports what came back: the document, title and author, header block sizes, object counts and the objects in each section. Exit 0 when an abstraction was loaded, 1 when not. (assisted by Claude-Code)
* ocda db: <doc>.ocda.db, and a script to run 4 testsRalph Amissah2026-09-091-1/+1
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The per document database is now written as <doc_uid>.ocda.db rather than <doc_uid>.abstraction.db. Shorter, and it says what is in the file: the object centric document abstraction, not "some abstraction". .ocda.db pairs with .ssp and cannot be mistaken for the collection search database (spine.search.db). test/run-tests.sh runs the four in sequence, one line of result each, and a summary. It re-runs itself inside nix shell "nixpkgs#sqlite" if sqlite3 is not on PATH, so this is all that is needed: SpinePOD=../../markup/sisudoc-spine-samples/markup/pod-samples/pod \ ./test/run-tests.sh ./bin/spine-ldc The tests are independent (with test-abstraction-ssp.sh run first): 1 test-abstraction-ssp.sh is first because it is the one that says whether the abstraction itself moved; if it fails the others are answering a different question than you think 2 test-abstraction-ssp-roundtrip.sh reads the committed reference set, so it is a statement about the current binary only if 1 passes 3 test-abstraction-db.sh and 4 test-abstraction-db-roundtrip.sh generate both artefacts themselves and depend on nothing committed (assisted by Claude-Code)
* ocda db: built from the .ssp, (tethered)Ralph Amissah2026-09-092-3/+27
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The ocda database is now built from the .ssp itself: the writer's lines are emitted, read straight back by ssp_in, and the objects that come out populate the database. Anything the .ssp does not carry, the database will not have either, by construction (rather than by test). [instead of as previously through a second walk over the in-memory abstraction] - spineAbstractionTxt is split: sspDocumentLines(doc) returns the whole .ssp as lines, and the file writer emits them. Output neutral, the reference test confirms. - spineAbstractionDb takes the abstraction as an argument rather than taking doc.abstraction. - sspRoundTripAbstraction(doc) in ssp_in is the join: lines out, lines in, abstraction returned. Both call sites in spine.d use it. - the header blocks and the image blobs still come from doc_matters (as: the .ssp does not carry image bytes). All (35) markup sample sourced databases built through the .ssp have byte identical SQL dumps to the one built directly before the change. That comparison also found one reader inaccuracy, which the .ssp round trip could not see because the writer omits the field either way: an absent identifier was restored as the ocn in every case, but for an object with no ocn it was empty ("a"~N identifiers are always written). Fixed; the two artefacts checking each other is what caught it. (assisted by Claude-Code)
* a reader for the ocda db, and its round tripRalph Amissah2026-09-092-0/+266
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | sisudoc.ocda.abstraction.db_in reads a <doc>.abstraction.db back into ObjGenericComposite[][string], the same value ssp_in returns from a .ssp, so a consumer need not know which artefact it was handed. It mixes in the .ssp reader for that shared document struct rather than declaring a second one. --db-round-trip=<file.abstraction.db> reads a database and emits it as .ssp on stdout, through sspObjectRecord as the other round trip does. Held against the .ssp written from the same document, this says whether the two artefacts really carry the same thing: not a count of fields, as test-abstraction-db.sh does, but the whole document reconstructed from the database and compared to the text. SpinePOD=... ./test/test-abstraction-db-roundtrip.sh ./bin/spine-ldc PASS: all 35 databases re-emit their document's .ssp exactly It passed on the first run over the whole sample set, which is evidence that the database is now field-complete against the .ssp rather than merely counting the same. Four tests with different checks: test-abstraction-ssp.sh the abstraction has not changed test-abstraction-db.sh the two serialisations agree, field by field test-abstraction-ssp-roundtrip.sh the .ssp can be read back whole test-abstraction-db-roundtrip.sh the .db can be read back whole (assisted by Claude-Code)
* ocda: clean heading text used for navigationRalph Amissah2026-09-093-20/+97
| | | | | | heading text used for navigation is normalised, and | escaped (assisted by Claude-Code)