aboutsummaryrefslogtreecommitdiffhomepage
path: root/src/sisudoc/spine.d
Commit message (Collapse)AuthorAgeFilesLines
* pod: generate the po4a configuration, --po4a-cfgRalph Amissah2026-09-221-7/+26
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A document's po4a configuration is derived rather than kept. Spine already knows both halves without looking: the manifest says which languages, and the markup says which files, through the << lines the parser follows. Derived, it cannot disagree with the document, where a hand-kept file eventually does: a language added to the manifest and forgotten in the Makefile is a translation nobody updates. It prints, rather than writing. Spine does not write into a source tree, and the file belongs beside the catalogues in the pod it describes: spine --po4a-cfg <pod> > <pod>/tools/po4a/po4a.cfg Three things the Makefile it replaces knew, which had to be found by running po4a rather than by reading: --keep 0 is the difference between regenerating a document and destroying it. po4a will not write a translated file below a completeness threshold, and that threshold defaults to 80%. A partly translated document is the normal state of one being worked on, and sisu covers it: an untranslated string falls back to the source text. Measured at the default, 181 of live-manual's 189 translated files would be discarded, and nine languages would collapse to eight surviving files, reading as the document reverting to its source language rather than as a failure. index.html.in is not markup and no << line names it, so a configuration built from the manifest and the insert list alone drops it, and with it index.html.in.pot and its nine .po files at the next regeneration. It is found by asking whether the source language has one. neverwrap is not set, and the measurement is recorded beside the code: setting it turns 889 of 1126 strings fuzzy in four languages, and po4a does not use a fuzzy translation, so Catalan's coverage would fall from about 46% to about 12%. Tidy formatting is not worth that. The run-complete banner is suppressed for this action. The configuration goes to stdout and the banner would otherwise be part of the file, which is what the round trips already avoid by leaving early; this one prints from inside output processing and cannot. (assisted by Claude-Code)
* pod: write <doc>.sisupod (single zstd frame)Ralph Amissah2026-09-221-0/+19
| | | | | | | | | | | | | | | | | | | | | | | | | | | | The pod archive spine writes is now named .sisupod and carries the same archive inside one zstd frame. The reader has accepted both since the previous commit, so every pod already published stays readable and this changes only what is written. free_culture 2863504 -> 1100180 live-manual 6972418 -> 692942 10.06x The wrap is at the pod call site and not in createZipFile, which also writes every epub and every odt. Best compression from whole-stream (rather than per-member compression). (for live-manual per-member deflate manages about 3x compared to 10x for whole pod content compression). The suffix is named as pod_archive_suffix. --pod-compression sets the level, 19 by default: a pod is written once and fetched many times (692942 bytes at 19, 837703 at 9, 2054365 at 1). (A value that is not a number warns and the default is used). Output built from a written .sisupod is byte identical to output built from the markup it came from. (assisted by Claude-Code)
* downloads: honour --allow-downloadsRalph Amissah2026-09-221-1/+11
| | | | | | | | | The flag has existed since a url argument could be fetched and nothing ever read it, so every url was fetched whether or not the flag was given. A refused url is now dropped from the arguments, which is what a failed download already did. (assisted by Claude-Code)
* some file renamesRalph Amissah2026-09-121-3/+3
|
* spine: a .ocda.db can be fetched, as a pod zip canRalph Amissah2026-09-121-1/+1
| | | | | | | | | | | | | | | | | A URL ending in .ocda.db is now downloaded and processed in place, through the path that already did it for .zip: the same curl call, the same size and timeout limits, the same refusal of local and private addresses, the same --allow-downloads guard, the same temp file cleanup. Only the pattern had to widen. rgx_url_zip ^https?://...[.]zip$ rgx_url_source ^https?://...([.]zip|[.]ocda[.]db)$ downloadZipUrl is no downloadSourceUrl and the download temp directory is spine-download rather than spine-zip-pod; (the extraction directory keeps its name, being still only for zips). (assisted by Claude-Code)
* ocda: the images an artefact carries, and the ones it only describesRalph Amissah2026-09-121-0/+8
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The document markup sample collection builds byte identically from either artefact: 1457 files, no differences, from markup, from the 35 .ssp files, and from the 35 .ocda.db files. A .ocda.db carries its images as blobs. dbReadFiles() reads them and they are written where the output writers look for images: a directory of this run's making, removed when the run ends, not the pod beside the artefact, (which could be a working published tree). The document's source path moves with it, image_dir_path being reached from the document's own file rather than from the pod. Each blob is checked against the sha256 the database recorded beside it. The two were written together, so a mismatch means the file is damaged. (the point of having recorded the digest). A .ssp only describes its images, so those are the ones in the pod it sits in, and they are checked against the digests it recorded. That asymmetry arises from of what the two artefacts are: the ocda.db is self-sufficient and can prove it is looking at the right image, the .ssp describes a pod it sits alongside. Found while wiring it: setting the pod directory alone was not enough, since image_dir_path is derived from the source file's path (../../image from media/text/<lang>/), so the source path has to be moved with it. (assisted by Claude-Code)
* spine: process a document from .ocda.db or .sspRalph Amissah2026-09-121-0/+54
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Process a document from its .ocda.db or .ssp abstraction. A .ssp or a .ocda.db given as an argument is now a document like any other: read rather than parsed, and handed to the output writers as the same doc they already take. The 35 document sample collection built from its .ssp files is byte identical to the same collection built from markup: 1457 files, no differences. Three things had to be rebuilt rather than read, all of them a value parsed rather than a structure derived: - classify_topic_register_arr and its expanded twin. The split is not a plain one, so the rule moved to sisudoc.ocda.meta.topic_register and both the yaml reader and this one call it. - creator_author_arr, which is creator.author split on the ", " it was joined with. - title_sub, a copy of title_subtitle made where the header is read, and what epub3 puts in dc:title id="subtitle". Fixed ordering bug found by the acceptance test. A heading's own anchor can be a bare number taken from its text and this can collide with the ocn of an unrelated object. Whole output comparison found that before the fix there was one wrong link in one epub's table of contents. From a .ocda.db, 1446 of 1457 files are identical. The 11 that are not are the images: a database carries its own image blobs and nothing yet extracts them, so the five sisu_markup images are not copied and the epub that embeds them differs. That is the next step and is not a defect in this one. --source and --pod2 are refused with a warning rather than half done, no artefact carrying the markup. (assisted by Claude-Code)
* ocda: the abstraction carries the document metadata output readsRalph Amissah2026-09-121-2/+4
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | .ssp format v0.1 -> v1.1, and the .ocda.db with it: they are two serialisations of one format and now move together. The gap this closes, measured against the writers' read set: of the 59 conf_make_meta fields outputs/io_out/ reads, 29 were in neither artefact. Five are conf.* and stay out, being site and run scoped (urls, papersize, the search db filename): the same abstraction published to two sites must take each site's. Of the rest, three are never assigned anywhere (title_short, publisher, original_publisher; read only by sqlite.d, always empty) and one is a copy of a field already carried (title_sub = title_subtitle), so 17 properties actually had to travel and now do: make breaks, footer, home_button_text meta title.edition, date.added_to_site, language.document_char, original.{title,source,language,language_char}, rights.copyright_{text,translation,illustrations, photographs,cover,audio,video} make splits by when it acts, which is worth keeping in mind: these three are read in outputs/io_out/ and must travel, while italics, bold, emphasis, substitute and headings are read at parse time and their effect is already in the objects. New @source block, and source.* rows in the database, so a reader can say which markup an abstraction came from rather than working from stale content in silence: language the document's own languages the pod's list, which is what the inter-language links in html need and neither artefact carried digest sha256 of the .sst; equals its digests.txt entry The database adds source.ssp_digest, the sha256 of the .ssp it was built from, since a file cannot hold its own hash. So the chain .sst -> .ssp -> .ocda.db is checkable end to end. The database's metadata table is now filled from the header blocks the .ssp gives back, not from a second list read off doc_matters. One list, in ssp.d: a property added there arrives in the database with nothing else changed, and one that is not in the .ssp cannot be in the database at all. That was the last place the two could drift; the objects stopped being able to on 2026-09-07. Version is checked on load. The major part must match, a newer minor is accepted (a minor bump only adds properties, and an unknown property line is ignored). A v0.1 artefact is now refused with a message saying to regenerate it, rather than loading half populated. test-abstraction-db.sh now compares the two artefacts' header blocks property by property, 36 per document on the wealth of networks, where it previously only checked that a schema.version row existed. Verified non-vacuous: dropping one property is reported. Reference regenerated, and the whole diff is this change and nothing else: 320 lines added, 35 removed over 35 files, being 35 format lines changed, 35 @source blocks (4 lines each), 35 language.document_char, 35 home_button_text (it has a default), 18 breaks, 16 footer, 4 date.added_to_site, 1 title.edition, 1 original.source. Existing outputs unaffected. (assisted by Claude-Code)
* --serial default behaviour (--parallel an option)Ralph Amissah2026-09-101-31/+25
| | | | | | | | | | serial processing, it turns out, is significantly faster and more efficient for tested use-cases, which came as a surprise. As the parallelization option buys nothing, serial processing is set as default. Parallel processing remains as an option (where available, as before). (assisted by Claude-Code)
* parallelise: show_abstraction & ocda_db as the restRalph Amissah2026-09-091-0/+9
| | | | | | | | | | | | | | | | | | | | | | | | Fix issue with consistency (flags run serial & parallel inconsistenly). Both write one file per document per language and share no handle, so they belong on the list with html, epub, text and sqlite_discrete. The guard above the list still takes out --pod, --pod2, --source and the shared sqlite db actions before it is reached, so those stay serial as they were; checked. However measurement tests show parallel turn out to result in a processing slowdown, on the (35) sample markup documents, run on multiple passes: run on 16 cores a slow down of about 20% for eleven times the cpu! --text 4.83-5.10 s wall 48-52 s user --text --serial 3.94-4.05 s wall 4.2-4.4 s user will make serial run the default. (assisted by Claude-Code)
* sqlite: schema version, & fail run on writes failRalph Amissah2026-09-091-0/+12
| | | | | | | | | | | | | | | | | | | | | | | | | Spine now declares sqlite_db_schema_version and stamps it into the database as PRAGMA user_version when the tables are created, in both the shared and the discrete DDL blocks. On opening an existing database it compares, and says once per run which version it found and which it writes. Failures are now tallied (shared, the output can run in parallel), reported one line each on stderr naming the operation, and main exits 1 without printing "run complete, ok". Two tests under test/, both taking the spine binary as their first argument and building their own database from data/pod unless $SpinePOD says otherwise: test-search-db-schema.sh names: every column the search form uses exists, and spine's declaration, the database's stamp and the search form's expectation all agree test-search-cgi.sh behaviour: the real search binary answers real requests against a fresh database, no web server involved, the probe values read out of whichever database it is given (assisted by Claude-Code)
* --ocda-db replaces --show-abstraction-dbRalph Amissah2026-09-091-15/+17
| | | | | | | | | ocda (object centric document abstraction) flag --ocda-db (or --abstraction-db) replaces --show-abstraction-db rename results in consequently large diff
* ocda loader: one way in whatever the sourceRalph Amissah2026-09-091-0/+18
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | sisudoc.ocda.abstraction.load names the five things a document can be read from, tells them apart, and loads the two that are self-describing artefacts: .sst / .ssm + images the markup source pod (dir) + images the same, bundled pod .zip the same, zipped .ssp + images the abstraction, as text .ocda.db the abstraction, sqlite, images inside abstractionSourceOf(path) is the detection, by name and for a directory by whether it holds pod.manifest. abstractionLoad(path) returns a LoadedAbstraction: the source kind, whether it was loaded, why not when it was not, and the document itself. The three source forms are deliberately not loaded here. Reading them is the parser's job (sisudoc.ocda.meta.metadoc spineAbstraction) and it needs the manifest, environment and configuration that spine.d assembles, none of which belongs in a loader. What this gives that case is the dispatch and a plain statement of where it is handled, rather than a silent empty result. spine --abstraction-source=<path> says what a path is and, for an artefact, loads it and reports what came back: the document, title and author, header block sizes, object counts and the objects in each section. Exit 0 when an abstraction was loaded, 1 when not. (assisted by Claude-Code)
* ocda db: <doc>.ocda.db, and a script to run 4 testsRalph Amissah2026-09-091-2/+2
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The per document database is now written as <doc_uid>.ocda.db rather than <doc_uid>.abstraction.db. Shorter, and it says what is in the file: the object centric document abstraction, not "some abstraction". .ocda.db pairs with .ssp and cannot be mistaken for the collection search database (spine.search.db). test/run-tests.sh runs the four in sequence, one line of result each, and a summary. It re-runs itself inside nix shell "nixpkgs#sqlite" if sqlite3 is not on PATH, so this is all that is needed: SpinePOD=../../markup/sisudoc-spine-samples/markup/pod-samples/pod \ ./test/run-tests.sh ./bin/spine-ldc The tests are independent (with test-abstraction-ssp.sh run first): 1 test-abstraction-ssp.sh is first because it is the one that says whether the abstraction itself moved; if it fails the others are answering a different question than you think 2 test-abstraction-ssp-roundtrip.sh reads the committed reference set, so it is a statement about the current binary only if 1 passes 3 test-abstraction-db.sh and 4 test-abstraction-db-roundtrip.sh generate both artefacts themselves and depend on nothing committed (assisted by Claude-Code)
* ocda db: built from the .ssp, (tethered)Ralph Amissah2026-09-091-4/+16
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The ocda database is now built from the .ssp itself: the writer's lines are emitted, read straight back by ssp_in, and the objects that come out populate the database. Anything the .ssp does not carry, the database will not have either, by construction (rather than by test). [instead of as previously through a second walk over the in-memory abstraction] - spineAbstractionTxt is split: sspDocumentLines(doc) returns the whole .ssp as lines, and the file writer emits them. Output neutral, the reference test confirms. - spineAbstractionDb takes the abstraction as an argument rather than taking doc.abstraction. - sspRoundTripAbstraction(doc) in ssp_in is the join: lines out, lines in, abstraction returned. Both call sites in spine.d use it. - the header blocks and the image blobs still come from doc_matters (as: the .ssp does not carry image bytes). All (35) markup sample sourced databases built through the .ssp have byte identical SQL dumps to the one built directly before the change. That comparison also found one reader inaccuracy, which the .ssp round trip could not see because the writer omits the field either way: an absent identifier was restored as the ocn in every case, but for an object with no ocn it was empty ("a"~N identifiers are always written). Fixed; the two artefacts checking each other is what caught it. (assisted by Claude-Code)
* a reader for the ocda db, and its round tripRalph Amissah2026-09-091-0/+41
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | sisudoc.ocda.abstraction.db_in reads a <doc>.abstraction.db back into ObjGenericComposite[][string], the same value ssp_in returns from a .ssp, so a consumer need not know which artefact it was handed. It mixes in the .ssp reader for that shared document struct rather than declaring a second one. --db-round-trip=<file.abstraction.db> reads a database and emits it as .ssp on stdout, through sspObjectRecord as the other round trip does. Held against the .ssp written from the same document, this says whether the two artefacts really carry the same thing: not a count of fields, as test-abstraction-db.sh does, but the whole document reconstructed from the database and compared to the text. SpinePOD=... ./test/test-abstraction-db-roundtrip.sh ./bin/spine-ldc PASS: all 35 databases re-emit their document's .ssp exactly It passed on the first run over the whole sample set, which is evidence that the database is now field-complete against the .ssp rather than merely counting the same. Four tests with different checks: test-abstraction-ssp.sh the abstraction has not changed test-abstraction-db.sh the two serialisations agree, field by field test-abstraction-ssp-roundtrip.sh the .ssp can be read back whole test-abstraction-db-roundtrip.sh the .db can be read back whole (assisted by Claude-Code)
* ocda: a reader for .ssp, and a round trip checkRalph Amissah2026-09-091-0/+45
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | sisudoc.ocda.abstraction.ssp_in reads a .ssp file back into ObjGenericComposite[][string], the same value the parser produces, so anything that consumes the abstraction can be fed from a .ssp instead of from markup. The three header blocks come back as key/value with their order preserved. --ssp-round-trip=<file.ssp> loads a file and emits it again on stdout, using sspObjectRecord, the writer's own definition of a record. So the check is against the writer, not against a second description of the format: ./bin/spine-ldc --ssp-round-trip=test/reference/abstraction/<doc>.ssp \ | diff test/reference/abstraction/<doc>.ssp - BUG as yet to FIX 27 of the 35 reference documents round trip byte identically. The other 8 fail on two defects in the *writer* that the round trip found, and which are left for a decision: 1. .heading_ancestors_text and .lev4_subtoc can carry a raw newline, because a heading's text may contain a line break. The value then spans two physical lines and the format's rule that a value runs to the end of the line is broken. 194 and 22 occurrences, in the seven live-manual translations. 2. .heading_ancestors_text joins its eight slots with "|" while the text in them may itself contain "|". 22 occurrences in revisiting_the_autonomous_contract. Both need an escape (or normalisation at source) and both change the .ssp, so require a decision and another reference regeneration. ocda: export the .ssp reader from the abstraction package package.d is the re-export surface for consumers that want to reach the abstraction without depending on the directory layout; the reader belongs there beside the writer. (assisted by Claude-Code)
* .ssp: doc structure related fixes (& to epub toc_nav)Ralph Amissah2026-08-281-0/+2
| | | | | | | | | | | | | | | | | | | | | | | | | | for document abstraction and its .ssp output, removed the requirement of including --abstraction & --serial flags to produce correct output (for: .dom_status, .dom_status_collapsed & .last_descendant) - meta_processing_xml_dom() includes show_abstraction, so --pod2 and --show-abstraction run the dom pass; last_descendant is derived from that pass via after_doc_get_descendants() The accumulators are now verified as eight wide locals of docAbstraction(), so each document starts clean and no two threads share one. - bug: the four dom accumulators were template scope (shared) and nine wide, while their end of document reset was eight wide, so the first document of a run differed from the rest and parallel runs raced on one buffer, (which also mis-nested epub toc_nav) test/ reference .ssp regenerated: accelerando only, trailing zero dropped. test-abstraction-ssp.sh now runs parallel and diffs output against a serial run. (assisted by Claude-Code)
* ocda + outputs split: module/import + dub.json fixupsRalph Amissah2026-05-251-35/+35
| | | | | | | | | | | Modules and imports rewritten to sisudoc.ocda.* and sisudoc.outputs.*; dub.json excludedSourceFiles and the spine:abstraction sub-package sourcePaths collapsed to ./src/sisudoc/ocda. Verified: nix build .#spine-overlay-ldc clean. (assisted by Claude-Code)
* org files out of sync, fixsisudoc-spine_v0.20.0Ralph Amissah2026-05-231-0/+0
| | | | (also cgi_sqlite_search_form.d did not belong here)
* decouple abstraction phase1:2Ralph Amissah2026-05-221-2/+2
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | phase1 step2: move SSP serialiser into sisudoc.abstraction package git mv src/sisudoc/io_out/create_abstraction_txt.d to src/sisudoc/abstraction/ssp.d Module rename: sisudoc.io_out.create_abstraction_txt -> sisudoc.abstraction.ssp Completes phase1: after this commit the sisudoc.abstraction package has zero outgoing edges into sisudoc.io_out. The library produces both the in-memory document object model AND the .ssp text serialisation without referencing any output-side module. The serialiser previously imported sisudoc.io_out.paths_output for the single purpose of constructing the .ssp output path. That import is dropped; the path construction is inlined as three lines of std.path (chainPath / asNormalizedPath / array) producing <output_path>/<language>/abstraction/<doc_uid_out>.ssp - byte-for-byte the same path the previous spineOutPaths!() call produced. Updated: - src/sisudoc/abstraction/ssp.d - module decl + inline path - src/sisudoc/abstraction/package.d - public import .ssp - src/sisudoc/spine.d - import sisudoc.abstraction.ssp (x2) Completes decouple abstraction phase1 (assisted by Claude-Code)
* decouple abstraction phase0:2Ralph Amissah2026-05-221-3/+3
| | | | | | | | | | | | | | | phase0 step2: move curation modules from meta/ to io_out/curate/ Curation modules moved to src/sisudoc/io_out/curate/, module declarations renamed sisudoc.io_out.curate.metadoc_curate* from sisudoc.meta.metadoc_curate* and updated spine.d imports. File contents are otherwise unchanged. Completes phase0: meta/ now has zero io_out imports - the abstraction core's outgoing deps are now only: meta/ internals + io_in/ + ext_depends/D-YAML (assisted by Claude-Code)
* add --pod2 flag, decouple --show-abstraction from --podRalph Amissah2026-04-221-3/+8
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Finer-grained control over when .ssp files are produced: --show-abstraction writes .ssp to OUTPUT/lang/abstraction/ independently of any pod flag --pod builds pod without .ssp bundled --pod2 builds pod with .ssp in media/abstraction/ Changes to spine.d: - show_abstraction() now only responds to its own flag and pod2, no longer triggered by source_or_pod - Add pod2 to opts init, getopt, OptActions - pod() returns true for both --pod and --pod2 - source_or_pod() includes pod2 Changes to source_pod.d: - Remove per-document pod directory (rmdirRecurse) before regeneration, ensuring clean slate on every run. This prevents stale content from previous runs (e.g. a --pod2 run followed by --pod would otherwise leave an outdated media/abstraction/ directory) - Gate abstraction directory creation and .ssp bundling on pod2 flag specifically Tested: --pod (no .ssp), --pod2 (.ssp in pod + zip), --show-abstraction (standalone .ssp), --pod after --pod2 (stale abstraction cleaned up). All 35 sample documents pass. Co-Authored-By: Anthropic Claude Opus 4.6 (1M context)
* include .ssp document abstraction in source podRalph Amissah2026-04-221-1/+1
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | - When --source/--pod is used, automatically generate the .ssp document abstraction and bundle it into the pod at media/abstraction/{doc_uid}.{lang}.ssp - This makes show_abstraction implicitly true when source_or_pod is active, so the .ssp file is generated before the pod assembler runs (abstraction runs before outputHub, and source_or_pod is the first task in outputHub). - Changes: paths_source.d: Add abstraction_root() path helper to _PodPaths struct, following the same pattern as image_root(). Produces paths like pod/media/abstraction/ for both zpod (inside zip) and filesystem_open_zpod (open directory). source_pod.d: - Create media/abstraction/ directory in podArchive_directory_tree - Bundle .ssp file in pod_zipMakeReady: reads from the abstraction output directory, copies to open pod directory, adds to zip archive, computes SHA-256 digest - Write .ssp digest in zipArchiveDigest alongside sstm and ssi digests spine.d: Make show_abstraction() return true when source_or_pod is active (previously only returned true for explicit --show-abstraction flag). - The .ssp is always included when building pods - no exclusion flag for this experimental feature to keep things simple. Not generated for non-pod outputs (--text, --html, etc.) unless --show-abstraction is explicitly passed. - Tested against all 35 sample documents - zero failures. Co-Authored-By: Anthropic Claude Opus 4.6 (1M context)
* document abstraction as per document sqlite dbRalph Amissah2026-04-221-0/+18
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | --show-abstraction-db flag to write per-document - SQLite database of document abstraction (Claude-Code primary assist) - Add a new output mode that serializes the in-memory document abstraction to a per-document SQLite database. This complements the .ssp text format (--show-abstraction) with a queryable database representation of the same data. - Schema: metadata table - key/value pairs for document metadata (title, creator, dates, rights, classify, identifiers, language, notes, make settings, doc_has counts) objects table - one row per document object with columns: section, seq (position within section), ocn, is_a, is_of_part, is_of_type, heading_level, identifier, parent_ocn, last_descendant_ocn, ancestors, indent/bullet/lang, has_* flags, segment/anchor tags, table/code properties, text content Indexed on: section, ocn, parent_ocn, is_a, heading_level - Uses prepared statements via d2sqlite3 (existing dependency) for safe and efficient insertion. Each document produces a standalone .abstraction.db file in the abstraction/ output directory. - New files: src/sisudoc/io_out/create_abstraction_db.d Follows the same pattern as create_abstraction_txt.d. Creates schema, populates metadata via key/value inserts, then iterates all sections writing objects with prepared statements within a single transaction. - Changes to spine.d: - Add "show-abstraction-db" to opts init, getopt, OptActions - Add to abstraction(), require_processing_files(), and meta_processing_general() gates - Insert call at both spineAbstraction sites - Tested against all 35 sample documents (including 9-language live-manual) - zero failures. Works standalone or combined with --show-abstraction and other output flags. - Example queries the database supports: SELECT ocn, heading_level, text FROM objects WHERE is_a = 'heading' AND section = 'body'; SELECT * FROM objects WHERE parent_ocn = 10; SELECT key, value FROM metadata WHERE key LIKE 'title.%'; Co-Authored-By: Anthropic Claude Opus 4.6 (1M context)
* .ssp document abstraction as PEG parsable textRalph Amissah2026-04-221-0/+18
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | --show-abstraction flag to write .ssp document abstraction files - Add a new output mode that serializes the in-memory document abstraction (produced by spineAbstraction) to a human-readable, line-oriented text format (.ssp). This captures the full object model after parsing and abstraction but before output generation. - The .ssp format uses unambiguous line prefixes: @section { } - section boundaries (head/toc/body/endnotes/...) [N] type - object declaration with OCN .name: value - object properties (only non-defaults) | content - text content lines % comment - comments - New files: src/sisudoc/io_out/create_abstraction_txt.d Serializer module following the same template pattern as metadoc_show_summary.d. Walks doc.abstraction() section by section, writing metadata preamble (@meta, @make, @doc_has) then each object with its properties and text content. Output goes to {output_path}/{lang}/abstraction/{doc}.ssp - Changes to spine.d: - Add "show-abstraction" to opts initialization, getopt, and OptActions struct - Add show_abstraction to abstraction(), require_processing_files(), and meta_processing_general() so the flag triggers full document processing - Insert call at both spineAbstraction sites (parallel and serial branches), gated by show_abstraction flag, following the same pattern as show_config/show_summary/show_make - Tested against all 35 sample documents (including multilingual live-manual in 9 languages) - zero failures. Works standalone (--show-abstraction) or combined with other output flags (--show-abstraction --html --text). No effect on existing code paths when the flag is not used. Co-Authored-By: Anthropic Claude Opus 4.6 (1M context)
* upkeep, update a few pathssisudoc-spine_v0.18.0Ralph Amissah2026-04-221-10/+10
|
* spine may be run against a zipped spine-pod urlRalph Amissah2026-04-131-2/+25
| | | | | | - claude contributed src - processes zip from url using (system installed) curl for download
* spine may be run against a document-markup zip podRalph Amissah2026-04-131-2/+178
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | - claude contributed src - Opens the zip with std.zip.ZipArchive (reads the whole file into memory) - Locates pod.manifest inside the archive to discover document paths and languages - Extracts markup files (.sst/.ssm/.ssi) as in-memory strings - Extracts images as in-memory byte arrays - Extracts conf/dr_document_make if present - Presents these to the existing pipeline as if they were read from the filesystem - Some security mitigations: - Zip Slip / Path Traversal: Reject entries containing `..` or starting with `/`; canonicalize resolved paths and verify they fall within extraction root - Zip Bomb: Check `ArchiveMember.size` before extracting; enforce per-file (50MB) and total size limits (500MB) - Entry Count: Limit number of entries (a pod should have at most ~100 files) - Path depth: limit (Maximum 10 path components). - Symlinks: Verify no symlinks in extracted content before processing (post-extraction recursive scan) - Filename Validation: Only allow expected characters; reject null bytes - Malformed Zips: Catch `ZipException` from `std.zip.ZipArchive` constructor - Cleanup on error
* 2026Ralph Amissah2026-01-091-1/+1
|
* text output, improve various (including no-ocn)Ralph Amissah2025-10-141-1/+3
| | | | - revisit links (fix later)
* a text output (and skel an outline)Ralph Amissah2025-10-031-4/+26
| | | | - spine --text [--output=output path] [markup source]
* terminal output verbosity levels, minor reworkRalph Amissah2025-09-251-27/+54
|
* spine.d tidyRalph Amissah2025-09-231-109/+88
|
* imports, make line searchableRalph Amissah2025-07-151-26/+24
|
* minorRalph Amissah2025-04-021-1/+0
|
* doc (metadata & abstraction) struct follow throughRalph Amissah2025-02-191-6/+2
|
* document (metadata & abstraction) structRalph Amissah2025-02-191-54/+44
| | | | | | - struct replaces tuple - some direct naming of structs returned (instead of use of auto) - minor
* 2025Ralph Amissah2025-01-011-1/+1
|
* pod zip fixesRalph Amissah2024-07-101-18/+13
| | | | | - serial processing (need to be built serially) - multilingual pods, copy all languages before zip
* 0.16.0 sisudoc (src/sisudoc sisudoc spine)sisudoc-spine_v0.16.0-devRalph Amissah2024-04-101-0/+1272
- src/sisudoc (replaces src/doc_reform) - sisudoc spine (used more)