| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A document's po4a configuration is derived rather than kept. Spine
already knows both halves without looking: the manifest says which
languages, and the markup says which files, through the << lines the
parser follows. Derived, it cannot disagree with the document, where a
hand-kept file eventually does: a language added to the manifest and
forgotten in the Makefile is a translation nobody updates.
It prints, rather than writing. Spine does not write into a source tree,
and the file belongs beside the catalogues in the pod it describes:
spine --po4a-cfg <pod> > <pod>/tools/po4a/po4a.cfg
Three things the Makefile it replaces knew, which had to be found by
running po4a rather than by reading:
--keep 0 is the difference between regenerating a document and
destroying it. po4a will not write a translated file below a
completeness threshold, and that threshold defaults to 80%. A partly
translated document is the normal state of one being worked on, and sisu
covers it: an untranslated string falls back to the source text.
Measured at the default, 181 of live-manual's 189 translated files would
be discarded, and nine languages would collapse to eight surviving
files, reading as the document reverting to its source language rather
than as a failure.
index.html.in is not markup and no << line names it, so a configuration
built from the manifest and the insert list alone drops it, and with it
index.html.in.pot and its nine .po files at the next regeneration. It is
found by asking whether the source language has one.
neverwrap is not set, and the measurement is recorded beside the code:
setting it turns 889 of 1126 strings fuzzy in four languages, and po4a
does not use a fuzzy translation, so Catalan's coverage would fall from
about 46% to about 12%. Tidy formatting is not worth that.
The run-complete banner is suppressed for this action. The configuration
goes to stdout and the banner would otherwise be part of the file, which
is what the round trips already avoid by leaving early; this one prints
from inside output processing and cannot.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The pod archive spine writes is now named .sisupod and carries the
same archive inside one zstd frame. The reader has accepted both
since the previous commit, so every pod already published stays
readable and this changes only what is written.
free_culture 2863504 -> 1100180
live-manual 6972418 -> 692942 10.06x
The wrap is at the pod call site and not in createZipFile, which
also writes every epub and every odt.
Best compression from whole-stream (rather than per-member compression).
(for live-manual per-member deflate manages about 3x compared to 10x for
whole pod content compression).
The suffix is named as pod_archive_suffix.
--pod-compression sets the level, 19 by default: a pod is written once
and fetched many times (692942 bytes at 19, 837703 at 9, 2054365 at 1).
(A value that is not a number warns and the default is used).
Output built from a written .sisupod is byte identical to output built
from the markup it came from.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
| |
The flag has existed since a url argument could be fetched and nothing
ever read it, so every url was fetched whether or not the flag was
given. A refused url is now dropped from the arguments, which is what a
failed download already did.
(assisted by Claude-Code)
|
| | |
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A URL ending in .ocda.db is now downloaded and processed in place,
through the path that already did it for .zip: the same curl call, the
same size and timeout limits, the same refusal of local and private
addresses, the same --allow-downloads guard, the same temp file cleanup.
Only the pattern had to widen.
rgx_url_zip ^https?://...[.]zip$
rgx_url_source ^https?://...([.]zip|[.]ocda[.]db)$
downloadZipUrl is no downloadSourceUrl and the download temp directory
is spine-download rather than spine-zip-pod; (the extraction directory
keeps its name, being still only for zips).
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The document markup sample collection builds byte identically from
either artefact: 1457 files, no differences, from markup, from the 35
.ssp files, and from the 35 .ocda.db files.
A .ocda.db carries its images as blobs. dbReadFiles() reads them and
they are written where the output writers look for images: a
directory of this run's making, removed when the run ends, not the
pod beside the artefact, (which could be a working published tree). The
document's source path moves with it, image_dir_path being reached from
the document's own file rather than from the pod.
Each blob is checked against the sha256 the database recorded beside it.
The two were written together, so a mismatch means the file is damaged.
(the point of having recorded the digest).
A .ssp only describes its images, so those are the ones in the pod it
sits in, and they are checked against the digests it recorded.
That asymmetry arises from of what the two artefacts are: the ocda.db is
self-sufficient and can prove it is looking at the right image, the .ssp
describes a pod it sits alongside.
Found while wiring it: setting the pod directory alone was not enough,
since image_dir_path is derived from the source file's path
(../../image from media/text/<lang>/), so the source path has to be
moved with it.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Process a document from its .ocda.db or .ssp abstraction. A .ssp or a
.ocda.db given as an argument is now a document like any other: read
rather than parsed, and handed to the output writers as the same doc
they already take.
The 35 document sample collection built from its .ssp files is byte
identical to the same collection built from markup: 1457 files, no
differences.
Three things had to be rebuilt rather than read, all of them a value
parsed rather than a structure derived:
- classify_topic_register_arr and its expanded twin. The split is not a
plain one, so the rule moved to sisudoc.ocda.meta.topic_register and
both the yaml reader and this one call it.
- creator_author_arr, which is creator.author split on the ", " it was
joined with.
- title_sub, a copy of title_subtitle made where the header is read, and
what epub3 puts in dc:title id="subtitle".
Fixed ordering bug found by the acceptance test. A heading's own anchor
can be a bare number taken from its text and this can collide with the
ocn of an unrelated object. Whole output comparison found that before
the fix there was one wrong link in one epub's table of contents.
From a .ocda.db, 1446 of 1457 files are identical. The 11 that are not
are the images: a database carries its own image blobs and nothing yet
extracts them, so the five sisu_markup images are not copied and the
epub that embeds them differs. That is the next step and is not a defect
in this one.
--source and --pod2 are refused with a warning rather than half
done, no artefact carrying the markup.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
.ssp format v0.1 -> v1.1, and the .ocda.db with it: they are two
serialisations of one format and now move together.
The gap this closes, measured against the writers' read set: of the 59
conf_make_meta fields outputs/io_out/ reads, 29 were in neither
artefact. Five are conf.* and stay out, being site and run scoped (urls,
papersize, the search db filename): the same abstraction published to
two sites must take each site's. Of the rest, three are never assigned
anywhere (title_short, publisher, original_publisher; read only by
sqlite.d, always empty) and one is a copy of a field already carried
(title_sub = title_subtitle), so 17 properties actually had to travel
and now do:
make breaks, footer, home_button_text
meta title.edition, date.added_to_site, language.document_char,
original.{title,source,language,language_char},
rights.copyright_{text,translation,illustrations,
photographs,cover,audio,video}
make splits by when it acts, which is worth keeping in mind: these three
are read in outputs/io_out/ and must travel, while italics, bold,
emphasis, substitute and headings are read at parse time and their
effect is already in the objects.
New @source block, and source.* rows in the database, so a reader can
say which markup an abstraction came from rather than working from
stale content in silence:
language the document's own
languages the pod's list, which is what the inter-language links
in html need and neither artefact carried
digest sha256 of the .sst; equals its digests.txt entry
The database adds source.ssp_digest, the sha256 of the .ssp it was built
from, since a file cannot hold its own hash. So the chain
.sst -> .ssp -> .ocda.db is checkable end to end.
The database's metadata table is now filled from the header blocks the
.ssp gives back, not from a second list read off doc_matters. One list,
in ssp.d: a property added there arrives in the database with nothing
else changed, and one that is not in the .ssp cannot be in the database
at all. That was the last place the two could drift; the objects stopped
being able to on 2026-09-07.
Version is checked on load. The major part must match, a newer minor is
accepted (a minor bump only adds properties, and an unknown property
line is ignored). A v0.1 artefact is now refused with a message saying
to regenerate it, rather than loading half populated.
test-abstraction-db.sh now compares the two artefacts' header blocks
property by property, 36 per document on the wealth of networks, where
it previously only checked that a schema.version row existed. Verified
non-vacuous: dropping one property is reported.
Reference regenerated, and the whole diff is this change and nothing
else: 320 lines added, 35 removed over 35 files, being 35 format lines
changed, 35 @source blocks (4 lines each), 35 language.document_char, 35
home_button_text (it has a default), 18 breaks, 16 footer,
4 date.added_to_site, 1 title.edition, 1 original.source.
Existing outputs unaffected.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
| |
serial processing, it turns out, is significantly faster and more
efficient for tested use-cases, which came as a surprise. As the
parallelization option buys nothing, serial processing is set as
default. Parallel processing remains as an option (where
available, as before).
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Fix issue with consistency (flags run serial & parallel
inconsistenly).
Both write one file per document per language and share no
handle, so they belong on the list with html, epub, text and
sqlite_discrete.
The guard above the list still takes out --pod, --pod2,
--source and the shared sqlite db actions before it is reached, so
those stay serial as they were; checked.
However measurement tests show parallel turn out to result in a
processing slowdown, on the (35) sample markup documents, run on
multiple passes: run on 16 cores a slow down of about 20% for
eleven times the cpu!
--text 4.83-5.10 s wall 48-52 s user
--text --serial 3.94-4.05 s wall 4.2-4.4 s user
will make serial run the default.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Spine now declares sqlite_db_schema_version and stamps it into the
database as PRAGMA user_version when the tables are created, in
both the shared and the discrete DDL blocks. On opening an
existing database it compares, and says once per run which version
it found and which it writes.
Failures are now tallied (shared, the output can run in parallel),
reported one line each on stderr naming the operation, and main
exits 1 without printing "run complete, ok".
Two tests under test/, both taking the spine binary as their first
argument and building their own database from data/pod unless
$SpinePOD says otherwise:
test-search-db-schema.sh names: every column the search form
uses exists, and spine's declaration, the database's stamp
and the search form's expectation all agree
test-search-cgi.sh behaviour: the real search binary answers
real requests against a fresh database, no web server
involved, the probe values read out of whichever database it
is given
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
| |
ocda (object centric document abstraction)
flag --ocda-db (or --abstraction-db)
replaces --show-abstraction-db
rename results in consequently large diff
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
sisudoc.ocda.abstraction.load names the five things a document can
be read from, tells them apart, and loads the two that are
self-describing artefacts:
.sst / .ssm + images the markup source
pod (dir) + images the same, bundled
pod .zip the same, zipped
.ssp + images the abstraction, as text
.ocda.db the abstraction, sqlite, images inside
abstractionSourceOf(path) is the detection, by name and for a
directory by whether it holds pod.manifest. abstractionLoad(path)
returns a LoadedAbstraction: the source kind, whether it was
loaded, why not when it was not, and the document itself.
The three source forms are deliberately not loaded here. Reading
them is the parser's job (sisudoc.ocda.meta.metadoc
spineAbstraction) and it needs the manifest, environment and
configuration that spine.d assembles, none of which belongs in a
loader. What this gives that case is the dispatch and a plain
statement of where it is handled, rather than a silent empty
result.
spine --abstraction-source=<path>
says what a path is and, for an artefact, loads it and reports
what came back: the document, title and author, header block
sizes, object counts and the objects in each section. Exit 0 when
an abstraction was loaded, 1 when not.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The per document database is now written as <doc_uid>.ocda.db
rather than <doc_uid>.abstraction.db. Shorter, and it says what is
in the file: the object centric document abstraction, not "some
abstraction". .ocda.db pairs with .ssp and cannot be mistaken for
the collection search database (spine.search.db).
test/run-tests.sh runs the four in sequence, one line of result
each, and a summary. It re-runs itself inside nix shell
"nixpkgs#sqlite" if sqlite3 is not on PATH, so this is all that
is needed:
SpinePOD=../../markup/sisudoc-spine-samples/markup/pod-samples/pod \
./test/run-tests.sh ./bin/spine-ldc
The tests are independent (with test-abstraction-ssp.sh run
first):
1 test-abstraction-ssp.sh is first because it is the
one that says whether the abstraction itself moved; if it
fails the others are answering a different question than you
think
2 test-abstraction-ssp-roundtrip.sh reads the committed
reference set, so it is a statement about the current binary
only if 1 passes
3 test-abstraction-db.sh and
4 test-abstraction-db-roundtrip.sh generate both artefacts
themselves and depend on nothing committed
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The ocda database is now built from the .ssp itself: the writer's
lines are emitted, read straight back by ssp_in, and the objects
that come out populate the database. Anything the .ssp does not
carry, the database will not have either, by construction (rather
than by test).
[instead of as previously through a second walk over the in-memory
abstraction]
- spineAbstractionTxt is split: sspDocumentLines(doc) returns the
whole .ssp as lines, and the file writer emits them. Output
neutral, the reference test confirms.
- spineAbstractionDb takes the abstraction as an argument rather
than taking doc.abstraction.
- sspRoundTripAbstraction(doc) in ssp_in is the join: lines out,
lines in, abstraction returned. Both call sites in spine.d use
it.
- the header blocks and the image blobs still come from
doc_matters (as: the .ssp does not carry image bytes).
All (35) markup sample sourced databases built through the .ssp
have byte identical SQL dumps to the one built directly before the
change.
That comparison also found one reader inaccuracy, which the .ssp
round trip could not see because the writer omits the field either
way: an absent identifier was restored as the ocn in every case,
but for an object with no ocn it was empty ("a"~N identifiers are
always written). Fixed; the two artefacts checking each other is
what caught it.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
sisudoc.ocda.abstraction.db_in reads a <doc>.abstraction.db back
into ObjGenericComposite[][string], the same value ssp_in returns
from a .ssp, so a consumer need not know which artefact it was
handed. It mixes in the .ssp reader for that shared document
struct rather than declaring a second one.
--db-round-trip=<file.abstraction.db> reads a database and emits
it as .ssp on stdout, through sspObjectRecord as the other round
trip does. Held against the .ssp written from the same document,
this says whether the two artefacts really carry the same thing:
not a count of fields, as test-abstraction-db.sh does, but the
whole document reconstructed from the database and compared to the
text.
SpinePOD=... ./test/test-abstraction-db-roundtrip.sh ./bin/spine-ldc
PASS: all 35 databases re-emit their document's .ssp exactly
It passed on the first run over the whole sample set, which is evidence
that the database is now field-complete against the .ssp rather than
merely counting the same.
Four tests with different checks:
test-abstraction-ssp.sh the abstraction has not changed
test-abstraction-db.sh the two serialisations agree, field by field
test-abstraction-ssp-roundtrip.sh the .ssp can be read back whole
test-abstraction-db-roundtrip.sh the .db can be read back whole
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
sisudoc.ocda.abstraction.ssp_in reads a .ssp file back into
ObjGenericComposite[][string], the same value the parser produces,
so anything that consumes the abstraction can be fed from a .ssp
instead of from markup. The three header blocks come back as
key/value with their order preserved.
--ssp-round-trip=<file.ssp> loads a file and emits it again on
stdout, using sspObjectRecord, the writer's own definition of a
record. So the check is against the writer, not against a second
description of the format:
./bin/spine-ldc --ssp-round-trip=test/reference/abstraction/<doc>.ssp \
| diff test/reference/abstraction/<doc>.ssp -
BUG as yet to FIX
27 of the 35 reference documents round trip byte identically. The
other 8 fail on two defects in the *writer* that the round trip
found, and which are left for a decision:
1. .heading_ancestors_text and .lev4_subtoc can carry a raw newline,
because a heading's text may contain a line break. The value then
spans two physical lines and the format's rule that a value runs to
the end of the line is broken. 194 and 22 occurrences, in the seven
live-manual translations.
2. .heading_ancestors_text joins its eight slots with "|" while the
text in them may itself contain "|". 22 occurrences in
revisiting_the_autonomous_contract.
Both need an escape (or normalisation at source) and both change the
.ssp, so require a decision and another reference regeneration.
ocda: export the .ssp reader from the abstraction package
package.d is the re-export surface for consumers that want to reach
the abstraction without depending on the directory layout; the reader
belongs there beside the writer.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
for document abstraction and its .ssp output, removed the
requirement of including --abstraction & --serial flags to produce
correct output (for: .dom_status, .dom_status_collapsed &
.last_descendant)
- meta_processing_xml_dom() includes show_abstraction, so --pod2
and --show-abstraction run the dom pass; last_descendant is
derived from that pass via after_doc_get_descendants()
The accumulators are now verified as eight wide locals of
docAbstraction(), so each document starts clean and no two threads
share one.
- bug: the four dom accumulators were template scope (shared) and
nine wide, while their end of document reset was eight wide, so
the first document of a run differed from the rest and parallel
runs raced on one buffer, (which also mis-nested epub toc_nav)
test/ reference .ssp regenerated: accelerando only, trailing zero
dropped.
test-abstraction-ssp.sh now runs parallel and diffs output against
a serial run.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
| |
Modules and imports rewritten to sisudoc.ocda.* and
sisudoc.outputs.*; dub.json excludedSourceFiles and the
spine:abstraction sub-package sourcePaths collapsed to
./src/sisudoc/ocda.
Verified: nix build .#spine-overlay-ldc clean.
(assisted by Claude-Code)
|
| |
|
|
| |
(also cgi_sqlite_search_form.d did not belong here)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
phase1 step2: move SSP serialiser into sisudoc.abstraction package
git mv src/sisudoc/io_out/create_abstraction_txt.d to
src/sisudoc/abstraction/ssp.d
Module rename: sisudoc.io_out.create_abstraction_txt
-> sisudoc.abstraction.ssp
Completes phase1: after this commit the sisudoc.abstraction package has
zero outgoing edges into sisudoc.io_out. The library produces both the
in-memory document object model AND the .ssp text serialisation without
referencing any output-side module.
The serialiser previously imported sisudoc.io_out.paths_output for the
single purpose of constructing the .ssp output path. That import is
dropped; the path construction is inlined as three lines of std.path
(chainPath / asNormalizedPath / array) producing
<output_path>/<language>/abstraction/<doc_uid_out>.ssp
- byte-for-byte the same path the previous spineOutPaths!() call
produced.
Updated:
- src/sisudoc/abstraction/ssp.d - module decl + inline path
- src/sisudoc/abstraction/package.d - public import .ssp
- src/sisudoc/spine.d - import sisudoc.abstraction.ssp (x2)
Completes decouple abstraction phase1
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
phase0 step2: move curation modules from meta/ to io_out/curate/
Curation modules moved to src/sisudoc/io_out/curate/, module
declarations renamed sisudoc.io_out.curate.metadoc_curate* from
sisudoc.meta.metadoc_curate* and updated spine.d imports. File contents
are otherwise unchanged.
Completes phase0: meta/ now has zero io_out imports - the abstraction
core's outgoing deps are now only:
meta/ internals + io_in/ + ext_depends/D-YAML
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Finer-grained control over when .ssp files are produced:
--show-abstraction writes .ssp to OUTPUT/lang/abstraction/
independently of any pod flag
--pod builds pod without .ssp bundled
--pod2 builds pod with .ssp in media/abstraction/
Changes to spine.d:
- show_abstraction() now only responds to its own flag and
pod2, no longer triggered by source_or_pod
- Add pod2 to opts init, getopt, OptActions
- pod() returns true for both --pod and --pod2
- source_or_pod() includes pod2
Changes to source_pod.d:
- Remove per-document pod directory (rmdirRecurse) before
regeneration, ensuring clean slate on every run. This
prevents stale content from previous runs (e.g. a --pod2
run followed by --pod would otherwise leave an outdated
media/abstraction/ directory)
- Gate abstraction directory creation and .ssp bundling on
pod2 flag specifically
Tested: --pod (no .ssp), --pod2 (.ssp in pod + zip),
--show-abstraction (standalone .ssp), --pod after --pod2
(stale abstraction cleaned up). All 35 sample documents pass.
Co-Authored-By: Anthropic Claude Opus 4.6 (1M context)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
- When --source/--pod is used, automatically generate the .ssp
document abstraction and bundle it into the pod at
media/abstraction/{doc_uid}.{lang}.ssp
- This makes show_abstraction implicitly true when source_or_pod
is active, so the .ssp file is generated before the pod
assembler runs (abstraction runs before outputHub, and
source_or_pod is the first task in outputHub).
- Changes:
paths_source.d:
Add abstraction_root() path helper to _PodPaths struct,
following the same pattern as image_root(). Produces
paths like pod/media/abstraction/ for both zpod (inside
zip) and filesystem_open_zpod (open directory).
source_pod.d:
- Create media/abstraction/ directory in
podArchive_directory_tree
- Bundle .ssp file in pod_zipMakeReady: reads from the
abstraction output directory, copies to open pod
directory, adds to zip archive, computes SHA-256 digest
- Write .ssp digest in zipArchiveDigest alongside sstm
and ssi digests
spine.d:
Make show_abstraction() return true when source_or_pod is
active (previously only returned true for explicit
--show-abstraction flag).
- The .ssp is always included when building pods - no exclusion
flag for this experimental feature to keep things simple.
Not generated for non-pod outputs (--text, --html, etc.)
unless --show-abstraction is explicitly passed.
- Tested against all 35 sample documents - zero failures.
Co-Authored-By: Anthropic Claude Opus 4.6 (1M context)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
--show-abstraction-db flag to write per-document
- SQLite database of document abstraction
(Claude-Code primary assist)
- Add a new output mode that serializes the in-memory document
abstraction to a per-document SQLite database. This complements
the .ssp text format (--show-abstraction) with a queryable
database representation of the same data.
- Schema:
metadata table - key/value pairs for document metadata
(title, creator, dates, rights, classify, identifiers,
language, notes, make settings, doc_has counts)
objects table - one row per document object with columns:
section, seq (position within section), ocn, is_a,
is_of_part, is_of_type, heading_level, identifier,
parent_ocn, last_descendant_ocn, ancestors,
indent/bullet/lang, has_* flags, segment/anchor tags,
table/code properties, text content
Indexed on: section, ocn, parent_ocn, is_a, heading_level
- Uses prepared statements via d2sqlite3 (existing dependency)
for safe and efficient insertion. Each document produces a
standalone .abstraction.db file in the abstraction/ output
directory.
- New files:
src/sisudoc/io_out/create_abstraction_db.d
Follows the same pattern as create_abstraction_txt.d.
Creates schema, populates metadata via key/value inserts,
then iterates all sections writing objects with prepared
statements within a single transaction.
- Changes to spine.d:
- Add "show-abstraction-db" to opts init, getopt, OptActions
- Add to abstraction(), require_processing_files(), and
meta_processing_general() gates
- Insert call at both spineAbstraction sites
- Tested against all 35 sample documents (including 9-language
live-manual) - zero failures. Works standalone or combined
with --show-abstraction and other output flags.
- Example queries the database supports:
SELECT ocn, heading_level, text FROM objects
WHERE is_a = 'heading' AND section = 'body';
SELECT * FROM objects WHERE parent_ocn = 10;
SELECT key, value FROM metadata WHERE key LIKE 'title.%';
Co-Authored-By: Anthropic Claude Opus 4.6 (1M context)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
--show-abstraction flag to write .ssp document abstraction files
- Add a new output mode that serializes the in-memory document
abstraction (produced by spineAbstraction) to a human-readable,
line-oriented text format (.ssp). This captures the full object
model after parsing and abstraction but before output generation.
- The .ssp format uses unambiguous line prefixes:
@section { } - section boundaries (head/toc/body/endnotes/...)
[N] type - object declaration with OCN
.name: value - object properties (only non-defaults)
| content - text content lines
% comment - comments
- New files:
src/sisudoc/io_out/create_abstraction_txt.d
Serializer module following the same template pattern as
metadoc_show_summary.d. Walks doc.abstraction() section by
section, writing metadata preamble (@meta, @make, @doc_has)
then each object with its properties and text content.
Output goes to {output_path}/{lang}/abstraction/{doc}.ssp
- Changes to spine.d:
- Add "show-abstraction" to opts initialization, getopt, and
OptActions struct
- Add show_abstraction to abstraction(), require_processing_files(),
and meta_processing_general() so the flag triggers full document
processing
- Insert call at both spineAbstraction sites (parallel and serial
branches), gated by show_abstraction flag, following the same
pattern as show_config/show_summary/show_make
- Tested against all 35 sample documents (including multilingual
live-manual in 9 languages) - zero failures. Works standalone
(--show-abstraction) or combined with other output flags
(--show-abstraction --html --text). No effect on existing code
paths when the flag is not used.
Co-Authored-By: Anthropic Claude Opus 4.6 (1M context)
|
| | |
|
| |
|
|
|
|
| |
- claude contributed src
- processes zip from url using (system
installed) curl for download
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
- claude contributed src
- Opens the zip with std.zip.ZipArchive (reads the whole file into
memory)
- Locates pod.manifest inside the archive to discover document paths
and languages
- Extracts markup files (.sst/.ssm/.ssi) as in-memory strings
- Extracts images as in-memory byte arrays
- Extracts conf/dr_document_make if present
- Presents these to the existing pipeline as if they were read from
the filesystem
- Some security mitigations:
- Zip Slip / Path Traversal: Reject entries containing `..` or
starting with `/`; canonicalize resolved paths and verify they
fall within extraction root
- Zip Bomb: Check `ArchiveMember.size` before extracting; enforce
per-file (50MB) and total size limits (500MB)
- Entry Count: Limit number of entries (a pod should have at most
~100 files)
- Path depth: limit (Maximum 10 path components).
- Symlinks: Verify no symlinks in extracted content before
processing (post-extraction recursive scan)
- Filename Validation: Only allow expected characters; reject null
bytes
- Malformed Zips: Catch `ZipException` from `std.zip.ZipArchive`
constructor
- Cleanup on error
|
| | |
|
| |
|
|
| |
- revisit links (fix later)
|
| |
|
|
| |
- spine --text [--output=output path] [markup source]
|
| | |
|
| | |
|
| | |
|
| | |
|
| | |
|
| |
|
|
|
|
| |
- struct replaces tuple
- some direct naming of structs returned
(instead of use of auto) - minor
|
| | |
|
| |
|
|
|
| |
- serial processing (need to be built serially)
- multilingual pods, copy all languages before zip
|
|
|
- src/sisudoc (replaces src/doc_reform)
- sisudoc spine (used more)
|