| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
- A poem is a container of verse, and stores its range of verse ocn,
- its verse are the citable units with ocn.
- A note in the last verse of a poem previously was not gathered into
the endnotes section, this now is fixed
Format 2.0 -> 2.1: the property is an addition, the poem is not a
citable object (but contans a range of objects), its verse are
(individual citable objects), and every reader checks the major version.
A 2.0 database reads as having no ranges.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
| |
The markup documents ~[* note ]~ and ~[+ note ]~ as editor's notes, each
a separately numbered series, and a bare ~[ note ]~ which sisu put in
the asterisk series. They are included as notes. Each series is numbered
through the document, *1, *2 ... and +1, +2 ..., apart from the author's
notes.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
| |
- meta_processing_general() in spine.d replaced by per-stage predicates.
- doc_matters.generated_time() commented out, unused since the odt and
epub clock stamps were removed.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A plain-text output that keeps the object numbers, written as CommonMark
with GFM pipe tables, one .md per document-language in <lang>/markdown/,
(linked from the metadata page) with images into that shared directory
itself, skipping any already there, so --markdown alone still gives a complete tree.
provided as an alternative to --text rather than a replacement
--markdown produces a document that renders, (both provide objext
numbering survives both. Every object carries an anchor and a
superscript number linking to itself, so a citation by object number
is a link in any markdown renderer.
--text may still be of value to read in a terminal
A note is written where it is referenced, with a link back to the
object it came from, rather than gathered into an endnotes
section.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
typst provides a new pdf build path replacing latex as the primary pdf
builder.
--typst (--typ) writes the .typ;
--pdf writes the .typ and compiles it by spawning `typst compile`.
(--pdf is redefined, being typst to pdf rather than latex).
--latex is unchanged and writes the .tex;
One .typ per document-language (not one per paper size). The paper and
the orientation are read from sys.inputs with a default, so the same
file compiles to every paper --set-papersize asks for, the pdfs keep
their original names.
Spine's paper names (derived from what suited latex) and typst's are
different vocabularies and are mapped; a paper typst does not have is
named and skipped.
With the flag --pdf spine invokes a typesetter to generate the pdf
output directly, which is new (the latex path writes a .tex and leaves
xelatex to the caller). If typst is not available on the machine it is
reported absent, naming the binary and printing the command, and the run
continues, the .typ is written.
Several useful features of an object-centric pdf are easily met, usually
being a few lines each.
- The object number goes in the margin as one `place` at a coordinate
inside the object's own block, carrying the label every citation
points at.
- A note is a real footnote at the point of reference, with the
document's own number and one link back to the object.
- The table of contents locates by object number rather than by page, so
it holds true across every paper and every setting.
That typst compiles silently across the sample set means every internal
link in every document resolves (--pdf over the collection with two
paper sizes gives 144 pdfs, none skipped and none failed, each the page
size its name claims and each tagged).
paths: .typ, pdf named in same way as .tex based built pdfs so site
linkage works either way
flake.nix: added a typst dev shell, for .typ pdf typesetter
- dsh-typst-pdf providing a pdf typsetter for spine's .typ.
- dsh-latex-pdf provides xelatex for .tex which can still be written.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
spine automatically builds some segments, including: (toc, endnotes,
glossary, bibliography, bookindex, blurb, _the_title) these are now
identified as reserved names and a user is now warned if any of these
names have been manually assigned to a heading by markup. A document
still builds but to disambiguate the ocn of the heading is attached to
the markup (reserved) name and this is seeded before a document not
after. It is read off the finished abstraction rather than reported by
the parser, and beside the ocn alignment check for the same reason: it
is a statement about a document rather than a step in building one.
WARNING reserved segment name: the_autonomous_contract... [en]
heading 1 at ocn 135 asks for "endnotes", which spine gives its
own generated section
it is named "endnotes-135" instead; ...
A warning: the document is correct and complete and the name it ends
up with works.
--strict makes it a failure for the run, as it does for ocn alignment,
and by the same reasoning: the outputs are written and can be looked at,
and the exit status is taken at the end.
Two of the thirty-six sample documents have reserved segment names,
"1~endnotes" heading. Output is unchanged: nothing here touches the
abstraction.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
need a character to split filenames on in certain circumstances.
there problems with use of a colon in filenames, which is legal on posix
but not on Windows."~" fits the bill better being legal on every
filesystem of interest and is unreserved in rfc 3986, needing no escaping in a url.
The reference abstraction is renamed, its content unchanged. Over
the sample collection two filenames move and nothing else does: the
abstraction and database digests are identical, the archive's
members and their sizes are unchanged, and only the member order
shifts, "~" collating after letters where ":" sorted before them. The
document's epub dc:identifier is a v5 uuid derived from the uid, so it
changes for those filenames.
Every document's uid changes in the search database (spine.search.db).
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
ensure that db uses the names it carries in manifest which produce
deterministic output (and fix divergence in db rendering of output
names, (which previously also looked for variable input from the
environment))
Output built from a database with --config naming the site configuration
is now byte identical to output built from the pod, on the render route
as it already was on the materialise route.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
dr_document_make in three identifier spellings becomes document_make,
which is what the file has been called since the rename.
The mixin import audit, a template declares its imports at template
scope and they are then visible in every scope it is mixed into (which
is how a template-level split once hijacked a UFCS lookup in code that
had not been touched). Disabling htmlSnippet's imports and rebuilding
names the consumers that were relying on them rather than on their own:
across twelve mixin sites, exactly one. metadata.d now imports the `to`
it uses. The templates keep their imports, their own functions needing
them, and narrowing those is a change of its own rather than a cleanup.
And the standing FIX in source_pod.d: the insert digest line for a non
pod source recorded the insert's path, where every other line in
digests.txt is a filename and a digest naming a file that is not in
the pod under that name cannot be checked against it.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
does this spine still produce this abstraction?
A database states an abstraction and carries the markup it was built
from, so re-parsing the one and comparing against the other says whether
the two still agree. Checked as follows:
- automatically and unskippably when a document is built from a
database, at the moment it is parsed and before anything has been
written from it.
- on --ocda-verify=<file>, an action of its own that materialises,
parses every language, compares, reports and writes nothing, the exit
status being the answer.
- on --no-verify as the escape, warning per document, for a newer spine
reading an older artefact where the abstraction is expected to differ
and the output is wanted anyway.
A build ends at the first failure rather than skipping the language.
A document rebuilt from a database now has its whole abstraction built.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The pod carries tools/po4a and the database (until now) did not, so a
pod written back out of a database came back a lossy copy, without its
translation catalogues.
the database now carries them, walked and name-checked exactly as the
pod writer walks and checks them, once on the last language. They are
compressed to save space, the test sample carrying 6.25 MB of catalogue
as read would have made that one document's database two thirds larger.
Files gains a compression column: NULL or 'none' for the bytes as read,
'zstd' for a frame. What to compress is decided by role, not by size
(source, conf, manifest and tools are text and are compressed; images
being compressed already are not). The same kind of file is always
stored the same way. Document objects are untouched and stay raw:
objects_fts is external content over objects.text and reads that column
directly.
bytes and sha256 remain those of the original file, so every digest
check, digests.txt line and source.digest rebuilt from these rows
works unchanged, and a reader that does not decompress can still say
what it is looking at.
dbReadFiles decompresses, so no caller learns how a blob is stored. It
asks pragma_table_info whether the column is there at all: a database
written before it reads as raw. A row that claims zstd and will not
decompress is named and left out, rather than handed back as a frame
where markup should be.
live-manual: 9,490,432 bytes without the catalogues, 9,908,224 with
them. Output from the materialised pod is identical to output from the
original pod, 679 files each side, no differences.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A database that carries its source can recreate the pod. --source and
--pod2 given a .ocda.db now write the pod to a directory of the run's
making and carry on with the pod, so everything after that point is
handling a pod like any other.
Nothing renders from the database. A materialiser that also rendered
would be a second path to every output format and the two would drift;
one that only writes files means the document is built by the same
code over the same bytes as the original, and identical output is a
consequence rather than an aspiration. Held against the original pod,
site configuration constant: every output file identical across ten
languages.
The pod's name comes from the database's filename, which inverts the
naming rule exactly: <doc>.ocda.db is named by doc_uid_out_no_lang,
the pod name and the document's filename joined by ":" when they
differ and the one name when they do not. So the materialised pod
recomputes the uid it was named by and every output file lands on the
name it had.
Names are checked before anything is created, and one bad name refuses
the artefact rather than skipping a file, as the zip reader does with
a zip. Markup that does not match the digest stored with it is refused
outright, where a mismatched image is written with a warning: a wrong
image makes a document that looks wrong, a wrong markup file makes one
that is wrong, in its text, with nothing downstream to notice.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The database already carried the images, because without them no output
can be produced. It now carries the markup as well, and with it the
document's configuration and the manifest as the author wrote it.
That is the difference between a serialised abstraction and a document
source. An abstraction can be rendered but not re-parsed, and a reader
who wants to correct a sentence needs the sentence as written. With
these a pod can be written back out of the sqlite-file.
Three new values in the existing role column, no schema change: the
format stays 2.0.
Source rows are named by their path within the pod,
media/text/<lang>/<file>, and not by bare filename as images are. Every
language of a document has a file of the same name, so bare names would
collide under UNIQUE(role, name) and nine of ten would be dropped
without a word. The path is also what a pod materialised from this
database has to be told.
Bytes stored as read with the digest over them, so that each row is
checkable against the line source.digest was built from.
Also build.spine_version, a file level row saying which spine wrote the
file: not a property of the document, and the one thing a file cannot be
asked for afterwards. The reader skips the build. prefix as it skips
schema. and translation., or it would come back inside the document
header and the round trip would differ.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The check runs once every language of a document has been abstracted.
Providing a warning by default. --strict makes divergence a failure,
taken at the end of the run so that outputs are complete and can be
examined. (the name --strict is general, so that later checks can be
added without a second flag).
Under --parallel the profiles are appended under synchronized, and the
comparison sorts a document's languages by name rather than taking
them in the order the threads finished, so two runs print the same
lines in the same order.
The outcome is noted per language in the database as
translation.ocn_aligned, for documents that have more than one
language.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Each language leaves a profile as it is abstracted: how many numbered
objects it has, and what kind of object each ocn is. The comparison is
against the language the manifest names first.
A count difference is reported alone. Past the first dropped or added
object every ocn names something else, and a kind comparison after it
would print hundreds of lines that are all the one fault.
The languages of a document share their object numbering, this being
the basis of ocn citation in a multi-language document: an ocn
names the same object in every language. Nothing enforces it.
It holds where translations follow the source object for object, and
stops holding otherwise (e.g. the moment a translator drops a paragraph,
merges two, or turns a heading into a sentence, at which point the
numbering no longer matches).
An ocn is not unique: a poem and its first verse share one, by design.
So the walk is in document order, the section order the .ssp is
written in rather than the abstraction's key order, which an
associative array does not promise, and the first object at an ocn is
the one recorded.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A database now holds every language of its document, so reading one
means naming a language. Three small pieces:
- spineDbLanguages says what languages a file holds. Its own template
rather than part of the reader (asking a file what it holds does not
instantiate the object setter and the markup regexes with it).
- abstractionLoad and spineDocFromArtefact take a language and pass it
to dbReadFile, which resolves it to a doc_id, takes that language
document where there is one, and lists the languages and stops where
there is a choice.
- spineArtefactLanguages turns an artefact into the documents to build:
one for a .ssp, and for a database the languages it holds filtered by
--lang. The artefact loop then has a language loop inside it, so a
single .ocda.db argument builds every language, as the pod it came
from would.
--db-round-trip takes --lang for the same reason, and ocda-db looses
--parallel (would cause a race).
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
| |
doc_uid_out ends in ".{lng}", which is right for an artefact written per
language and wrong for one that holds every language of a document.
Rather than strip the suffix at each place that needs the bare name,
build it by the same branches without the language, so the two names
cannot drift apart.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
each file still holds one language, here groundwork for one database per
document. The shape changes here and nothing merges yet: a database is
still written per language, so every doc_id is 1. Real and testable on
its own, where the writer and the reader together are not.
documents, a row per language, is what lets one file hold a document's
whole set. Everything that tells one language's rows from another's keys
on documents.id.
metadata is keyed on (doc_id, key) (no longer on key alone). Every
language has a title and a creator, and a key-only primary key refuses
the second one. schema.name and schema.version describe the file rather
than a document in it, so they are written with a null doc_id and are
the only rows that are.
objects gain doc_id and its uniqueness widens from (section, seq) to
(doc_id, section, seq). ('body', 0) exists once per language, so the
narrow constraint was the thing that would have refused a second
language outright. idx_objects_section leads with doc_id, or reading one
language scans them all.
objects.id stays a global INTEGER PRIMARY KEY, so object_images,
object_links, object_anchors and object_subtoc keep their schema and
their keys, and objects_fts keeps content_rowid='id'. The reader's four
sweeps filter through objects rather than gaining a column of their own.
outline and citable name the language and order by it first. A view over
a file that may hold several languages and does not say which reads as
one document and is several.
The DDL is IF NOT EXISTS throughout, since a second language will open a
file that already has its schema.
dbReadFile takes an optional language and means "the only document in
it" without one. Given none where there are several it reports the
languages and stops, rather than returning the first: the round trip
compares byte for byte, and a quietly wrong answer there would read as a
spine fault.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
format 2.0, source.digest covers every markup file, the sha256 over one
line per markup file, "<sha256> <filename>", sorted by filename.
source.digest was the sha256 of the master file as read, before any
insert. This is right for a .sst, which is the whole document. For a
.ssm with inserts .ssi (containing the substantive part of a documents
text) this is close to meaningless.
Considered change of meaning (rather than an addition) and given the
major part of the format version with it: 1.1 becomes 2.0, in the .ssp
header line and in the database's schema.version row together.
A 1.x reader refuses a 2.0 artefact, which is what that check is for.
- Sorted, so the value does not depend on the order a filesystem hands
back a directory.
- By filename (rather than by path), so it does not depend on where the
pod sits, which is the property the pod rebuild comparison rests on:
a pod unzipped elsewhere must still give the same abstraction.
Within one language every markup file lives in one directory, so a
filename identifies it.
Each line is checkable on its own against a digests.txt line or a files
row, which the single value was not. The whole is reproducible with
sha256sum and sort, and was verified that way for a one file document
and for a twenty file one.
An unreadable file is named in the digest rather than skipped. The parse
has failed elsewhere by then, and a digest that quietly left a file out
would claim the document is something it is not.
The reference abstractions are regenerated. Two lines change in each and
no others: the version, and the digest.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Both now take the number from one constant now set at 10,000, the writer
checks it before the archive is written.
Previously the number set for the pod reader was less than 500 members.
The writer had no limit, (so spine could write a pod it would then
refuse to read, reporting too many entries: a good file that looks
corrupt, and only for a document with enough parts).
500 was set when a pod held markup and images. A pod carrying
translation catalogues has a different arithmetic, and the count follows
from how many files a document is made of times how many languages it
has, not from how large the document is:
the_wealth_of_networks 213,405 words, 1 file per language
12 languages -> 49 entries
live-manual 24,724 words, 20 files per language
12 languages -> 506 entries
The big book is not what runs into this; the modular manual is.
live-manual stood two languages from an unreadable artefact.
10,000 covers a hundred-insert manual in thirty languages, about 6,200
entries, with room. It is deliberately well clear of any real document
rather than snug above the largest one known: the two errors are not
comparable, since too low refuses a legitimate document with a message
that reads as corruption, while too high defers to a size cap a moment
later. The count is the weakest of the three guards and is not what
bounds resource use; the per entry and total size caps do that, both
before a byte is written.
Refusing also removes any archive an earlier language left. The writer
runs once per language and only the last pass holds every language, so
it is the last that goes over, and the passes before it wrote smaller
archives that passed. Without this, refusal left a pod missing a
language: an artefact that reads perfectly well and is wrong. A missing
file is an error someone notices.
The arithmetic is recorded beside the constant so the next person can
re-derive the number rather than guess at it.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The pod archive spine writes is now named .sisupod and carries the
same archive inside one zstd frame. The reader has accepted both
since the previous commit, so every pod already published stays
readable and this changes only what is written.
free_culture 2863504 -> 1100180
live-manual 6972418 -> 692942 10.06x
The wrap is at the pod call site and not in createZipFile, which
also writes every epub and every odt.
Best compression from whole-stream (rather than per-member compression).
(for live-manual per-member deflate manages about 3x compared to 10x for
whole pod content compression).
The suffix is named as pod_archive_suffix.
--pod-compression sets the level, 19 by default: a pod is written once
and fetched many times (692942 bytes at 19, 837703 at 9, 2054365 at 1).
(A value that is not a number warns and the default is used).
Output built from a written .sisupod is byte identical to output built
from the markup it came from.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
sisupod using zstd compression recognised by first bytes.
<doc>.sisupod: spine reads one before it writes one, so when the name
changes (from .zip to .ssiupod) every pod already published stays
readable.
A .sisupod is one zstd frame wrapping the archive spine already builds.
The reader unwraps the bytes and hands the same archive to the same
parser, so every guard downstream is untouched: entry names, per-entry
and total size, path depth, escape and symlinks.
Which container it is comes from the first four bytes (rather than the
name). A pod published as a plain .zip reads as it always did, a
.sisupod reads, and either one renamed reads too. The suffix is still
recognised, both spellings, for the argument and for a url.
libzstd is declared rather than bound: provides the whole surface of
fifteen extern C prototypes (there is nothing to generate and no
upstream tree to track, which is the arrangement sqlite3 already has).
dub links it with "libs": [ "zstd" ]; nix needs zstd.out rather than
zstd, whose default output is the binaries and carries no library at
all.
A frame declares its uncompressed size in its own header, and for a pod
fetched over https that number is attacker controlled. The declared size
is checked against a ceiling before a buffer is asked for, a frame that
will not declare one is refused, and what comes out is checked against
what was promised. The ceiling is the extraction limit the archive
reader already applies, so the two bounds agree.
Measured: output built from a .sisupod is byte identical to output built
from the same pod's .zip, 359 files over text, html, epub, odt and .ssp
for three documents, live-manual's ten languages included. A truncated
frame is refused and the document skipped.
free_culture 2863504 -> 1099736 2.60x
the_wealth_of_networks 4294022 -> 1172649 3.66x
live-manual 6972418 -> 675970 10.31x
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Three faults in four lines of the odt metadata.
meta:creation-date and dc:date were taken from generated_time, the run's
own clock, so two runs over unchanged markup produced different odt
bytes. They now carry the document's own instant, so an odt is a
function of its source. That was the last clock stamp in any writer:
every output format is now reproducible.
generated_time is a human-readable run stamp, of the form
"2026-9-17 [38/3] 18:39:35", with unpadded numbers and an iso week in
brackets. It is not a valid xsd:dateTime, which is what ODF asks for in
both of those fields, so what was written there was malformed as well as
unstable.
dc:language now carries the document's own language tag.
The instant is doc_matters.modified_utc, moved there from epub3.d
in this commit: epub3 needs the same value for its dcterms:modified,
and one definition cannot drift from the other. generated_time now
has no caller and is left in place.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
| |
Spine now looks for conf/document_make, and bundles that name.
(the dr_ prefix was a sisu-era name to distinguish it)
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Spine names a language as the directory under media/text/ is
named, which for a region is the posix form with an underscore,
pt_BR. A lang, xml:lang or dc:language value is a BCP 47 (RFC
3066) tag and wants a hyphen, pt-BR. Spine put its own code
there unchanged, so a document in a region language produced an
invalid attribute in every one of its xhtml files: 138 epubcheck
errors for a document of ten segments and a navigation file.
src.language_tag is the same language as a tag. The attributes
and the dc metadata now read it. Output paths keep src.language,
since a path is named after the directory, not after the tag.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
bugfix: content image names are now checked in a pass of their own,
before anything is created or written, and one bad name refuses every
image in the artefact rather than skipping the one, which is what the
zip reader does with a zip. The write then re-checks containment.
(prior to this an image name read out of a .ocda.db was written without
checks, and such a database can be fetched over https. A name could
climb with "..", and being handed to chainPath an absolute name dropped
everything before it, so "/x" did not land under the image directory at
all. The zip reader has guarded against this since it was written; the
database reader did not).
carried_names.d holds both rules: a bare filename, as an image is
carried, and a relative path, for the markup and conf a database is to
carry.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A file beginning with a utf-8 byte order mark had the mark read as part
of its first yaml key, so "title:" arrived as "title:" and every
value under that key was dropped silently: the document kept its text
and lost its title. Downstream that showed as an empty dc:title and an
empty <title> element in the epub, which epubcheck reports as two
errors.
The mark is now kept out of the header and body split, and so out of
the parse. It is not stripped when the file is read, because that text
is what source.digest is taken over and the digest names the file as it
sits on disk.
One document in the sample set begins with a mark. It is left as it is:
it is the case this guards against.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Four things, not visible in an error count:
- Images had no alt text, even where markup carried the text, and the
field alt text is present, incorrectly used in code; fixed as the alt.
{ sm_tux.png 64x80 "Gnu/Linux - a better way" }image
an image with nothing to say for itself gets alt=""
- Images with no dimensions were incorrectly written width="0"
height="0", fixed
- Accessibility metadata, which EPUB Accessibility 1.1 requires:
schema:accessMode, accessibilityFeature and accessibilityHazard, plus
accessModeSufficient and accessibilitySummary, which epubcheck does
not test for; used by
- readers looking for a book it can use,
- anyone distributing into the EU.
Every value is derived from the document, not asserted: a document
with no images says so, one whose images all carry alt text claims
alternativeText and textual sufficiency, one where an image says
nothing claims neither.
Nothing claims WCAG conformance, which would be a claim about an
evaluation that has not happened.
- DPUB-ARIA roles beside the structure spine already knew about.
doc-toc, doc-endnotes, doc-bibliography, doc-glossary and doc-index on
the sections, doc-noteref on a note reference and doc-footnote on the
note it points at. "doc_endnotes" as a class name means something to a
stylesheet and nothing to a reading system. Not doc-biblioentry: DPUB-
ARIA 1.1 deprecated it. These are plain ARIA, so the html gets them
too.
Also dealt with here:
- Spine makes certain sections such as endnotes, bookindex, ... and a
heading that has a name for itself cannot share the same name, if it
does the name is now disambiguated, disambiguated. (disambiguation
similar to that used for a repeated anchor tag)
- Endnote anchors fixed to use "id" (being one unique tag in the
document that is target of every reference to it)instead of obsolete
form (they were <a name="note_1">. "name" on an <a>).
Book index markers keep "name": the same marker sits on every object
an entry names, 296 times in one document, and those are markers of
where a term occurs rather than distinct targets.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
epubcheck 5.3.0 over the 35 sample epubs went from 1302 remaining errors
to none. The six causes for those remaining errors fixed here.
most of the count was one of them:
- Tables (1123 errors, 86%) where the writer emitted obsolete
attributes, removed in 2014; table also now in div instead of a
paragraph. also fixed:
- Navigation (115). Every level 4 heading's nav entry ended "#0", the
ocn of an object that has no ocn and so no id.
- Duplicate ids (37). (from three unrelated sources).
- The publication identifier (27 warnings). dc:identifier was a
hardcoded hex string, not a UUID, and the same one in all 35 files,
so two documents in one library collided. It is now UUID v5 over the
document's own uid: a real RFC 4122 identifier, stable across builds
and distinct per document.
- Image manifest entries (15). media-type was "image/" plus the file
extension, giving "image/jpg", which is not a media type; epubcheck
reads it as a foreign resource and wants a fallback. There is now a
mapping, which also covers svg and webp. Item ids came from the file
basename, and "2bits_02_01-100.png" gave an id starting with a digit,
which is not an XML name; ids are now prefixed and sanitised.
- An ocn resolving to no segment (9). A poem block and its first verse
share one ocn; the verses are written into a segment and the block
itself is not.
Also dropped from the package document: an xmlns:xsi that nothing
used, and a prefix declaration for the rendition vocabulary, which is
a reserved prefix and needed no declaring. The package now carries
xml:lang.
test-epub-validity.sh now allows no fatal and no error. Two NAV-011
warnings remain and are described in it. Three reference .ssp files
change, all of them the anchor tag disambiguation above.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
| |
in error regex permitted a url to go past an endnote close delimiter,
fixed
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A URL ending in .ocda.db is now downloaded and processed in place,
through the path that already did it for .zip: the same curl call, the
same size and timeout limits, the same refusal of local and private
addresses, the same --allow-downloads guard, the same temp file cleanup.
Only the pattern had to widen.
rgx_url_zip ^https?://...[.]zip$
rgx_url_source ^https?://...([.]zip|[.]ocda[.]db)$
downloadZipUrl is no downloadSourceUrl and the download temp directory
is spine-download rather than spine-zip-pod; (the extraction directory
keeps its name, being still only for zips).
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The document markup sample collection builds byte identically from
either artefact: 1457 files, no differences, from markup, from the 35
.ssp files, and from the 35 .ocda.db files.
A .ocda.db carries its images as blobs. dbReadFiles() reads them and
they are written where the output writers look for images: a
directory of this run's making, removed when the run ends, not the
pod beside the artefact, (which could be a working published tree). The
document's source path moves with it, image_dir_path being reached from
the document's own file rather than from the pod.
Each blob is checked against the sha256 the database recorded beside it.
The two were written together, so a mismatch means the file is damaged.
(the point of having recorded the digest).
A .ssp only describes its images, so those are the ones in the pod it
sits in, and they are checked against the digests it recorded.
That asymmetry arises from of what the two artefacts are: the ocda.db is
self-sufficient and can prove it is looking at the right image, the .ssp
describes a pod it sits alongside.
Found while wiring it: setting the pod directory alone was not enough,
since image_dir_path is derived from the source file's path
(../../image from media/text/<lang>/), so the source path has to be
moved with it.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Process a document from its .ocda.db or .ssp abstraction. A .ssp or a
.ocda.db given as an argument is now a document like any other: read
rather than parsed, and handed to the output writers as the same doc
they already take.
The 35 document sample collection built from its .ssp files is byte
identical to the same collection built from markup: 1457 files, no
differences.
Three things had to be rebuilt rather than read, all of them a value
parsed rather than a structure derived:
- classify_topic_register_arr and its expanded twin. The split is not a
plain one, so the rule moved to sisudoc.ocda.meta.topic_register and
both the yaml reader and this one call it.
- creator_author_arr, which is creator.author split on the ", " it was
joined with.
- title_sub, a copy of title_subtitle made where the header is read, and
what epub3 puts in dc:title id="subtitle".
Fixed ordering bug found by the acceptance test. A heading's own anchor
can be a bare number taken from its text and this can collide with the
ocn of an unrelated object. Whole output comparison found that before
the fix there was one wrong link in one epub's table of contents.
From a .ocda.db, 1446 of 1457 files are identical. The 11 that are not
are the images: a database carries its own image blobs and nothing yet
extracts them, so the five sisu_markup images are not copied and the
epub that embeds them differs. That is the next step and is not a defect
in this one.
--source and --pod2 are refused with a warning rather than half
done, no artefact carrying the markup.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
ST_DocumentMatters was declared inside spineAbstraction(), closing over
the parser's locals, so only the parser could build it. A document
loaded from a .ssp or a .ocda.db has to arrive at the same value from
what the artefact carries, and had nothing to build.
It is moved to sisudoc.ocda.meta.doc_matters as docMattersMake(), taking
the seven things a document is described by: the run (program_info,
opt_action), where it is and what it is called (manifest), its header
and the site config (conf_make_meta), its counts and indexes
(ST_DocHas), its source digests, and the files its markup inserted. The
parser fills them from what it has just parsed; the loader will fill
them from the artefact.
Pure refactor: the struct's members are unchanged, only where it is
declared and where its inputs come from.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The segment a cross reference to a heading lands on, gixed upstream in
ocda, where the field is created, rather than inferred by the loader.
A heading above level 4 opens no html segment of its own, so a link to
it has to land on the level 4 heading that follows. The build loop
cannot know that when it reads the heading, and says so in a comment of
its own: "for html segname need following lv4 not yet known". It
back-fills tag_assoc when the level 4 heading arrives (lv0to3_tags), and
the answer then lives only in that map.
tags.segment_lv4_is now holds it, resolved in a pass over the finished
head and body sections: walking backwards, the last level 4 heading seen
is the next one for everything above it, which is the same answer the
back-fill gives. Emitted in the .ssp only where it differs from
.segment_html_is, which is every heading at level 4 or below, so it is
sparse: 317 lines over the 35 document reference set, none removed.
Carried in the database as a column of its own, and read back by both
readers.
docHasFromAbstraction reads it instead of working it out. With that, and
on top of the two defect fixes, a document loaded from an artefact and
one parsed from markup agree:
key sets identical, all 35 documents
values 3 entries differ of some 30,000, all of them
_the_title, where the abstraction supplies an epub
segment the parser leaves unset
link targets all 1,651 agree, against 89 differing before any
of this work and 1 after the defect fixes alone
Format stays v1.1. That version is new in this same run of work and
nothing outside spine has read it, so this belongs in it rather than in
a bump of its own.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
docHasFromAbstraction() builds ST_DocHas from the objects and the
@doc_has block, so a document read from a .ssp or a .ocda.db can reach
the value the parser reaches. Everything it builds is an index over
properties the artefact already carries, not a recomputation of
something it does not: the counts come from the header, imagelist from
the .image records, the segment name lists and the tag associations from
.anchor, .segment, .segment_epub, .heading_lev_anchor and
.segment_*_is, and section_keys_sequenced from which sections are
non-empty plus the run's own flags.
*The two defect fixes this now sits on did most of the closing.*
Measured before them and after, over some 30,000 tag_associations
entries and the 1,651 keys a document actually links to:
imagelist 8 documents differed, now none. The parser was
the one that was wrong, and rgx.image being
anchored to image markup brought it into line
with what the .image records always said.
key sets four keys differed (_part_eof, "0", the empty
key, and "toc" the other way about), now none.
_part_eof came back the moment @tail was
carried, and the rest with it.
values 89 link targets resolved differently, now 1.
The one left is free_culture's ocn 5, and it is the case this cannot
reach: a heading above level 4 takes its segment by back-filling when
the next level 4 heading arrives, so the object never holds the answer
and reading the objects in order only approximates it. Fixed properly in
the commit that follows, in ocda where the field is made, rather than
guessed at here.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
ST_DocHas in metadoc_object_setter, filled by the parser, in place of
DocHas_, a struct of accessors nested inside docAbstraction() that
closed over the parser's locals. The member names are the ones the
output writers already use, so nothing downstream changes.
The point is that a struct closing over parser locals can only ever be
built by the parser. A document loaded from a .ssp or a .ocda.db has to
arrive at the same value with only what the artefact carries, and now
there is something for it to fill.
Two small things fall out. imagelist is a string[] rather than the lazy
uniq range it was, and images() is therefore imagelist.length rather
than counting the commas in the range's string representation, which
carried a "TODO not ideal rethink". Same answer unless a filename
contains a comma; the reference test agrees on all 35.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
reading a .ocda.db was slower than parsing markup source. Taking the
largest markup document sample War and peace, 12,135 objects, optimised
build, before this commit and after:
parse markup 0.87 s 0.87 s
load .ssp 0.047 s 0.047 s
load .ocda.db 1.09 s 0.128 s
The database goes from being 1.25x slower than parsing the document to
6.8x faster.
Two things fixed in the reader were:
First, four queries per object. object_images, object_links,
object_anchors and object_subtoc were queried per object as the objects
were built, each statement compiled fresh from a concatenated string.
For war and peace that is 48,540 statement preparations to collect
138 rows, which is all those four tables hold between them. They are now
four ordered sweeps, kept by object id, so the cost is what the tables
hold rather than what the document holds.
Second, and even more consequentially: d2sqlite3's row["name"] is
indexForName, a linear scan over the statement's columns that calls
sqlite3_column_name and allocates a D string for every column it passes.
At some fifty named reads per object over forty columns that is around a
thousand of those per object, twelve million for the document. The
column name to index map is now resolved once per statement and the
reads are an integer index.
Also here, since it was measured while doing this: sspReadFile no longer
hashes the file it read unless asked (with_digest). Only the database
writer wants that digest and it has the lines already, so every other
read was paying for it.
Existing outputs unaffected.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
.ssp format v0.1 -> v1.1, and the .ocda.db with it: they are two
serialisations of one format and now move together.
The gap this closes, measured against the writers' read set: of the 59
conf_make_meta fields outputs/io_out/ reads, 29 were in neither
artefact. Five are conf.* and stay out, being site and run scoped (urls,
papersize, the search db filename): the same abstraction published to
two sites must take each site's. Of the rest, three are never assigned
anywhere (title_short, publisher, original_publisher; read only by
sqlite.d, always empty) and one is a copy of a field already carried
(title_sub = title_subtitle), so 17 properties actually had to travel
and now do:
make breaks, footer, home_button_text
meta title.edition, date.added_to_site, language.document_char,
original.{title,source,language,language_char},
rights.copyright_{text,translation,illustrations,
photographs,cover,audio,video}
make splits by when it acts, which is worth keeping in mind: these three
are read in outputs/io_out/ and must travel, while italics, bold,
emphasis, substitute and headings are read at parse time and their
effect is already in the objects.
New @source block, and source.* rows in the database, so a reader can
say which markup an abstraction came from rather than working from
stale content in silence:
language the document's own
languages the pod's list, which is what the inter-language links
in html need and neither artefact carried
digest sha256 of the .sst; equals its digests.txt entry
The database adds source.ssp_digest, the sha256 of the .ssp it was built
from, since a file cannot hold its own hash. So the chain
.sst -> .ssp -> .ocda.db is checkable end to end.
The database's metadata table is now filled from the header blocks the
.ssp gives back, not from a second list read off doc_matters. One list,
in ssp.d: a property added there arrives in the database with nothing
else changed, and one that is not in the .ssp cannot be in the database
at all. That was the last place the two could drift; the objects stopped
being able to on 2026-09-07.
Version is checked on load. The major part must match, a newer minor is
accepted (a minor bump only adds properties, and an unknown property
line is ignored). A v0.1 artefact is now refused with a message saying
to regenerate it, rather than loading half populated.
test-abstraction-db.sh now compares the two artefacts' header blocks
property by property, 36 per document on the wealth of networks, where
it previously only checked that a schema.version row existed. Verified
non-vacuous: dropping one property is reported.
Reference regenerated, and the whole diff is this change and nothing
else: 320 lines added, 35 removed over 35 files, being 35 format lines
changed, 35 @source blocks (4 lines each), 35 language.document_char, 35
home_button_text (it has a default), 18 breaks, 16 footer,
4 date.added_to_site, 1 title.edition, 1 original.source.
Existing outputs unaffected.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
| |
The abstraction has nine sections and includes "tail". Both artefact
writers hardcoded a list of eight, so "tail" was silently dropped.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
| |
correct circumstance where regex was incorrectly able to match image
shaped words as well as identified and marked up image
(internal markup).
take all matches.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
abstraction directory cleared once, by whichever language is first
The .ssp writer clears stale files out of
pod/<doc>/media/abstraction/, and that one directory is shared by
every language of a document. The clearing is now done once per
directory per run, by whichever language reaches it first, with
the lock held across it. A language that finds the directory
already prepared has passed through that same lock before writing,
so the clearing it skipped had completed before its own write
began: no .ssp produced on a given run can be removed during it.
(removes possibility of a race condition on parallelisation)
(assisted by Claude-Code)
|
| |
|
|
|
|
|
| |
The metadata page gains a line between the markup source and the
source digests.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
| |
ocda (object centric document abstraction)
flag --ocda-db (or --abstraction-db)
replaces --show-abstraction-db
rename results in consequently large diff
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
pod/ holds all document source representations:
pod/<doc>/ source tree
pod/<doc>/media/abstraction/<uid>.ssp abstraction, as text
pod/<doc>.zip tree, zipped, .ssp included
pod/<doc>.digests.txt sha256s of what is in them
pod/<uid>.ocda.db abstraction, as sqlite db
<lang>/abstraction/ is gone.
pod/<doc>/media/abstraction/<uid>.ssp preferred as having the
images (found within the pod tree) which .ssp needs to reproduce a
document but does not carry on its own.
The .ocda.db sits carries the images as well and (like the
pod.zip) can be used to reproduce a document directly.
digests.txt now covers the database as well as the zip, the source
and the .ssp; (as does the metadata html page).
Two ordering issues addressed:
- the pod builder clean-slates pod/<doc>/ before regenerating it,
and the .ssp is now written before that runs. The clean slate
now leaves media/abstraction/ alone, and the .ssp writer clears
that directory itself on the first language of a run, so a .ssp
for a language the document no longer is removed and cannot be
bundled.
- for a multi-language document the .ssp files accumulate one
language at a time and are bundled on the last, which is why the
directory cannot simply be emptied by whichever (language) gets
there first.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
sisudoc.ocda.abstraction.load names the five things a document can
be read from, tells them apart, and loads the two that are
self-describing artefacts:
.sst / .ssm + images the markup source
pod (dir) + images the same, bundled
pod .zip the same, zipped
.ssp + images the abstraction, as text
.ocda.db the abstraction, sqlite, images inside
abstractionSourceOf(path) is the detection, by name and for a
directory by whether it holds pod.manifest. abstractionLoad(path)
returns a LoadedAbstraction: the source kind, whether it was
loaded, why not when it was not, and the document itself.
The three source forms are deliberately not loaded here. Reading
them is the parser's job (sisudoc.ocda.meta.metadoc
spineAbstraction) and it needs the manifest, environment and
configuration that spine.d assembles, none of which belongs in a
loader. What this gives that case is the dispatch and a plain
statement of where it is handled, rather than a silent empty
result.
spine --abstraction-source=<path>
says what a path is and, for an artefact, loads it and reports
what came back: the document, title and author, header block
sizes, object counts and the objects in each section. Exit 0 when
an abstraction was loaded, 1 when not.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The per document database is now written as <doc_uid>.ocda.db
rather than <doc_uid>.abstraction.db. Shorter, and it says what is
in the file: the object centric document abstraction, not "some
abstraction". .ocda.db pairs with .ssp and cannot be mistaken for
the collection search database (spine.search.db).
test/run-tests.sh runs the four in sequence, one line of result
each, and a summary. It re-runs itself inside nix shell
"nixpkgs#sqlite" if sqlite3 is not on PATH, so this is all that
is needed:
SpinePOD=../../markup/sisudoc-spine-samples/markup/pod-samples/pod \
./test/run-tests.sh ./bin/spine-ldc
The tests are independent (with test-abstraction-ssp.sh run
first):
1 test-abstraction-ssp.sh is first because it is the
one that says whether the abstraction itself moved; if it
fails the others are answering a different question than you
think
2 test-abstraction-ssp-roundtrip.sh reads the committed
reference set, so it is a statement about the current binary
only if 1 passes
3 test-abstraction-db.sh and
4 test-abstraction-db-roundtrip.sh generate both artefacts
themselves and depend on nothing committed
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The ocda database is now built from the .ssp itself: the writer's
lines are emitted, read straight back by ssp_in, and the objects
that come out populate the database. Anything the .ssp does not
carry, the database will not have either, by construction (rather
than by test).
[instead of as previously through a second walk over the in-memory
abstraction]
- spineAbstractionTxt is split: sspDocumentLines(doc) returns the
whole .ssp as lines, and the file writer emits them. Output
neutral, the reference test confirms.
- spineAbstractionDb takes the abstraction as an argument rather
than taking doc.abstraction.
- sspRoundTripAbstraction(doc) in ssp_in is the join: lines out,
lines in, abstraction returned. Both call sites in spine.d use
it.
- the header blocks and the image blobs still come from
doc_matters (as: the .ssp does not carry image bytes).
All (35) markup sample sourced databases built through the .ssp
have byte identical SQL dumps to the one built directly before the
change.
That comparison also found one reader inaccuracy, which the .ssp
round trip could not see because the writer omits the field either
way: an absent identifier was restored as the ocn in every case,
but for an object with no ocn it was empty ("a"~N identifiers are
always written). Fixed; the two artefacts checking each other is
what caught it.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
sisudoc.ocda.abstraction.db_in reads a <doc>.abstraction.db back
into ObjGenericComposite[][string], the same value ssp_in returns
from a .ssp, so a consumer need not know which artefact it was
handed. It mixes in the .ssp reader for that shared document
struct rather than declaring a second one.
--db-round-trip=<file.abstraction.db> reads a database and emits
it as .ssp on stdout, through sspObjectRecord as the other round
trip does. Held against the .ssp written from the same document,
this says whether the two artefacts really carry the same thing:
not a count of fields, as test-abstraction-db.sh does, but the
whole document reconstructed from the database and compared to the
text.
SpinePOD=... ./test/test-abstraction-db-roundtrip.sh ./bin/spine-ldc
PASS: all 35 databases re-emit their document's .ssp exactly
It passed on the first run over the whole sample set, which is evidence
that the database is now field-complete against the .ssp rather than
merely counting the same.
Four tests with different checks:
test-abstraction-ssp.sh the abstraction has not changed
test-abstraction-db.sh the two serialisations agree, field by field
test-abstraction-ssp-roundtrip.sh the .ssp can be read back whole
test-abstraction-db-roundtrip.sh the .db can be read back whole
(assisted by Claude-Code)
|
| |
|
|
|
|
| |
heading text used for navigation is normalised, and | escaped
(assisted by Claude-Code)
|