| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
|
|
| |
- adjust metadata, cleanup fixes
- no <h0> realign title on h1 along with bespoke css
- show group bullets
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
- A poem is a container of verse, and stores its range of verse ocn,
- its verse are the citable units with ocn.
- A note in the last verse of a poem previously was not gathered into
the endnotes section, this now is fixed
Format 2.0 -> 2.1: the property is an addition, the poem is not a
citable object (but contans a range of objects), its verse are
(individual citable objects), and every reader checks the major version.
A 2.0 database reads as having no ranges.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
search database: show notes other than regular numbered (foot)notes.
make sure all types of notes for all object types set the objects
has.inline_notes_reg flag required for them to be rendered
search results' html (doc_objects.body) show notes only where an
object's has.inline_notes_reg flag is set, which occurred for numbered
notes only
- only numbered notes, regular footnotes were included
- and notes in some objects were omitted (quotes and verse)
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
| |
The markup documents ~[* note ]~ and ~[+ note ]~ as editor's notes, each
a separately numbered series, and a bare ~[ note ]~ which sisu put in
the asterisk series. They are included as notes. Each series is numbered
through the document, *1, *2 ... and +1, +2 ..., apart from the author's
notes.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
| |
- meta_processing_general() in spine.d replaced by per-stage predicates.
- doc_matters.generated_time() commented out, unused since the odt and
epub clock stamps were removed.
(assisted by Claude-Code)
|
| |
|
|
|
|
| |
--version prints the: version, compiler, and platform, and exits.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
| |
unrecognised options are named on stderr before any work is done, and
--strict makes such fatal (exit 1). (prior to this commit a mistyped
or retired flag was ignored without acknowledgement). This applies to
options only, an argument that is not a source, such as a README beside
pods being processed, is still passed over quietly.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A plain-text output that keeps the object numbers, written as CommonMark
with GFM pipe tables, one .md per document-language in <lang>/markdown/,
(linked from the metadata page) with images into that shared directory
itself, skipping any already there, so --markdown alone still gives a complete tree.
provided as an alternative to --text rather than a replacement
--markdown produces a document that renders, (both provide objext
numbering survives both. Every object carries an anchor and a
superscript number linking to itself, so a citation by object number
is a link in any markdown renderer.
--text may still be of value to read in a terminal
A note is written where it is referenced, with a link back to the
object it came from, rather than gathered into an endnotes
section.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
typst provides a new pdf build path replacing latex as the primary pdf
builder.
--typst (--typ) writes the .typ;
--pdf writes the .typ and compiles it by spawning `typst compile`.
(--pdf is redefined, being typst to pdf rather than latex).
--latex is unchanged and writes the .tex;
One .typ per document-language (not one per paper size). The paper and
the orientation are read from sys.inputs with a default, so the same
file compiles to every paper --set-papersize asks for, the pdfs keep
their original names.
Spine's paper names (derived from what suited latex) and typst's are
different vocabularies and are mapped; a paper typst does not have is
named and skipped.
With the flag --pdf spine invokes a typesetter to generate the pdf
output directly, which is new (the latex path writes a .tex and leaves
xelatex to the caller). If typst is not available on the machine it is
reported absent, naming the binary and printing the command, and the run
continues, the .typ is written.
Several useful features of an object-centric pdf are easily met, usually
being a few lines each.
- The object number goes in the margin as one `place` at a coordinate
inside the object's own block, carrying the label every citation
points at.
- A note is a real footnote at the point of reference, with the
document's own number and one link back to the object.
- The table of contents locates by object number rather than by page, so
it holds true across every paper and every setting.
That typst compiles silently across the sample set means every internal
link in every document resolves (--pdf over the collection with two
paper sizes gives 144 pdfs, none skipped and none failed, each the page
size its name claims and each tagged).
paths: .typ, pdf named in same way as .tex based built pdfs so site
linkage works either way
flake.nix: added a typst dev shell, for .typ pdf typesetter
- dsh-typst-pdf providing a pdf typsetter for spine's .typ.
- dsh-latex-pdf provides xelatex for .tex which can still be written.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
flags names tidied so that each names the artefact it acts on.
--ocda-db and --abstraction-db do the same thing. --ocda-db is the
official flag (--abstraction-db remains but undocumented)
--db-round-trip becomes --ocda-db-round-trip, so that it agrees with
--ssp-round-trip: each names the artefact it reads rather than one
naming a container and the other a format.
--abstraction-source is untouched named after he stage as it takes .sst,
a pod, .ssp or .ocda.db, rather than any one artefact.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
spine automatically builds some segments, including: (toc, endnotes,
glossary, bibliography, bookindex, blurb, _the_title) these are now
identified as reserved names and a user is now warned if any of these
names have been manually assigned to a heading by markup. A document
still builds but to disambiguate the ocn of the heading is attached to
the markup (reserved) name and this is seeded before a document not
after. It is read off the finished abstraction rather than reported by
the parser, and beside the ocn alignment check for the same reason: it
is a statement about a document rather than a step in building one.
WARNING reserved segment name: the_autonomous_contract... [en]
heading 1 at ocn 135 asks for "endnotes", which spine gives its
own generated section
it is named "endnotes-135" instead; ...
A warning: the document is correct and complete and the name it ends
up with works.
--strict makes it a failure for the run, as it does for ocn alignment,
and by the same reasoning: the outputs are written and can be looked at,
and the exit status is taken at the end.
Two of the thirty-six sample documents have reserved segment names,
"1~endnotes" heading. Output is unchanged: nothing here touches the
abstraction.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
need a character to split filenames on in certain circumstances.
there problems with use of a colon in filenames, which is legal on posix
but not on Windows."~" fits the bill better being legal on every
filesystem of interest and is unreserved in rfc 3986, needing no escaping in a url.
The reference abstraction is renamed, its content unchanged. Over
the sample collection two filenames move and nothing else does: the
abstraction and database digests are identical, the archive's
members and their sizes are unchanged, and only the member order
shifts, "~" collating after letters where ":" sorted before them. The
document's epub dc:identifier is a v5 uuid derived from the uid, so it
changes for those filenames.
Every document's uid changes in the search database (spine.search.db).
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
--help now names the five source forms before the options, and after
them states the contract: what a .ssp and a .ocda.db each carry and do
not, that both record the digest of the markup they were built from,
which five actions need that markup and what happens when they cannot
have it, and the three ways to verify.
And the version, 0.24.1 to 0.25.0, for abstraction format 2.0 and what
it made possible: (one database per document holding every language of
it, carrying the markup, images, configuration, manifest and
translation catalogues it was built from, with the implications that
carries)
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
ensure that db uses the names it carries in manifest which produce
deterministic output (and fix divergence in db rendering of output
names, (which previously also looked for variable input from the
environment))
Output built from a database with --config naming the site configuration
is now byte identical to output built from the pod, on the render route
as it already was on the materialise route.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
dr_document_make in three identifier spellings becomes document_make,
which is what the file has been called since the rename.
The mixin import audit, a template declares its imports at template
scope and they are then visible in every scope it is mixed into (which
is how a template-level split once hijacked a UFCS lookup in code that
had not been touched). Disabling htmlSnippet's imports and rebuilding
names the consumers that were relying on them rather than on their own:
across twelve mixin sites, exactly one. metadata.d now imports the `to`
it uses. The templates keep their imports, their own functions needing
them, and narrowing those is a change of its own rather than a cleanup.
And the standing FIX in source_pod.d: the insert digest line for a non
pod source recorded the insert's path, where every other line in
digests.txt is a filename and a digest naming a file that is not in
the pod under that name cannot be checked against it.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
does this spine still produce this abstraction?
A database states an abstraction and carries the markup it was built
from, so re-parsing the one and comparing against the other says whether
the two still agree. Checked as follows:
- automatically and unskippably when a document is built from a
database, at the moment it is parsed and before anything has been
written from it.
- on --ocda-verify=<file>, an action of its own that materialises,
parses every language, compares, reports and writes nothing, the exit
status being the answer.
- on --no-verify as the escape, warning per document, for a newer spine
reading an older artefact where the abstraction is expected to differ
and the output is wanted anyway.
A build ends at the first failure rather than skipping the language.
A document rebuilt from a database now has its whole abstraction built.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
actions requiring source first materialise the pod
One route per argument, decided by what was asked for. --source, --pod,
--pod2, --show-abstraction and --ocda-db need original markup, so an
.ocda.db given with any of them is written back out as a pod and parsed.
Asked only to render, it is read as an abstraction as before.
To ensure consistency, where a pod is materialized by an argument, that
pod is used for the whole run, including for html and epub that could
have been built from the loaded abstraction instead.
--show-abstraction and --ocda-db join the markup side deliberately.
They could be satisfied by re-serialising the abstraction already
loaded, and were. But an artefact is a statement about the source:
written from a loaded abstraction it says only that the loader is
self consistent, where written from the markup it says what this
spine (whatever current version) makes of that document today.
The chain is checkable against itself: --ocda-db from a database now
goes database, pod, parse, database, and returns the same 9,908,224
bytes it started from. --show-abstraction likewise re-emits the .ssp
files byte for byte.
A .ssp carries no markup by construction and a database written
before the format carried it has none, so both are refused for those
actions, once and by name, and the rest of the run goes on.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
actions requiring source first materialise the pod
One route per argument, decided by what was asked for. --source, --pod,
--pod2, --show-abstraction and --ocda-db need original markup, so an
.ocda.db given with any of them is written back out as a pod and parsed.
Asked only to render, it is read as an abstraction as before.
To ensure consistency, where a pod is materialized by an argument, that
pod is used for the whole run, including for html and epub that could
have been built from the loaded abstraction instead.
--show-abstraction and --ocda-db join the markup side deliberately.
They could be satisfied by re-serialising the abstraction already
loaded, and were. But an artefact is a statement about the source:
written from a loaded abstraction it says only that the loader is
self consistent, where written from the markup it says what this
spine (whatever current version) makes of that document today.
The chain is checkable against itself: --ocda-db from a database now
goes database, pod, parse, database, and returns the same 9,908,224
bytes it started from. --show-abstraction likewise re-emits the .ssp
files byte for byte.
A .ssp carries no markup by construction and a database written
before the format carried it has none, so both are refused for those
actions, once and by name, and the rest of the run goes on.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The pod carries tools/po4a and the database (until now) did not, so a
pod written back out of a database came back a lossy copy, without its
translation catalogues.
the database now carries them, walked and name-checked exactly as the
pod writer walks and checks them, once on the last language. They are
compressed to save space, the test sample carrying 6.25 MB of catalogue
as read would have made that one document's database two thirds larger.
Files gains a compression column: NULL or 'none' for the bytes as read,
'zstd' for a frame. What to compress is decided by role, not by size
(source, conf, manifest and tools are text and are compressed; images
being compressed already are not). The same kind of file is always
stored the same way. Document objects are untouched and stay raw:
objects_fts is external content over objects.text and reads that column
directly.
bytes and sha256 remain those of the original file, so every digest
check, digests.txt line and source.digest rebuilt from these rows
works unchanged, and a reader that does not decompress can still say
what it is looking at.
dbReadFiles decompresses, so no caller learns how a blob is stored. It
asks pragma_table_info whether the column is there at all: a database
written before it reads as raw. A row that claims zstd and will not
decompress is named and left out, rather than handed back as a frame
where markup should be.
live-manual: 9,490,432 bytes without the catalogues, 9,908,224 with
them. Output from the materialised pod is identical to output from the
original pod, 679 files each side, no differences.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A database that carries its source can recreate the pod. --source and
--pod2 given a .ocda.db now write the pod to a directory of the run's
making and carry on with the pod, so everything after that point is
handling a pod like any other.
Nothing renders from the database. A materialiser that also rendered
would be a second path to every output format and the two would drift;
one that only writes files means the document is built by the same
code over the same bytes as the original, and identical output is a
consequence rather than an aspiration. Held against the original pod,
site configuration constant: every output file identical across ten
languages.
The pod's name comes from the database's filename, which inverts the
naming rule exactly: <doc>.ocda.db is named by doc_uid_out_no_lang,
the pod name and the document's filename joined by ":" when they
differ and the one name when they do not. So the materialised pod
recomputes the uid it was named by and every output file lands on the
name it had.
Names are checked before anything is created, and one bad name refuses
the artefact rather than skipping a file, as the zip reader does with
a zip. Markup that does not match the digest stored with it is refused
outright, where a mismatched image is written with a warning: a wrong
image makes a document that looks wrong, a wrong markup file makes one
that is wrong, in its text, with nothing downstream to notice.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The database already carried the images, because without them no output
can be produced. It now carries the markup as well, and with it the
document's configuration and the manifest as the author wrote it.
That is the difference between a serialised abstraction and a document
source. An abstraction can be rendered but not re-parsed, and a reader
who wants to correct a sentence needs the sentence as written. With
these a pod can be written back out of the sqlite-file.
Three new values in the existing role column, no schema change: the
format stays 2.0.
Source rows are named by their path within the pod,
media/text/<lang>/<file>, and not by bare filename as images are. Every
language of a document has a file of the same name, so bare names would
collide under UNIQUE(role, name) and nine of ten would be dropped
without a word. The path is also what a pod materialised from this
database has to be told.
Bytes stored as read with the digest over them, so that each row is
checkable against the line source.digest was built from.
Also build.spine_version, a file level row saying which spine wrote the
file: not a property of the document, and the one thing a file cannot be
asked for afterwards. The reader skips the build. prefix as it skips
schema. and translation., or it would come back inside the document
header and the round trip would differ.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The check runs once every language of a document has been abstracted.
Providing a warning by default. --strict makes divergence a failure,
taken at the end of the run so that outputs are complete and can be
examined. (the name --strict is general, so that later checks can be
added without a second flag).
Under --parallel the profiles are appended under synchronized, and the
comparison sorts a document's languages by name rather than taking
them in the order the threads finished, so two runs print the same
lines in the same order.
The outcome is noted per language in the database as
translation.ocn_aligned, for documents that have more than one
language.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Each language leaves a profile as it is abstracted: how many numbered
objects it has, and what kind of object each ocn is. The comparison is
against the language the manifest names first.
A count difference is reported alone. Past the first dropped or added
object every ocn names something else, and a kind comparison after it
would print hundreds of lines that are all the one fault.
The languages of a document share their object numbering, this being
the basis of ocn citation in a multi-language document: an ocn
names the same object in every language. Nothing enforces it.
It holds where translations follow the source object for object, and
stops holding otherwise (e.g. the moment a translator drops a paragraph,
merges two, or turns a heading into a sentence, at which point the
numbering no longer matches).
An ocn is not unique: a poem and its first verse share one, by design.
So the walk is in document order, the section order the .ssp is
written in rather than the abstraction's key order, which an
associative array does not promise, and the first object at an ocn is
the one recorded.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
| |
One file for the document (including all languages) with digest of the
document. digests.txt lists it once, beside the shared images, taken
when the last language has been written.
digests.txt carries all digests.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A database now holds every language of its document, so reading one
means naming a language. Three small pieces:
- spineDbLanguages says what languages a file holds. Its own template
rather than part of the reader (asking a file what it holds does not
instantiate the object setter and the markup regexes with it).
- abstractionLoad and spineDocFromArtefact take a language and pass it
to dbReadFile, which resolves it to a doc_id, takes that language
document where there is one, and lists the languages and stops where
there is a choice.
- spineArtefactLanguages turns an artefact into the documents to build:
one for a .ssp, and for a database the languages it holds filtered by
--lang. The artefact loop then has a language loop inside it, so a
single .ocda.db argument builds every language, as the pod it came
from would.
--db-round-trip takes --lang for the same reason, and ocda-db looses
--parallel (would cause a race).
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The file is named for the document, not for one of its languages (and
is written a language at a time: ten calls fill one file for a ten
language document).
The removal on entry was right when a call wrote the whole file. It
now happens on the first language of the document, taken from the
manifest rather than from whatever order the loop ran in, so nine
languages are no longer written and thrown away.
Also a partial unique index for the file level metadata rows. In SQL
no null equals any other null, so UNIQUE(doc_id, key) does not
constrain the rows written with a null doc_id, and schema.name and
schema.version were inserted once per language: ten rows each, with
nothing for INSERT OR REPLACE to replace.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
| |
doc_uid_out ends in ".{lng}", which is right for an artefact written per
language and wrong for one that holds every language of a document.
Rather than strip the suffix at each place that needs the bare name,
build it by the same branches without the language, so the two names
cannot drift apart.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
each file still holds one language, here groundwork for one database per
document. The shape changes here and nothing merges yet: a database is
still written per language, so every doc_id is 1. Real and testable on
its own, where the writer and the reader together are not.
documents, a row per language, is what lets one file hold a document's
whole set. Everything that tells one language's rows from another's keys
on documents.id.
metadata is keyed on (doc_id, key) (no longer on key alone). Every
language has a title and a creator, and a key-only primary key refuses
the second one. schema.name and schema.version describe the file rather
than a document in it, so they are written with a null doc_id and are
the only rows that are.
objects gain doc_id and its uniqueness widens from (section, seq) to
(doc_id, section, seq). ('body', 0) exists once per language, so the
narrow constraint was the thing that would have refused a second
language outright. idx_objects_section leads with doc_id, or reading one
language scans them all.
objects.id stays a global INTEGER PRIMARY KEY, so object_images,
object_links, object_anchors and object_subtoc keep their schema and
their keys, and objects_fts keeps content_rowid='id'. The reader's four
sweeps filter through objects rather than gaining a column of their own.
outline and citable name the language and order by it first. A view over
a file that may hold several languages and does not say which reads as
one document and is several.
The DDL is IF NOT EXISTS throughout, since a second language will open a
file that already has its schema.
dbReadFile takes an optional language and means "the only document in
it" without one. Given none where there are several it reports the
languages and stops, rather than returning the first: the round trip
compares byte for byte, and a quietly wrong answer there would read as a
spine fault.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
format 2.0, source.digest covers every markup file, the sha256 over one
line per markup file, "<sha256> <filename>", sorted by filename.
source.digest was the sha256 of the master file as read, before any
insert. This is right for a .sst, which is the whole document. For a
.ssm with inserts .ssi (containing the substantive part of a documents
text) this is close to meaningless.
Considered change of meaning (rather than an addition) and given the
major part of the format version with it: 1.1 becomes 2.0, in the .ssp
header line and in the database's schema.version row together.
A 1.x reader refuses a 2.0 artefact, which is what that check is for.
- Sorted, so the value does not depend on the order a filesystem hands
back a directory.
- By filename (rather than by path), so it does not depend on where the
pod sits, which is the property the pod rebuild comparison rests on:
a pod unzipped elsewhere must still give the same abstraction.
Within one language every markup file lives in one directory, so a
filename identifies it.
Each line is checkable on its own against a digests.txt line or a files
row, which the single value was not. The whole is reproducible with
sha256sum and sort, and was verified that way for a one file document
and for a twenty file one.
An unreadable file is named in the digest rather than skipped. The parse
has failed elsewhere by then, and a digest that quietly left a file out
would claim the document is something it is not.
The reference abstractions are regenerated. Two lines change in each and
no others: the version, and the digest.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A document's po4a configuration is derived rather than kept. Spine
already knows both halves without looking: the manifest says which
languages, and the markup says which files, through the << lines the
parser follows. Derived, it cannot disagree with the document, where a
hand-kept file eventually does: a language added to the manifest and
forgotten in the Makefile is a translation nobody updates.
It prints, rather than writing. Spine does not write into a source tree,
and the file belongs beside the catalogues in the pod it describes:
spine --po4a-cfg <pod> > <pod>/tools/po4a/po4a.cfg
Three things the Makefile it replaces knew, which had to be found by
running po4a rather than by reading:
--keep 0 is the difference between regenerating a document and
destroying it. po4a will not write a translated file below a
completeness threshold, and that threshold defaults to 80%. A partly
translated document is the normal state of one being worked on, and sisu
covers it: an untranslated string falls back to the source text.
Measured at the default, 181 of live-manual's 189 translated files would
be discarded, and nine languages would collapse to eight surviving
files, reading as the document reverting to its source language rather
than as a failure.
index.html.in is not markup and no << line names it, so a configuration
built from the manifest and the insert list alone drops it, and with it
index.html.in.pot and its nine .po files at the next regeneration. It is
found by asking whether the source language has one.
neverwrap is not set, and the measurement is recorded beside the code:
setting it turns 889 of 1126 strings fuzzy in four languages, and po4a
does not use a fuzzy translation, so Catalan's coverage would fall from
about 46% to about 12%. Tidy formatting is not worth that.
The run-complete banner is suppressed for this action. The configuration
goes to stdout and the banner would otherwise be part of the file, which
is what the round trips already avoid by leaving early; this one prints
from inside output processing and cannot.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A pod now carries its catalogues, so a translation can be updated
rather than only retyped: providing the .pot and .po (the means of
maintaining it).
A directory walk, where every other thing the pod writer carries is
named. A catalogue set is whatever the translator has, and enumerating
.pot and .po files would mean the writer deriving the document's
languages a second time, from a second place, to say what it already
reads from the manifest.
Done once, on the last language, as the markup blocks are. The tree is
not language specific, so a pass for each language would re-read all of
it, 5.5 MB ten times over for live-manual, into an archive the next pass
overwrites.
Every name is checked before it is carried, by the rule the pod reader
applies to an archive it is handed. The writer is reading a directory
someone else may have filled, which is the reason the reader checks.
The cost is small, and it is the whole argument for compressing the
archive as a stream rather than per member:
uncompressed content 6972418 -> 13167489 +89%
the .sisupod 692942 -> 891388 +28.6%
5.5 MB of catalogues cost 198 KB, because the .po files largely repeat
markup already in the archive and whole-stream compression sees it.
Per-member deflate could not.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Both now take the number from one constant now set at 10,000, the writer
checks it before the archive is written.
Previously the number set for the pod reader was less than 500 members.
The writer had no limit, (so spine could write a pod it would then
refuse to read, reporting too many entries: a good file that looks
corrupt, and only for a document with enough parts).
500 was set when a pod held markup and images. A pod carrying
translation catalogues has a different arithmetic, and the count follows
from how many files a document is made of times how many languages it
has, not from how large the document is:
the_wealth_of_networks 213,405 words, 1 file per language
12 languages -> 49 entries
live-manual 24,724 words, 20 files per language
12 languages -> 506 entries
The big book is not what runs into this; the modular manual is.
live-manual stood two languages from an unreadable artefact.
10,000 covers a hundred-insert manual in thirty languages, about 6,200
entries, with room. It is deliberately well clear of any real document
rather than snug above the largest one known: the two errors are not
comparable, since too low refuses a legitimate document with a message
that reads as corruption, while too high defers to a size cap a moment
later. The count is the weakest of the three guards and is not what
bounds resource use; the per entry and total size caps do that, both
before a byte is written.
Refusing also removes any archive an earlier language left. The writer
runs once per language and only the last pass holds every language, so
it is the last that goes over, and the passes before it wrote smaller
archives that passed. Without this, refusal left a pod missing a
language: an artefact that reads perfectly well and is wrong. A missing
file is an error someone notices.
The arithmetic is recorded beside the constant so the next person can
re-derive the number rather than guess at it.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The pod archive spine writes is now named .sisupod and carries the
same archive inside one zstd frame. The reader has accepted both
since the previous commit, so every pod already published stays
readable and this changes only what is written.
free_culture 2863504 -> 1100180
live-manual 6972418 -> 692942 10.06x
The wrap is at the pod call site and not in createZipFile, which
also writes every epub and every odt.
Best compression from whole-stream (rather than per-member compression).
(for live-manual per-member deflate manages about 3x compared to 10x for
whole pod content compression).
The suffix is named as pod_archive_suffix.
--pod-compression sets the level, 19 by default: a pod is written once
and fetched many times (692942 bytes at 19, 837703 at 9, 2054365 at 1).
(A value that is not a number warns and the default is used).
Output built from a written .sisupod is byte identical to output built
from the markup it came from.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
sisupod using zstd compression recognised by first bytes.
<doc>.sisupod: spine reads one before it writes one, so when the name
changes (from .zip to .ssiupod) every pod already published stays
readable.
A .sisupod is one zstd frame wrapping the archive spine already builds.
The reader unwraps the bytes and hands the same archive to the same
parser, so every guard downstream is untouched: entry names, per-entry
and total size, path depth, escape and symlinks.
Which container it is comes from the first four bytes (rather than the
name). A pod published as a plain .zip reads as it always did, a
.sisupod reads, and either one renamed reads too. The suffix is still
recognised, both spellings, for the argument and for a url.
libzstd is declared rather than bound: provides the whole surface of
fifteen extern C prototypes (there is nothing to generate and no
upstream tree to track, which is the arrangement sqlite3 already has).
dub links it with "libs": [ "zstd" ]; nix needs zstd.out rather than
zstd, whose default output is the binaries and carries no library at
all.
A frame declares its uncompressed size in its own header, and for a pod
fetched over https that number is attacker controlled. The declared size
is checked against a ceiling before a buffer is asked for, a frame that
will not declare one is refused, and what comes out is checked against
what was promised. The ceiling is the extraction limit the archive
reader already applies, so the two bounds agree.
Measured: output built from a .sisupod is byte identical to output built
from the same pod's .zip, 359 files over text, html, epub, odt and .ssp
for three documents, live-manual's ten languages included. A truncated
frame is refused and the document skipped.
free_culture 2863504 -> 1099736 2.60x
the_wealth_of_networks 4294022 -> 1172649 3.66x
live-manual 6972418 -> 675970 10.31x
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Three faults in four lines of the odt metadata.
meta:creation-date and dc:date were taken from generated_time, the run's
own clock, so two runs over unchanged markup produced different odt
bytes. They now carry the document's own instant, so an odt is a
function of its source. That was the last clock stamp in any writer:
every output format is now reproducible.
generated_time is a human-readable run stamp, of the form
"2026-9-17 [38/3] 18:39:35", with unpadded numbers and an iso week in
brackets. It is not a valid xsd:dateTime, which is what ODF asks for in
both of those fields, so what was written there was malformed as well as
unstable.
dc:language now carries the document's own language tag.
The instant is doc_matters.modified_utc, moved there from epub3.d
in this commit: epub3 needs the same value for its dcterms:modified,
and one definition cannot drift from the other. generated_time now
has no caller and is left in place.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
| |
Every metadata page body element carried lang="en" xml:lang="en" as a
literal. It now carries the document's own language tag.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
timestamp dcterms:modified from the document, not the clock
EPUB3 requires exactly one dcterms:modified on the package
It is now read in order: SOURCE_DATE_EPOCH where the environment sets
it, which is the reproducible-builds convention and lets a build pin
every artefact of one run; then the document's own date.modified; then
date.published. The clock remains the last resort, for a document that
says nothing about when it is from.
A document date is loose markup. "1991", "2006-03" and "2021-12-00" all
occur in the sample set, and a month or day of 00 is not a date. A
missing or zero part becomes 01, and a day out of range for its month
falls back to the first, so what is written is always a valid
CCYY-MM-DDThh:mm:ssZ.
Over the sample set two consecutive runs now produce all 36 epubs byte
identical, and epubcheck reports no change: 0 fatal, 0 errors.
(previously taken from the clock at build time. Two runs over unchanged
markup therefore produced different epub bytes, so no comparison of epub
output could tell a real change from the seconds between two builds. It
also meant an epub was a function of when it was built rather than of
what it was built from.)
(assisted by Claude-Code)
|
| |
|
|
|
|
|
| |
Spine now looks for conf/document_make, and bundles that name.
(the dr_ prefix was a sisu-era name to distinguish it)
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Spine names a language as the directory under media/text/ is
named, which for a region is the posix form with an underscore,
pt_BR. A lang, xml:lang or dc:language value is a BCP 47 (RFC
3066) tag and wants a hyphen, pt-BR. Spine put its own code
there unchanged, so a document in a region language produced an
invalid attribute in every one of its xhtml files: 138 epubcheck
errors for a document of ten segments and a navigation file.
src.language_tag is the same language as a tag. The attributes
and the dc metadata now read it. Output paths keep src.language,
since a path is named after the directory, not after the tag.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
| |
Fixes for three faults in the insert (.ssi) digests.
The key was the language captured from whichever path resolved the
insert. It did not change across the loop over languages, so
every language's inserts were filed under one language. The key is now
the language being digested, as the digest of the primary file above it
already did.
- digest is taken after the check that the file exists.
- digest is now taken inside the guard.
- non-pod path now correctly keyed on the document's own language.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
| |
The flag has existed since a url argument could be fetched and nothing
ever read it, so every url was fetched whether or not the flag was
given. A refused url is now dropped from the arguments, which is what a
failed download already did.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
bugfix: content image names are now checked in a pass of their own,
before anything is created or written, and one bad name refuses every
image in the artefact rather than skipping the one, which is what the
zip reader does with a zip. The write then re-checks containment.
(prior to this an image name read out of a .ocda.db was written without
checks, and such a database can be fetched over https. A name could
climb with "..", and being handed to chainPath an absolute name dropped
everything before it, so "/x" did not land under the image directory at
all. The zip reader has guarded against this since it was written; the
database reader did not).
carried_names.d holds both rules: a bare filename, as an image is
carried, and a relative path, for the markup and conf a database is to
carry.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
A file beginning with a utf-8 byte order mark had the mark read as part
of its first yaml key, so "title:" arrived as "title:" and every
value under that key was dropped silently: the document kept its text
and lost its title. Downstream that showed as an empty dc:title and an
empty <title> element in the epub, which epubcheck reports as two
errors.
The mark is now kept out of the header and body split, and so out of
the parse. It is not stripped when the file is read, because that text
is what source.digest is taken over and the digest names the file as it
sits on disk.
One document in the sample set begins with a mark. It is left as it is:
it is the case this guards against.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Four things, not visible in an error count:
- Images had no alt text, even where markup carried the text, and the
field alt text is present, incorrectly used in code; fixed as the alt.
{ sm_tux.png 64x80 "Gnu/Linux - a better way" }image
an image with nothing to say for itself gets alt=""
- Images with no dimensions were incorrectly written width="0"
height="0", fixed
- Accessibility metadata, which EPUB Accessibility 1.1 requires:
schema:accessMode, accessibilityFeature and accessibilityHazard, plus
accessModeSufficient and accessibilitySummary, which epubcheck does
not test for; used by
- readers looking for a book it can use,
- anyone distributing into the EU.
Every value is derived from the document, not asserted: a document
with no images says so, one whose images all carry alt text claims
alternativeText and textual sufficiency, one where an image says
nothing claims neither.
Nothing claims WCAG conformance, which would be a claim about an
evaluation that has not happened.
- DPUB-ARIA roles beside the structure spine already knew about.
doc-toc, doc-endnotes, doc-bibliography, doc-glossary and doc-index on
the sections, doc-noteref on a note reference and doc-footnote on the
note it points at. "doc_endnotes" as a class name means something to a
stylesheet and nothing to a reading system. Not doc-biblioentry: DPUB-
ARIA 1.1 deprecated it. These are plain ARIA, so the html gets them
too.
Also dealt with here:
- Spine makes certain sections such as endnotes, bookindex, ... and a
heading that has a name for itself cannot share the same name, if it
does the name is now disambiguated, disambiguated. (disambiguation
similar to that used for a repeated anchor tag)
- Endnote anchors fixed to use "id" (being one unique tag in the
document that is target of every reference to it)instead of obsolete
form (they were <a name="note_1">. "name" on an <a>).
Book index markers keep "name": the same marker sits on every object
an entry names, 296 times in one document, and those are markers of
where a term occurs rather than distinct targets.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
epubcheck 5.3.0 over the 35 sample epubs went from 1302 remaining errors
to none. The six causes for those remaining errors fixed here.
most of the count was one of them:
- Tables (1123 errors, 86%) where the writer emitted obsolete
attributes, removed in 2014; table also now in div instead of a
paragraph. also fixed:
- Navigation (115). Every level 4 heading's nav entry ended "#0", the
ocn of an object that has no ocn and so no id.
- Duplicate ids (37). (from three unrelated sources).
- The publication identifier (27 warnings). dc:identifier was a
hardcoded hex string, not a UUID, and the same one in all 35 files,
so two documents in one library collided. It is now UUID v5 over the
document's own uid: a real RFC 4122 identifier, stable across builds
and distinct per document.
- Image manifest entries (15). media-type was "image/" plus the file
extension, giving "image/jpg", which is not a media type; epubcheck
reads it as a foreign resource and wants a fallback. There is now a
mapping, which also covers svg and webp. Item ids came from the file
basename, and "2bits_02_01-100.png" gave an id starting with a digit,
which is not an XML name; ids are now prefixed and sanitised.
- An ocn resolving to no segment (9). A poem block and its first verse
share one ocn; the verses are written into a segment and the block
itself is not.
Also dropped from the package document: an xmlns:xsi that nothing
used, and a prefix declaration for the rendition vocabulary, which is
a reserved prefix and needed no declaring. The package now carries
xml:lang.
test-epub-validity.sh now allows no fatal and no error. Two NAV-011
warnings remain and are described in it. Three reference .ssp files
change, all of them the anchor tag disambiguation above.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
| |
in error regex permitted a url to go past an endnote close delimiter,
fixed
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
| |
fixes related to:
- named anchor id
- subtitles, emit only where they exist
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
| |
- wrapper set to <div>, with class unchanged, and the six stylesheets
select div.code where they selected p.code.
- invented element <codeline> removed
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
| |
- _part_eof.xhtml was declared in the manifest and never written.
- image_sys/bullet_09.png was referenced as a background image, a
throwback to the Ruby version of sisu, never used here.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
| |
One document head per file, and ids must be unique. No remaining fatal
errors in the epub in the marup sample set.
(assisted by Claude-Code)
|