<feed xmlns='http://www.w3.org/2005/Atom'>
<title>sisudoc-spine/org/document_source_conversions.org, branch main</title>
<subtitle>SiSU Spine: document publishing and search (in D) 2015</subtitle>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/'/>
<entry>
<title>ocda db: carry catalogues &amp; text blobs compressed</title>
<updated>2026-09-22T20:17:59+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-21T17:33:44+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=4943c93919f349e2d47e6a74f498779fb13f7862'/>
<id>4943c93919f349e2d47e6a74f498779fb13f7862</id>
<content type='text'>
The pod carries tools/po4a and the database (until now) did not, so a
pod written back out of a database came back a lossy copy, without its
translation catalogues.

the database now carries them, walked and name-checked exactly as the
pod writer walks and checks them, once on the last language. They are
compressed to save space, the test sample carrying 6.25 MB of catalogue
as read would have made that one document's database two thirds larger.
Files gains a compression column: NULL or 'none' for the bytes as read,
'zstd' for a frame. What to compress is decided by role, not by size
(source, conf, manifest and tools are text and are compressed; images
being compressed already are not). The same kind of file is always
stored the same way. Document objects are untouched and stay raw:
objects_fts is external content over objects.text and reads that column
directly.

bytes and sha256 remain those of the original file, so every digest
check, digests.txt line and source.digest rebuilt from these rows
works unchanged, and a reader that does not decompress can still say
what it is looking at.

dbReadFiles decompresses, so no caller learns how a blob is stored. It
asks pragma_table_info whether the column is there at all: a database
written before it reads as raw. A row that claims zstd and will not
decompress is named and left out, rather than handed back as a frame
where markup should be.

live-manual: 9,490,432 bytes without the catalogues, 9,908,224 with
them. Output from the materialised pod is identical to output from the
original pod, 679 files each side, no differences.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
The pod carries tools/po4a and the database (until now) did not, so a
pod written back out of a database came back a lossy copy, without its
translation catalogues.

the database now carries them, walked and name-checked exactly as the
pod writer walks and checks them, once on the last language. They are
compressed to save space, the test sample carrying 6.25 MB of catalogue
as read would have made that one document's database two thirds larger.
Files gains a compression column: NULL or 'none' for the bytes as read,
'zstd' for a frame. What to compress is decided by role, not by size
(source, conf, manifest and tools are text and are compressed; images
being compressed already are not). The same kind of file is always
stored the same way. Document objects are untouched and stay raw:
objects_fts is external content over objects.text and reads that column
directly.

bytes and sha256 remain those of the original file, so every digest
check, digests.txt line and source.digest rebuilt from these rows
works unchanged, and a reader that does not decompress can still say
what it is looking at.

dbReadFiles decompresses, so no caller learns how a blob is stored. It
asks pragma_table_info whether the column is there at all: a database
written before it reads as raw. A row that claims zstd and will not
decompress is named and left out, rather than handed back as a frame
where markup should be.

live-manual: 9,490,432 bytes without the catalogues, 9,908,224 with
them. Output from the materialised pod is identical to output from the
original pod, 679 files each side, no differences.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>ocda db: materialise a pod from carried source</title>
<updated>2026-09-22T19:30:59+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-21T17:16:28+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=36c4b1721ce99865d10b9dd74c57290869efb383'/>
<id>36c4b1721ce99865d10b9dd74c57290869efb383</id>
<content type='text'>
A database that carries its source can recreate the pod. --source and
--pod2 given a .ocda.db now write the pod to a directory of the run's
making and carry on with the pod, so everything after that point is
handling a pod like any other.

Nothing renders from the database. A materialiser that also rendered
would be a second path to every output format and the two would drift;
one that only writes files means the document is built by the same
code over the same bytes as the original, and identical output is a
consequence rather than an aspiration. Held against the original pod,
site configuration constant: every output file identical across ten
languages.

The pod's name comes from the database's filename, which inverts the
naming rule exactly: &lt;doc&gt;.ocda.db is named by doc_uid_out_no_lang,
the pod name and the document's filename joined by ":" when they
differ and the one name when they do not. So the materialised pod
recomputes the uid it was named by and every output file lands on the
name it had.

Names are checked before anything is created, and one bad name refuses
the artefact rather than skipping a file, as the zip reader does with
a zip. Markup that does not match the digest stored with it is refused
outright, where a mismatched image is written with a warning: a wrong
image makes a document that looks wrong, a wrong markup file makes one
that is wrong, in its text, with nothing downstream to notice.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
A database that carries its source can recreate the pod. --source and
--pod2 given a .ocda.db now write the pod to a directory of the run's
making and carry on with the pod, so everything after that point is
handling a pod like any other.

Nothing renders from the database. A materialiser that also rendered
would be a second path to every output format and the two would drift;
one that only writes files means the document is built by the same
code over the same bytes as the original, and identical output is a
consequence rather than an aspiration. Held against the original pod,
site configuration constant: every output file identical across ten
languages.

The pod's name comes from the database's filename, which inverts the
naming rule exactly: &lt;doc&gt;.ocda.db is named by doc_uid_out_no_lang,
the pod name and the document's filename joined by ":" when they
differ and the one name when they do not. So the materialised pod
recomputes the uid it was named by and every output file lands on the
name it had.

Names are checked before anything is created, and one bad name refuses
the artefact rather than skipping a file, as the zip reader does with
a zip. Markup that does not match the digest stored with it is refused
outright, where a mismatched image is written with a warning: a wrong
image makes a document that looks wrong, a wrong markup file makes one
that is wrong, in its text, with nothing downstream to notice.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
</feed>
