| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
- A poem is a container of verse, and stores its range of verse ocn,
- its verse are the citable units with ocn.
- A note in the last verse of a poem previously was not gathered into
the endnotes section, this now is fixed
Format 2.0 -> 2.1: the property is an addition, the poem is not a
citable object (but contans a range of objects), its verse are
(individual citable objects), and every reader checks the major version.
A 2.0 database reads as having no ranges.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The pod carries tools/po4a and the database (until now) did not, so a
pod written back out of a database came back a lossy copy, without its
translation catalogues.
the database now carries them, walked and name-checked exactly as the
pod writer walks and checks them, once on the last language. They are
compressed to save space, the test sample carrying 6.25 MB of catalogue
as read would have made that one document's database two thirds larger.
Files gains a compression column: NULL or 'none' for the bytes as read,
'zstd' for a frame. What to compress is decided by role, not by size
(source, conf, manifest and tools are text and are compressed; images
being compressed already are not). The same kind of file is always
stored the same way. Document objects are untouched and stay raw:
objects_fts is external content over objects.text and reads that column
directly.
bytes and sha256 remain those of the original file, so every digest
check, digests.txt line and source.digest rebuilt from these rows
works unchanged, and a reader that does not decompress can still say
what it is looking at.
dbReadFiles decompresses, so no caller learns how a blob is stored. It
asks pragma_table_info whether the column is there at all: a database
written before it reads as raw. A row that claims zstd and will not
decompress is named and left out, rather than handed back as a frame
where markup should be.
live-manual: 9,490,432 bytes without the catalogues, 9,908,224 with
them. Output from the materialised pod is identical to output from the
original pod, 679 files each side, no differences.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The database already carried the images, because without them no output
can be produced. It now carries the markup as well, and with it the
document's configuration and the manifest as the author wrote it.
That is the difference between a serialised abstraction and a document
source. An abstraction can be rendered but not re-parsed, and a reader
who wants to correct a sentence needs the sentence as written. With
these a pod can be written back out of the sqlite-file.
Three new values in the existing role column, no schema change: the
format stays 2.0.
Source rows are named by their path within the pod,
media/text/<lang>/<file>, and not by bare filename as images are. Every
language of a document has a file of the same name, so bare names would
collide under UNIQUE(role, name) and nine of ten would be dropped
without a word. The path is also what a pod materialised from this
database has to be told.
Bytes stored as read with the digest over them, so that each row is
checkable against the line source.digest was built from.
Also build.spine_version, a file level row saying which spine wrote the
file: not a property of the document, and the one thing a file cannot be
asked for afterwards. The reader skips the build. prefix as it skips
schema. and translation., or it would come back inside the document
header and the round trip would differ.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The check runs once every language of a document has been abstracted.
Providing a warning by default. --strict makes divergence a failure,
taken at the end of the run so that outputs are complete and can be
examined. (the name --strict is general, so that later checks can be
added without a second flag).
Under --parallel the profiles are appended under synchronized, and the
comparison sorts a document's languages by name rather than taking
them in the order the threads finished, so two runs print the same
lines in the same order.
The outcome is noted per language in the database as
translation.ocn_aligned, for documents that have more than one
language.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The file is named for the document, not for one of its languages (and
is written a language at a time: ten calls fill one file for a ten
language document).
The removal on entry was right when a call wrote the whole file. It
now happens on the first language of the document, taken from the
manifest rather than from whatever order the loop ran in, so nine
languages are no longer written and thrown away.
Also a partial unique index for the file level metadata rows. In SQL
no null equals any other null, so UNIQUE(doc_id, key) does not
constrain the rows written with a null doc_id, and schema.name and
schema.version were inserted once per language: ten rows each, with
nothing for INSERT OR REPLACE to replace.
(assisted by Claude-Code)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
each file still holds one language, here groundwork for one database per
document. The shape changes here and nothing merges yet: a database is
still written per language, so every doc_id is 1. Real and testable on
its own, where the writer and the reader together are not.
documents, a row per language, is what lets one file hold a document's
whole set. Everything that tells one language's rows from another's keys
on documents.id.
metadata is keyed on (doc_id, key) (no longer on key alone). Every
language has a title and a creator, and a key-only primary key refuses
the second one. schema.name and schema.version describe the file rather
than a document in it, so they are written with a null doc_id and are
the only rows that are.
objects gain doc_id and its uniqueness widens from (section, seq) to
(doc_id, section, seq). ('body', 0) exists once per language, so the
narrow constraint was the thing that would have refused a second
language outright. idx_objects_section leads with doc_id, or reading one
language scans them all.
objects.id stays a global INTEGER PRIMARY KEY, so object_images,
object_links, object_anchors and object_subtoc keep their schema and
their keys, and objects_fts keeps content_rowid='id'. The reader's four
sweeps filter through objects rather than gaining a column of their own.
outline and citable name the language and order by it first. A view over
a file that may hold several languages and does not say which reads as
one document and is several.
The DDL is IF NOT EXISTS throughout, since a second language will open a
file that already has its schema.
dbReadFile takes an optional language and means "the only document in
it" without one. Given none where there are several it reports the
languages and stops, rather than returning the first: the round trip
compares byte for byte, and a quietly wrong answer there would read as a
spine fault.
(assisted by Claude-Code)
|
| | |
|
| |
|