<feed xmlns='http://www.w3.org/2005/Atom'>
<title>sisudoc-spine/src/sisudoc/ocda/abstraction/ssp_in.d, branch main</title>
<subtitle>SiSU Spine: document publishing and search (in D) 2015</subtitle>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/'/>
<entry>
<title>abstraction 2.1: a poem states its object range</title>
<updated>2026-10-01T03:14:23+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-30T20:18:59+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=993a3f655359f98dc26d3d3222e07792e2ccddf7'/>
<id>993a3f655359f98dc26d3d3222e07792e2ccddf7</id>
<content type='text'>
- A poem is a container of verse, and stores its range of verse ocn,
  - its verse are the citable units with ocn.

- A note in the last verse of a poem previously was not gathered into
  the endnotes section, this now is fixed

Format 2.0 -&gt; 2.1: the property is an addition, the poem is not a
citable object (but contans a range of objects), its verse are
(individual citable objects), and every reader checks the major version.
A 2.0 database reads as having no ranges.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
- A poem is a container of verse, and stores its range of verse ocn,
  - its verse are the citable units with ocn.

- A note in the last verse of a poem previously was not gathered into
  the endnotes section, this now is fixed

Format 2.0 -&gt; 2.1: the property is an addition, the poem is not a
citable object (but contans a range of objects), its verse are
(individual citable objects), and every reader checks the major version.
A 2.0 database reads as having no ranges.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>abstraction: format 2.0, source.digests</title>
<updated>2026-09-22T18:23:53+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-21T14:52:39+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=e1c9056903c2478066d5da8934d704f4d1935da4'/>
<id>e1c9056903c2478066d5da8934d704f4d1935da4</id>
<content type='text'>
format 2.0, source.digest covers every markup file, the sha256 over one
line per markup file, "&lt;sha256&gt; &lt;filename&gt;", sorted by filename.

source.digest was the sha256 of the master file as read, before any
insert. This is right for a .sst, which is the whole document. For a
.ssm with inserts .ssi (containing the substantive part of a documents
text) this is close to meaningless.

Considered change of meaning (rather than an addition) and given the
major part of the format version with it: 1.1 becomes 2.0, in the .ssp
header line and in the database's schema.version row together.
A 1.x reader refuses a 2.0 artefact, which is what that check is for.

- Sorted, so the value does not depend on the order a filesystem hands
  back a directory.
- By filename (rather than by path), so it does not depend on where the
  pod sits, which is the property the pod rebuild comparison rests on:
  a pod unzipped elsewhere must still give the same abstraction.
  Within one language every markup file lives in one directory, so a
  filename identifies it.

Each line is checkable on its own against a digests.txt line or a files
row, which the single value was not. The whole is reproducible with
sha256sum and sort, and was verified that way for a one file document
and for a twenty file one.

An unreadable file is named in the digest rather than skipped. The parse
has failed elsewhere by then, and a digest that quietly left a file out
would claim the document is something it is not.

The reference abstractions are regenerated. Two lines change in each and
no others: the version, and the digest.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
format 2.0, source.digest covers every markup file, the sha256 over one
line per markup file, "&lt;sha256&gt; &lt;filename&gt;", sorted by filename.

source.digest was the sha256 of the master file as read, before any
insert. This is right for a .sst, which is the whole document. For a
.ssm with inserts .ssi (containing the substantive part of a documents
text) this is close to meaningless.

Considered change of meaning (rather than an addition) and given the
major part of the format version with it: 1.1 becomes 2.0, in the .ssp
header line and in the database's schema.version row together.
A 1.x reader refuses a 2.0 artefact, which is what that check is for.

- Sorted, so the value does not depend on the order a filesystem hands
  back a directory.
- By filename (rather than by path), so it does not depend on where the
  pod sits, which is the property the pod rebuild comparison rests on:
  a pod unzipped elsewhere must still give the same abstraction.
  Within one language every markup file lives in one directory, so a
  filename identifies it.

Each line is checkable on its own against a digests.txt line or a files
row, which the single value was not. The whole is reproducible with
sha256sum and sort, and was verified that way for a one file document
and for a twenty file one.

An unreadable file is named in the digest rather than skipped. The parse
has failed elsewhere by then, and a digest that quietly left a file out
would claim the document is something it is not.

The reference abstractions are regenerated. Two lines change in each and
no others: the version, and the digest.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>spine: process a document from .ocda.db or .ssp</title>
<updated>2026-09-12T16:14:39+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-11T17:57:47+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=f41bf1cc525c70cced764a5e4df70d499d82f9c9'/>
<id>f41bf1cc525c70cced764a5e4df70d499d82f9c9</id>
<content type='text'>
Process a document from its .ocda.db or .ssp abstraction. A .ssp or a
.ocda.db given as an argument is now a document like any other: read
rather than parsed, and handed to the output writers as the same doc
they already take.

The 35 document sample collection built from its .ssp files is byte
identical to the same collection built from markup: 1457 files, no
differences.

Three things had to be rebuilt rather than read, all of them a value
parsed rather than a structure derived:

- classify_topic_register_arr and its expanded twin. The split is not a
  plain one, so the rule moved to sisudoc.ocda.meta.topic_register and
  both the yaml reader and this one call it.
- creator_author_arr, which is creator.author split on the ", " it was
  joined with.
- title_sub, a copy of title_subtitle made where the header is read, and
  what epub3 puts in dc:title id="subtitle".

Fixed ordering bug found by the acceptance test. A heading's own anchor
can be a bare number taken from its text and this can collide with the
ocn of an unrelated object. Whole output comparison found that before
the fix there was one wrong link in one epub's table of contents.

From a .ocda.db, 1446 of 1457 files are identical. The 11 that are not
are the images: a database carries its own image blobs and nothing yet
extracts them, so the five sisu_markup images are not copied and the
epub that embeds them differs. That is the next step and is not a defect
in this one.

--source and --pod2 are refused with a warning rather than half
done, no artefact carrying the markup.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
Process a document from its .ocda.db or .ssp abstraction. A .ssp or a
.ocda.db given as an argument is now a document like any other: read
rather than parsed, and handed to the output writers as the same doc
they already take.

The 35 document sample collection built from its .ssp files is byte
identical to the same collection built from markup: 1457 files, no
differences.

Three things had to be rebuilt rather than read, all of them a value
parsed rather than a structure derived:

- classify_topic_register_arr and its expanded twin. The split is not a
  plain one, so the rule moved to sisudoc.ocda.meta.topic_register and
  both the yaml reader and this one call it.
- creator_author_arr, which is creator.author split on the ", " it was
  joined with.
- title_sub, a copy of title_subtitle made where the header is read, and
  what epub3 puts in dc:title id="subtitle".

Fixed ordering bug found by the acceptance test. A heading's own anchor
can be a bare number taken from its text and this can collide with the
ocn of an unrelated object. Whole output comparison found that before
the fix there was one wrong link in one epub's table of contents.

From a .ocda.db, 1446 of 1457 files are identical. The 11 that are not
are the images: a database carries its own image blobs and nothing yet
extracts them, so the five sisu_markup images are not copied and the
epub that embeds them differs. That is the next step and is not a defect
in this one.

--source and --pod2 are refused with a warning rather than half
done, no artefact carrying the markup.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>ocda: heading cross reference linking</title>
<updated>2026-09-12T16:14:39+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-11T16:32:13+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=4e6774904ce811b074d8b3694b04dc5c1e8219e4'/>
<id>4e6774904ce811b074d8b3694b04dc5c1e8219e4</id>
<content type='text'>
The segment a cross reference to a heading lands on, gixed upstream in
ocda, where the field is created, rather than inferred by the loader.

A heading above level 4 opens no html segment of its own, so a link to
it has to land on the level 4 heading that follows. The build loop
cannot know that when it reads the heading, and says so in a comment of
its own: "for html segname need following lv4 not yet known". It
back-fills tag_assoc when the level 4 heading arrives (lv0to3_tags), and
the answer then lives only in that map.

tags.segment_lv4_is now holds it, resolved in a pass over the finished
head and body sections: walking backwards, the last level 4 heading seen
is the next one for everything above it, which is the same answer the
back-fill gives. Emitted in the .ssp only where it differs from
.segment_html_is, which is every heading at level 4 or below, so it is
sparse: 317 lines over the 35 document reference set, none removed.
Carried in the database as a column of its own, and read back by both
readers.

docHasFromAbstraction reads it instead of working it out. With that, and
on top of the two defect fixes, a document loaded from an artefact and
one parsed from markup agree:

  key sets       identical, all 35 documents
  values         3 entries differ of some 30,000, all of them
                 _the_title, where the abstraction supplies an epub
                 segment the parser leaves unset
  link targets   all 1,651 agree, against 89 differing before any
                 of this work and 1 after the defect fixes alone

Format stays v1.1. That version is new in this same run of work and
nothing outside spine has read it, so this belongs in it rather than in
a bump of its own.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
The segment a cross reference to a heading lands on, gixed upstream in
ocda, where the field is created, rather than inferred by the loader.

A heading above level 4 opens no html segment of its own, so a link to
it has to land on the level 4 heading that follows. The build loop
cannot know that when it reads the heading, and says so in a comment of
its own: "for html segname need following lv4 not yet known". It
back-fills tag_assoc when the level 4 heading arrives (lv0to3_tags), and
the answer then lives only in that map.

tags.segment_lv4_is now holds it, resolved in a pass over the finished
head and body sections: walking backwards, the last level 4 heading seen
is the next one for everything above it, which is the same answer the
back-fill gives. Emitted in the .ssp only where it differs from
.segment_html_is, which is every heading at level 4 or below, so it is
sparse: 317 lines over the 35 document reference set, none removed.
Carried in the database as a column of its own, and read back by both
readers.

docHasFromAbstraction reads it instead of working it out. With that, and
on top of the two defect fixes, a document loaded from an artefact and
one parsed from markup agree:

  key sets       identical, all 35 documents
  values         3 entries differ of some 30,000, all of them
                 _the_title, where the abstraction supplies an epub
                 segment the parser leaves unset
  link targets   all 1,651 agree, against 89 differing before any
                 of this work and 1 after the defect fixes alone

Format stays v1.1. That version is new in this same run of work and
nothing outside spine has read it, so this belongs in it rather than in
a bump of its own.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>ocda: reading .ocda.db now faster than parsing markup</title>
<updated>2026-09-12T16:14:39+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-11T14:55:34+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=0ef5f8b0db9e79b6fa059cbfe8685d2063100ea5'/>
<id>0ef5f8b0db9e79b6fa059cbfe8685d2063100ea5</id>
<content type='text'>
reading a .ocda.db was slower than parsing markup source. Taking the
largest markup document sample War and peace, 12,135 objects, optimised
build, before this commit and after:

  parse markup    0.87 s   0.87 s
  load .ssp       0.047 s  0.047 s
  load .ocda.db   1.09 s   0.128 s

The database goes from being 1.25x slower than parsing the document to
6.8x faster.

Two things fixed in the reader were:

First, four queries per object. object_images, object_links,
object_anchors and object_subtoc were queried per object as the objects
were built, each statement compiled fresh from a concatenated string.
For war and peace that is 48,540 statement preparations to collect
138 rows, which is all those four tables hold between them. They are now
four ordered sweeps, kept by object id, so the cost is what the tables
hold rather than what the document holds.

Second, and even more consequentially: d2sqlite3's row["name"] is
indexForName, a linear scan over the statement's columns that calls
sqlite3_column_name and allocates a D string for every column it passes.
At some fifty named reads per object over forty columns that is around a
thousand of those per object, twelve million for the document. The
column name to index map is now resolved once per statement and the
reads are an integer index.

Also here, since it was measured while doing this: sspReadFile no longer
hashes the file it read unless asked (with_digest). Only the database
writer wants that digest and it has the lines already, so every other
read was paying for it.

Existing outputs unaffected.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
reading a .ocda.db was slower than parsing markup source. Taking the
largest markup document sample War and peace, 12,135 objects, optimised
build, before this commit and after:

  parse markup    0.87 s   0.87 s
  load .ssp       0.047 s  0.047 s
  load .ocda.db   1.09 s   0.128 s

The database goes from being 1.25x slower than parsing the document to
6.8x faster.

Two things fixed in the reader were:

First, four queries per object. object_images, object_links,
object_anchors and object_subtoc were queried per object as the objects
were built, each statement compiled fresh from a concatenated string.
For war and peace that is 48,540 statement preparations to collect
138 rows, which is all those four tables hold between them. They are now
four ordered sweeps, kept by object id, so the cost is what the tables
hold rather than what the document holds.

Second, and even more consequentially: d2sqlite3's row["name"] is
indexForName, a linear scan over the statement's columns that calls
sqlite3_column_name and allocates a D string for every column it passes.
At some fifty named reads per object over forty columns that is around a
thousand of those per object, twelve million for the document. The
column name to index map is now resolved once per statement and the
reads are an integer index.

Also here, since it was measured while doing this: sspReadFile no longer
hashes the file it read unless asked (with_digest). Only the database
writer wants that digest and it has the lines already, so every other
read was paying for it.

Existing outputs unaffected.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>ocda: the abstraction carries the document metadata output reads</title>
<updated>2026-09-12T16:14:39+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-11T14:23:25+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=b6b2ff06054edb20c880aff26b8b02d8f093fcc2'/>
<id>b6b2ff06054edb20c880aff26b8b02d8f093fcc2</id>
<content type='text'>
.ssp format v0.1 -&gt; v1.1, and the .ocda.db with it: they are two
serialisations of one format and now move together.

The gap this closes, measured against the writers' read set: of the 59
conf_make_meta fields outputs/io_out/ reads, 29 were in neither
artefact. Five are conf.* and stay out, being site and run scoped (urls,
papersize, the search db filename): the same abstraction published to
two sites must take each site's. Of the rest, three are never assigned
anywhere (title_short, publisher, original_publisher; read only by
sqlite.d, always empty) and one is a copy of a field already carried
(title_sub = title_subtitle), so 17 properties actually had to travel
and now do:

  make      breaks, footer, home_button_text
  meta      title.edition, date.added_to_site, language.document_char,
            original.{title,source,language,language_char},
            rights.copyright_{text,translation,illustrations,
            photographs,cover,audio,video}

make splits by when it acts, which is worth keeping in mind: these three
are read in outputs/io_out/ and must travel, while italics, bold,
emphasis, substitute and headings are read at parse time and their
effect is already in the objects.

New @source block, and source.* rows in the database, so a reader can
say which markup an abstraction came from rather than working from
stale content in silence:

  language    the document's own
  languages   the pod's list, which is what the inter-language links
              in html need and neither artefact carried
  digest      sha256 of the .sst; equals its digests.txt entry

The database adds source.ssp_digest, the sha256 of the .ssp it was built
from, since a file cannot hold its own hash. So the chain
.sst -&gt; .ssp -&gt; .ocda.db is checkable end to end.

The database's metadata table is now filled from the header blocks the
.ssp gives back, not from a second list read off doc_matters. One list,
in ssp.d: a property added there arrives in the database with nothing
else changed, and one that is not in the .ssp cannot be in the database
at all. That was the last place the two could drift; the objects stopped
being able to on 2026-09-07.

Version is checked on load. The major part must match, a newer minor is
accepted (a minor bump only adds properties, and an unknown property
line is ignored). A v0.1 artefact is now refused with a message saying
to regenerate it, rather than loading half populated.

test-abstraction-db.sh now compares the two artefacts' header blocks
property by property, 36 per document on the wealth of networks, where
it previously only checked that a schema.version row existed. Verified
non-vacuous: dropping one property is reported.

Reference regenerated, and the whole diff is this change and nothing
else: 320 lines added, 35 removed over 35 files, being 35 format lines
changed, 35 @source blocks (4 lines each), 35 language.document_char, 35
home_button_text (it has a default), 18 breaks, 16 footer,
4 date.added_to_site, 1 title.edition, 1 original.source.

Existing outputs unaffected.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
.ssp format v0.1 -&gt; v1.1, and the .ocda.db with it: they are two
serialisations of one format and now move together.

The gap this closes, measured against the writers' read set: of the 59
conf_make_meta fields outputs/io_out/ reads, 29 were in neither
artefact. Five are conf.* and stay out, being site and run scoped (urls,
papersize, the search db filename): the same abstraction published to
two sites must take each site's. Of the rest, three are never assigned
anywhere (title_short, publisher, original_publisher; read only by
sqlite.d, always empty) and one is a copy of a field already carried
(title_sub = title_subtitle), so 17 properties actually had to travel
and now do:

  make      breaks, footer, home_button_text
  meta      title.edition, date.added_to_site, language.document_char,
            original.{title,source,language,language_char},
            rights.copyright_{text,translation,illustrations,
            photographs,cover,audio,video}

make splits by when it acts, which is worth keeping in mind: these three
are read in outputs/io_out/ and must travel, while italics, bold,
emphasis, substitute and headings are read at parse time and their
effect is already in the objects.

New @source block, and source.* rows in the database, so a reader can
say which markup an abstraction came from rather than working from
stale content in silence:

  language    the document's own
  languages   the pod's list, which is what the inter-language links
              in html need and neither artefact carried
  digest      sha256 of the .sst; equals its digests.txt entry

The database adds source.ssp_digest, the sha256 of the .ssp it was built
from, since a file cannot hold its own hash. So the chain
.sst -&gt; .ssp -&gt; .ocda.db is checkable end to end.

The database's metadata table is now filled from the header blocks the
.ssp gives back, not from a second list read off doc_matters. One list,
in ssp.d: a property added there arrives in the database with nothing
else changed, and one that is not in the .ssp cannot be in the database
at all. That was the last place the two could drift; the objects stopped
being able to on 2026-09-07.

Version is checked on load. The major part must match, a newer minor is
accepted (a minor bump only adds properties, and an unknown property
line is ignored). A v0.1 artefact is now refused with a message saying
to regenerate it, rather than loading half populated.

test-abstraction-db.sh now compares the two artefacts' header blocks
property by property, 36 per document on the wealth of networks, where
it previously only checked that a schema.version row existed. Verified
non-vacuous: dropping one property is reported.

Reference regenerated, and the whole diff is this change and nothing
else: 320 lines added, 35 removed over 35 files, being 35 format lines
changed, 35 @source blocks (4 lines each), 35 language.document_char, 35
home_button_text (it has a default), 18 breaks, 16 footer,
4 date.added_to_site, 1 title.edition, 1 original.source.

Existing outputs unaffected.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>ocda db: built from the .ssp, (tethered)</title>
<updated>2026-09-09T21:56:14+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-08T12:28:24+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=9af1dda9b03f745aa2b7700419574e2e3e59fde2'/>
<id>9af1dda9b03f745aa2b7700419574e2e3e59fde2</id>
<content type='text'>
The ocda database is now built from the .ssp itself: the writer's
lines are emitted, read straight back by ssp_in, and the objects
that come out populate the database. Anything the .ssp does not
carry, the database will not have either, by construction (rather
than by test).
[instead of as previously through a second walk over the in-memory
abstraction]

- spineAbstractionTxt is split: sspDocumentLines(doc) returns the
  whole .ssp as lines, and the file writer emits them. Output
  neutral, the reference test confirms.
- spineAbstractionDb takes the abstraction as an argument rather
  than taking doc.abstraction.
- sspRoundTripAbstraction(doc) in ssp_in is the join: lines out,
  lines in, abstraction returned. Both call sites in spine.d use
  it.
- the header blocks and the image blobs still come from
  doc_matters (as: the .ssp does not carry image bytes).

All (35) markup sample sourced databases built through the .ssp
have byte identical SQL dumps to the one built directly before the
change.

That comparison also found one reader inaccuracy, which the .ssp
round trip could not see because the writer omits the field either
way: an absent identifier was restored as the ocn in every case,
but for an object with no ocn it was empty ("a"~N identifiers are
always written). Fixed; the two artefacts checking each other is
what caught it.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
The ocda database is now built from the .ssp itself: the writer's
lines are emitted, read straight back by ssp_in, and the objects
that come out populate the database. Anything the .ssp does not
carry, the database will not have either, by construction (rather
than by test).
[instead of as previously through a second walk over the in-memory
abstraction]

- spineAbstractionTxt is split: sspDocumentLines(doc) returns the
  whole .ssp as lines, and the file writer emits them. Output
  neutral, the reference test confirms.
- spineAbstractionDb takes the abstraction as an argument rather
  than taking doc.abstraction.
- sspRoundTripAbstraction(doc) in ssp_in is the join: lines out,
  lines in, abstraction returned. Both call sites in spine.d use
  it.
- the header blocks and the image blobs still come from
  doc_matters (as: the .ssp does not carry image bytes).

All (35) markup sample sourced databases built through the .ssp
have byte identical SQL dumps to the one built directly before the
change.

That comparison also found one reader inaccuracy, which the .ssp
round trip could not see because the writer omits the field either
way: an absent identifier was restored as the ocn in every case,
but for an object with no ocn it was empty ("a"~N identifiers are
always written). Fixed; the two artefacts checking each other is
what caught it.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>ocda: clean heading text used for navigation</title>
<updated>2026-09-09T21:45:07+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-08T02:30:03+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=da575bf8a80cc2900bd5984718485cdacca0f527'/>
<id>da575bf8a80cc2900bd5984718485cdacca0f527</id>
<content type='text'>
heading text used for navigation is normalised, and | escaped

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
heading text used for navigation is normalised, and | escaped

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
<entry>
<title>ocda: a reader for .ssp, and a round trip check</title>
<updated>2026-09-09T21:30:06+00:00</updated>
<author>
<name>Ralph Amissah</name>
<email>ralph.amissah@gmail.com</email>
</author>
<published>2026-09-07T15:18:52+00:00</published>
<link rel='alternate' type='text/html' href='https://git.sisudoc.org/projects/sisudoc-spine/commit/?id=28b27ec45b9a592be3187d995b4355f483e53cc1'/>
<id>28b27ec45b9a592be3187d995b4355f483e53cc1</id>
<content type='text'>
sisudoc.ocda.abstraction.ssp_in reads a .ssp file back into
ObjGenericComposite[][string], the same value the parser produces,
so anything that consumes the abstraction can be fed from a .ssp
instead of from markup. The three header blocks come back as
key/value with their order preserved.

--ssp-round-trip=&lt;file.ssp&gt; loads a file and emits it again on
stdout, using sspObjectRecord, the writer's own definition of a
record. So the check is against the writer, not against a second
description of the format:

  ./bin/spine-ldc --ssp-round-trip=test/reference/abstraction/&lt;doc&gt;.ssp \
    | diff test/reference/abstraction/&lt;doc&gt;.ssp -

BUG as yet to FIX

27 of the 35 reference documents round trip byte identically. The
other 8 fail on two defects in the *writer* that the round trip
found, and which are left for a decision:

1. .heading_ancestors_text and .lev4_subtoc can carry a raw newline,
   because a heading's text may contain a line break. The value then
   spans two physical lines and the format's rule that a value runs to
   the end of the line is broken. 194 and 22 occurrences, in the seven
   live-manual translations.
2. .heading_ancestors_text joins its eight slots with "|" while the
   text in them may itself contain "|". 22 occurrences in
   revisiting_the_autonomous_contract.

Both need an escape (or normalisation at source) and both change the
.ssp, so require a decision and another reference regeneration.

ocda: export the .ssp reader from the abstraction package

package.d is the re-export surface for consumers that want to reach
the abstraction without depending on the directory layout; the reader
belongs there beside the writer.

(assisted by Claude-Code)
</content>
<content type='xhtml'>
<div xmlns='http://www.w3.org/1999/xhtml'>
<pre>
sisudoc.ocda.abstraction.ssp_in reads a .ssp file back into
ObjGenericComposite[][string], the same value the parser produces,
so anything that consumes the abstraction can be fed from a .ssp
instead of from markup. The three header blocks come back as
key/value with their order preserved.

--ssp-round-trip=&lt;file.ssp&gt; loads a file and emits it again on
stdout, using sspObjectRecord, the writer's own definition of a
record. So the check is against the writer, not against a second
description of the format:

  ./bin/spine-ldc --ssp-round-trip=test/reference/abstraction/&lt;doc&gt;.ssp \
    | diff test/reference/abstraction/&lt;doc&gt;.ssp -

BUG as yet to FIX

27 of the 35 reference documents round trip byte identically. The
other 8 fail on two defects in the *writer* that the round trip
found, and which are left for a decision:

1. .heading_ancestors_text and .lev4_subtoc can carry a raw newline,
   because a heading's text may contain a line break. The value then
   spans two physical lines and the format's rule that a value runs to
   the end of the line is broken. 194 and 22 occurrences, in the seven
   live-manual translations.
2. .heading_ancestors_text joins its eight slots with "|" while the
   text in them may itself contain "|". 22 occurrences in
   revisiting_the_autonomous_contract.

Both need an escape (or normalisation at source) and both change the
.ssp, so require a decision and another reference regeneration.

ocda: export the .ssp reader from the abstraction package

package.d is the re-export surface for consumers that want to reach
the abstraction without depending on the directory layout; the reader
belongs there beside the writer.

(assisted by Claude-Code)
</pre>
</div>
</content>
</entry>
</feed>
