diff options
| author | Ralph Amissah <ralph.amissah@gmail.com> | 2026-09-21 13:16:28 -0400 |
|---|---|---|
| committer | Ralph Amissah <ralph.amissah@gmail.com> | 2026-09-22 15:30:59 -0400 |
| commit | 36c4b1721ce99865d10b9dd74c57290869efb383 (patch) | |
| tree | 112df80c451b230688adfbde772f8731ad857ac9 /src/sisudoc | |
| parent | test: the carried markup, by rebuilding its digest (diff) | |
ocda db: materialise a pod from carried source
A database that carries its source can recreate the pod. --source and
--pod2 given a .ocda.db now write the pod to a directory of the run's
making and carry on with the pod, so everything after that point is
handling a pod like any other.
Nothing renders from the database. A materialiser that also rendered
would be a second path to every output format and the two would drift;
one that only writes files means the document is built by the same
code over the same bytes as the original, and identical output is a
consequence rather than an aspiration. Held against the original pod,
site configuration constant: every output file identical across ten
languages.
The pod's name comes from the database's filename, which inverts the
naming rule exactly: <doc>.ocda.db is named by doc_uid_out_no_lang,
the pod name and the document's filename joined by ":" when they
differ and the one name when they do not. So the materialised pod
recomputes the uid it was named by and every output file lands on the
name it had.
Names are checked before anything is created, and one bad name refuses
the artefact rather than skipping a file, as the zip reader does with
a zip. Markup that does not match the digest stored with it is refused
outright, where a mismatched image is written with a warning: a wrong
image makes a document that looks wrong, a wrong markup file makes one
that is wrong, in its text, with nothing downstream to notice.
(assisted by Claude-Code)
Diffstat (limited to 'src/sisudoc')
| -rw-r--r-- | src/sisudoc/ocda/abstraction/pod_from_db.d | 231 | ||||
| -rw-r--r-- | src/sisudoc/spine.d | 53 |
2 files changed, 284 insertions, 0 deletions
diff --git a/src/sisudoc/ocda/abstraction/pod_from_db.d b/src/sisudoc/ocda/abstraction/pod_from_db.d new file mode 100644 index 0000000..0bdb3bc --- /dev/null +++ b/src/sisudoc/ocda/abstraction/pod_from_db.d @@ -0,0 +1,231 @@ +/+ +- Name: SisuDoc Spine, Doc Reform [a part of] + - Description: documents, structuring, processing, publishing, search + - static content generator + + - Author: Ralph Amissah + [ralph.amissah@gmail.com] + + - Copyright: (C) 2015 (continuously updated, current 2026) Ralph Amissah, All Rights Reserved. + + - License: AGPL 3 or later: + + Spine (SiSU), a framework for document structuring, publishing and + search + + Copyright (C) Ralph Amissah + + This program is free software: you can redistribute it and/or modify it + under the terms of the GNU AFERO General Public License as published by the + Free Software Foundation, either version 3 of the License, or (at your + option) any later version. + + This program is distributed in the hope that it will be useful, but WITHOUT + ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or + FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for + more details. + + You should have received a copy of the GNU General Public License along with + this program. If not, see [https://www.gnu.org/licenses/]. + + If you have Internet connection, the latest version of the AGPL should be + available at these locations: + [https://www.fsf.org/licensing/licenses/agpl.html] + [https://www.gnu.org/licenses/agpl.html] + + - Spine (by Doc Reform, related to SiSU) uses standard: + - docReform markup syntax + - standard SiSU markup syntax with modified headers and minor modifications + - docReform object numbering + - standard SiSU object citation numbering & system + + - Homepages: + [https://www.sisudoc.org] + [https://www.doc-reform.org] + + - Git + [https://git.sisudoc.org/] + ++/ +/++ + a pod, written back out of a database<br><br> + . + the inverse of the carrying: markup, configuration, manifest and images + out of a .ocda.db and onto the filesystem as a pod tree<br><br> + . + [sisudoc.ocda.abstraction.pod_from_db] ++/ +module sisudoc.ocda.abstraction.pod_from_db; +@safe: +/+ ↓ a database that carries its source can give the pod back. + . + This writes the tree and stops. It does not parse, and it does not produce + output: what it leaves behind is a pod directory like any other, and what + happens next is whatever would have happened to the original. That is the + intent. A materialiser that also rendered would be a second path to every + output format, and would provide opportunity for the two to drift; a + materialiser that only writes files means the document is built by the + same code over the same bytes, and "identical output" is guaranteed as a + consequence (rather than an aspirational). + . + What is written: + . + <dest>/<pod>/pod.manifest role='manifest' + <dest>/<pod>/conf/document_make role='conf' + <dest>/<pod>/media/text/<lang>/<file> role='source' + <dest>/<pod>/media/image/<file> role='image' + . + The pod's name comes from the database's own filename: <doc>.ocda.db is + named by doc_uid_out_no_lang, which is the pod name and the document's + filename joined by ":" when they differ and the one name when they do not. + Splitting on ":" inverts that rule exactly, so a pod written here + recomputes the uid the database was named by, and every output file lands + on the name it had before. Rename the database and the pod is named + accordingly, which is the behaviour to expect of a file named for its + document. + . + Every name is checked before anything is created, and every file is checked + against the digest stored with it. A database can be downloaded, so its + names are attacker controlled: a name that climbs out of the pod is + refused, and one bad name refuses the whole artefact rather than skipping + one file, which is what the zip reader does with a zip. A digest that does + not match is refused outright here, unlike the images, which are written + with a warning: an image that is wrong makes a document that looks wrong, + while markup that is wrong makes a document that *is* wrong, silently and + in its text. ++/ +template spinePodFromDb() { + /+ ↓ narrow imports: this is mixed in, and both std.file and std.stdio + define write, which is an ambiguity the mixing scope inherits + +/ + import std.array : array; + import std.conv : to; + import std.file : exists, mkdirRecurse, fileWrite = write; + import std.path : baseName, chainPath, dirName; + import std.stdio : writeln; + import sisudoc.ocda.io_in.carried_names; + import sisudoc.ocda.abstraction.db_in : spineAbstractionDbRead; + mixin spineCarriedNames; + mixin spineAbstractionDbRead _dbr; + struct ST_PodMaterialised { + string pod_dir; // where the pod was written, "" when it was not + string note; // why not, when it was not + size_t files; // how many were written + bool ok; + } + /+ ↓ the pod's name, from the database's own filename +/ + string podNameFromDbPath(string _db_file) { + import std.algorithm : endsWith, findSplit; + import sisudoc.ocda.meta.defaults : InternalMarkup; + mixin InternalMarkup _mkup_; + auto _mkup = _mkup_.InlineMarkup(); + string _stem = _db_file.baseName; + foreach (_sfx; [".ocda.db", ".db"]) { + if (_stem.endsWith(_sfx)) { _stem = _stem[0 .. $ - _sfx.length]; break; } + } + if (auto _s = _stem.findSplit(_mkup.uid_sep)) { return _s[0]; } + return _stem; + } + /+ ↓ where each role is written within the pod. + image is the one role whose rows are named by bare filename, because that + is how the abstraction refers to an image; the rest carry their path + within the pod and are written at it. + +/ + private string _relPathFor(string _role, string _name) { + switch (_role) { + case "image": return "media/image/" ~ _name; + case "source": + case "conf": + case "manifest": return _name; + default: return ""; + } + } + @trusted ST_PodMaterialised podFromDb(O)( + string _db_file, + string _dest_root, + O _opt_action, + ) { + import std.digest : toHexString; + import std.digest.sha : sha256Of; + ST_PodMaterialised _out; + if (!_db_file.exists) { + _out.note = "no such file"; + return _out; + } + _dbr.ST_ArtefactFile[] _files; + foreach (_role; ["manifest", "conf", "source", "image"]) { + _files ~= _dbr.dbReadFiles(_db_file, _role); + } + if (_files.length == 0) { + _out.note = "carries no source: written by a spine older than the" + ~ " format that carries markup, or written from an artefact"; + return _out; + } + bool _has_source = false; + foreach (_f; _files) { if (_f.role == "source") { _has_source = true; } } + if (!_has_source) { + _out.note = "carries images but no markup, so no pod can be written" + ~ " from it"; + return _out; + } + string _pod_dir = (_dest_root.chainPath(podNameFromDbPath(_db_file)) + .array).to!string; + /+ ↓ every name checked before anything is created or written +/ + foreach (_f; _files) { + string _rel = _relPathFor(_f.role, _f.name); + if (_rel.length == 0) { + _out.note = "carries a file of a role spine does not write: " + ~ _f.role; + return _out; + } + string _bad = (_f.role == "image") + ? validateCarriedFileName(_f.name) + : validateCarriedPath(_f.name); + if (_bad.length > 0) { + _out.note = "carries a file spine will not write: " ~ _bad; + return _out; + } + } + /+ ↓ and every digest, before anything is created or written. markup that + does not match what was recorded with it is refused, not warned about: + the difference would be in the text of the document and nothing + downstream would notice. + +/ + foreach (_f; _files) { + if (_f.sha256.length == 0) { continue; } + string _got = _f.data.sha256Of.toHexString.to!string; + if (_got != _f.sha256) { + _out.note = _f.role ~ " " ~ _f.name + ~ " does not match the digest recorded with it (" + ~ _f.sha256 ~ " expected, " ~ _got ~ " found)"; + return _out; + } + } + foreach (_f; _files) { + string _path = (_pod_dir.chainPath(_relPathFor(_f.role, _f.name)) + .array).to!string; + /+ ↓ the check on the check: the name rules above already forbid a + path that climbs, so this can only fire if they were loosened + +/ + if (!(carriedPathIsWithin(_pod_dir, _path))) { + _out.note = _f.name ~ " resolves outside the pod"; + return _out; + } + try { + _path.dirName.mkdirRecurse; + fileWrite(_path, _f.data); + } catch (Exception ex) { + _out.note = "could not write " ~ _f.name ~ ": " ~ ex.msg; + return _out; + } + _out.files += 1; + } + _out.pod_dir = _pod_dir; + _out.ok = true; + if (_opt_action.vox_gt_1) { + writeln(" pod from ", _db_file.baseName, ": ", _out.files, + " file(s) in ", _out.pod_dir); + } + return _out; + } +} diff --git a/src/sisudoc/spine.d b/src/sisudoc/spine.d index 1286c24..bb7dbb8 100644 --- a/src/sisudoc/spine.d +++ b/src/sisudoc/spine.d @@ -1094,6 +1094,51 @@ string program_name = "spine"; _resolved_args ~= arg; } } + /+ ↓ a database asked for as a source: write the pod back out of it and carry on + with the pod. + . + --source and --pod2 need the original markup file, included in a 2.0 + database. The pod is written to a directory of this run's making and takes + the place of the argument, so everything after this point is handling a + pod, and the document is built by the same code over the same bytes as the + original. That is what makes "identical output" a consequence. + . + Before the config discovery below, because that walks the argument looking + for a .dr/ above it, and the argument it should walk is the pod rather than + the database. + . + A database without original markup still refuses, as it has to: no artefact + written by an older spine carries any. + +/ + string[] _pod_materialisations; + if (_opt_action.source_or_pod) { + import sisudoc.ocda.abstraction.pod_from_db; + import std.process : thisProcessID; + mixin spinePodFromDb; + string[] _args_after; + foreach (arg; _resolved_args) { + if (!(arg.endsWith(".ocda.db") || arg.endsWith(".db")) + || arg.match(rgx.flag_action) + ) { + _args_after ~= arg; + continue; + } + string _root = (tempDir.chainPath("spine-pod-" + ~ arg.baseName ~ "-" ~ thisProcessID.to!string).array).to!string; + auto _mat = podFromDb!()(arg, _root, _opt_action); + if (_mat.ok) { + _pod_materialisations ~= _root; + _args_after ~= _mat.pod_dir; + if (_opt_action.vox_gt_1) { + writeln("pod from database: ", arg.baseName, " -> ", _mat.pod_dir); + } + } else { + stderr.writeln("WARNING: --source and --pod2 need the markup, and ", + arg.baseName, " ", _mat.note, "; skipped"); + } + } + _resolved_args = _args_after; + } ConfComposite _siteConfig; if ( _opt_action.require_processing_files @@ -1872,6 +1917,14 @@ string program_name = "spine"; foreach (ref _dlr; _url_downloads) { cleanupDownload(_dlr); } + /+ ↓ and any pod written out of a database for this run +/ + foreach (_root; _pod_materialisations) { + try { + if (_root.exists) { _root.rmdirRecurse; } + } catch (Exception ex) { + stderr.writeln("WARNING: could not remove ", _root, ": ", ex.msg); + } + } /+ ↓ --strict and a document whose languages diverged: fatal to the run. At the end rather than where the check runs, so that the outputs of the run are complete and can be looked at. The warning has already |
