diff options
| author | Ralph Amissah <ralph.amissah@gmail.com> | 2026-09-11 14:26:09 -0400 |
|---|---|---|
| committer | Ralph Amissah <ralph.amissah@gmail.com> | 2026-09-12 12:14:40 -0400 |
| commit | 25d3ab8e6524697ab3f2972f6dba1fbe26fcb0ec (patch) | |
| tree | e524e09405ea3956b9838e03b0c46e9b0575f631 /src/sisudoc/ocda | |
| parent | ocda: the images an artefact carries, and the ones it only describes (diff) | |
spine: a .ocda.db can be fetched, as a pod zip can
A URL ending in .ocda.db is now downloaded and processed in place,
through the path that already did it for .zip: the same curl call, the
same size and timeout limits, the same refusal of local and private
addresses, the same --allow-downloads guard, the same temp file cleanup.
Only the pattern had to widen.
rgx_url_zip ^https?://...[.]zip$
rgx_url_source ^https?://...([.]zip|[.]ocda[.]db)$
downloadZipUrl is no downloadSourceUrl and the download temp directory
is spine-download rather than spine-zip-pod; (the extraction directory
keeps its name, being still only for zips).
(assisted by Claude-Code)
Diffstat (limited to 'src/sisudoc/ocda')
| -rw-r--r-- | src/sisudoc/ocda/abstraction/load.d | 34 | ||||
| -rw-r--r-- | src/sisudoc/ocda/io_in/read_zip_pod.d | 15 |
2 files changed, 23 insertions, 26 deletions
diff --git a/src/sisudoc/ocda/abstraction/load.d b/src/sisudoc/ocda/abstraction/load.d index ebdd1a5..ea387d1 100644 --- a/src/sisudoc/ocda/abstraction/load.d +++ b/src/sisudoc/ocda/abstraction/load.d @@ -50,38 +50,30 @@ module sisudoc.ocda.abstraction.load; @safe: /+ ↓ one way in, whatever the document is being read from - spine's pipeline is markup -> abstraction -> output. once the abstraction is serialised, there is more than one thing an abstraction can be read from, and a consumer should not have to know which it was handed: - + . .sst / .ssm + images the markup source pod (dir) + images the same, bundled pod .zip the same, zipped .ssp + images the abstraction, as text .ocda.db the abstraction, sqlite, images inside - - A source may also be remote. spine already takes a URL argument ending - in .zip: it is downloaded to a temp file (downloadZipUrl in - ocda/io_in/read_zip_pod.d, guarded by rgx_url_zip and by --allow- - downloads) and the local path is processed in its place. The same should - be offered for .ocda.db, which is the one other artefact that is - self-sufficient enough to be fetched on its own: it carries its images. - What that needs is small and is not done here: - - - rgx_url_zip is `^https?://...[.]zip$`; a sibling pattern for - `[.]ocda[.]db$`, or one pattern covering both - - the same download path, which is already generic apart from that - regex, and the same temp-file cleanup - - dispatch, which this module already does once the file is local - - A .ssp URL would want its images too, so it is a pod or a database that - travels, not a .ssp on its own. - + . + A source may also be remote. A URL ending in .zip or in .ocda.db is + downloaded to a temp file (downloadSourceUrl in + ocda/io_in/read_zip_pod.d, guarded by rgx_url_source and by + --allow-downloads) and the local path is processed in its place. + + Those two travel because each is self-sufficient: a pod holds its + source and images, a .ocda.db holds its abstraction and images. A .ssp + does not, describing its images by name and digest without carrying + them, so it would arrive without them and is not fetched. + . This module names those sources, tells them apart, and loads the two that are self-describing artefacts, returning the value the parser produces. - + . The three source forms are deliberately *not* loaded here. They are the parser's job (sisudoc.ocda.meta.metadoc, spineAbstraction), and it needs the environment, the options and the configuration that spine.d diff --git a/src/sisudoc/ocda/io_in/read_zip_pod.d b/src/sisudoc/ocda/io_in/read_zip_pod.d index 322382d..ea1fa1b 100644 --- a/src/sisudoc/ocda/io_in/read_zip_pod.d +++ b/src/sisudoc/ocda/io_in/read_zip_pod.d @@ -269,7 +269,12 @@ template spineExtractZipPod() { enum size_t MAX_DOWNLOAD_SIZE = 200 * 1024 * 1024; /+ 200 MB download limit +/ enum int DOWNLOAD_TIMEOUT = 120; /+ seconds +/ - static auto rgx_url_zip = ctRegex!(`^https?://[a-zA-Z0-9._:/-]+[.]zip$`); + /+ ↓ the sources that can be fetched whole: a pod zip, and a .ocda.db, + which is the one other artefact self-sufficient enough to travel on + its own, carrying its images. A .ssp is not here: it describes its + images without carrying them, so it would arrive without them. + +/ + static auto rgx_url_source = ctRegex!(`^https?://[a-zA-Z0-9._:/-]+([.]zip|[.]ocda[.]db)$`); struct DownloadResult { string local_path; /+ path to downloaded temp file +/ @@ -282,7 +287,7 @@ template spineExtractZipPod() { && (arg[0..8] == "https://" || arg[0..7] == "http://"); } - @trusted DownloadResult downloadZipUrl(string url) { + @trusted DownloadResult downloadSourceUrl(string url) { import std.process : execute, environment; DownloadResult result; result.ok = false; @@ -295,8 +300,8 @@ template spineExtractZipPod() { stderr.writeln("WARNING: downloading over insecure http: ", url); } /+ ↓ validate URL format +/ - if (!(url.matchFirst(rgx_url_zip))) { - result.error_msg = "URL does not match expected zip URL pattern: " ~ url; + if (!(url.matchFirst(rgx_url_source))) { + result.error_msg = "URL is not a .zip pod or a .ocda.db: " ~ url; return result; } /+ ↓ reject URLs that could target internal services +/ @@ -336,7 +341,7 @@ template spineExtractZipPod() { return result; } /+ ↓ create temp directory for download +/ - string tmp_base = tempDir.buildPath("spine-zip-pod"); + string tmp_base = tempDir.buildPath("spine-download"); try { if (!exists(tmp_base)) mkdirRecurse(tmp_base); |
