aboutsummaryrefslogtreecommitdiffhomepage
path: root/src/sisudoc/ocda/io_in/read_zip_pod.d
diff options
context:
space:
mode:
authorRalph Amissah <ralph.amissah@gmail.com>2026-09-21 09:46:26 -0400
committerRalph Amissah <ralph.amissah@gmail.com>2026-09-22 14:03:39 -0400
commit373a5a267cfe849e184fda0f9baac3bf81f77223 (patch)
tree29b3131488679bc2c38471e88d3adb40e956245f /src/sisudoc/ocda/io_in/read_zip_pod.d
parentpod: write <doc>.sisupod (single zstd frame) (diff)
pod: set entry limit for the writer and reader
Both now take the number from one constant now set at 10,000, the writer checks it before the archive is written. Previously the number set for the pod reader was less than 500 members. The writer had no limit, (so spine could write a pod it would then refuse to read, reporting too many entries: a good file that looks corrupt, and only for a document with enough parts). 500 was set when a pod held markup and images. A pod carrying translation catalogues has a different arithmetic, and the count follows from how many files a document is made of times how many languages it has, not from how large the document is: the_wealth_of_networks 213,405 words, 1 file per language 12 languages -> 49 entries live-manual 24,724 words, 20 files per language 12 languages -> 506 entries The big book is not what runs into this; the modular manual is. live-manual stood two languages from an unreadable artefact. 10,000 covers a hundred-insert manual in thirty languages, about 6,200 entries, with room. It is deliberately well clear of any real document rather than snug above the largest one known: the two errors are not comparable, since too low refuses a legitimate document with a message that reads as corruption, while too high defers to a size cap a moment later. The count is the weakest of the three guards and is not what bounds resource use; the per entry and total size caps do that, both before a byte is written. Refusing also removes any archive an earlier language left. The writer runs once per language and only the last pass holds every language, so it is the last that goes over, and the passes before it wrote smaller archives that passed. Without this, refusal left a pod missing a language: an artefact that reads perfectly well and is wrong. A missing file is an error someone notices. The arithmetic is recorded beside the constant so the next person can re-derive the number rather than guess at it. (assisted by Claude-Code)
Diffstat (limited to 'src/sisudoc/ocda/io_in/read_zip_pod.d')
-rw-r--r--src/sisudoc/ocda/io_in/read_zip_pod.d6
1 files changed, 5 insertions, 1 deletions
diff --git a/src/sisudoc/ocda/io_in/read_zip_pod.d b/src/sisudoc/ocda/io_in/read_zip_pod.d
index d0eb65a..18a957c 100644
--- a/src/sisudoc/ocda/io_in/read_zip_pod.d
+++ b/src/sisudoc/ocda/io_in/read_zip_pod.d
@@ -64,11 +64,15 @@ template spineExtractZipPod() {
import std.stdio;
import std.string : indexOf;
import sisudoc.ocda.zstd;
+ import sisudoc.ocda.io_in.carried_names : MAX_POD_ENTRY_COUNT;
/+ security limits for zip extraction +/
enum size_t MAX_ENTRY_SIZE = 50 * 1024 * 1024; /+ 50 MB per entry +/
enum size_t MAX_TOTAL_SIZE = 500 * 1024 * 1024; /+ 500 MB total +/
- enum size_t MAX_ENTRY_COUNT = 500; /+ max entries in archive +/
+ /+ ↓ the entry count now comes from carried_names, so that the writer is
+ held to the same number the reader enforces.
+ +/
+ alias MAX_ENTRY_COUNT = MAX_POD_ENTRY_COUNT;
enum size_t MAX_PATH_DEPTH = 10; /+ max path components +/
/+ allowed entry name pattern: alphanumeric, dots, dashes, underscores, forward slashes +/