diff options
| author | Ralph Amissah <ralph.amissah@gmail.com> | 2026-09-20 23:34:25 -0400 |
|---|---|---|
| committer | Ralph Amissah <ralph.amissah@gmail.com> | 2026-09-21 17:16:47 -0400 |
| commit | 3ac7f9801255b0440647e83119524efb95f0ada7 (patch) | |
| tree | 76e1ced617a8f44eb0f100a6fcc291890f92ae0b /src | |
| parent | things nix, including flake.nix update (diff) | |
markup: keep a utf-8 bom out of the yaml header
A file beginning with a utf-8 byte order mark had the mark read as part
of its first yaml key, so "title:" arrived as "title:" and every
value under that key was dropped silently: the document kept its text
and lost its title. Downstream that showed as an empty dc:title and an
empty <title> element in the epub, which epubcheck reports as two
errors.
The mark is now kept out of the header and body split, and so out of
the parse. It is not stripped when the file is read, because that text
is what source.digest is taken over and the digest names the file as it
sits on disk.
One document in the sample set begins with a mark. It is left as it is:
it is the case this guards against.
(assisted by Claude-Code)
Diffstat (limited to 'src')
| -rw-r--r-- | src/sisudoc/ocda/io_in/read_source_files.d | 15 |
1 files changed, 14 insertions, 1 deletions
diff --git a/src/sisudoc/ocda/io_in/read_source_files.d b/src/sisudoc/ocda/io_in/read_source_files.d index f99bb9b..4ceb039 100644 --- a/src/sisudoc/ocda/io_in/read_source_files.d +++ b/src/sisudoc/ocda/io_in/read_source_files.d @@ -188,7 +188,20 @@ template spineRawMarkupContent() { @trusted final private char[][] header0Content1(in string src_text) { // cast(char[]) /+ split string on _first_ match of "^:?A~\s" into [header, content] array/tuple +/ char[][] header_and_content; - auto m = (cast(char[]) src_text).matchFirst(rgx.heading_a); + /+ ↓ a utf-8 byte order mark at the start of the file is kept out of + the split, and so out of the yaml header. Left in, it becomes + part of the first key, so "title:" arrives as "title:" and + every value under that key is silently lost: the document keeps + its text and loses its title. It is not stripped at read time, + because source_txt_str is what source.digest is taken over and + that digest names the file as it sits on disk. + +/ + string _src = src_text; + enum string _bom = ""; + if (_src.length >= _bom.length && _src[0.._bom.length] == _bom) { + _src = _src[_bom.length..$]; + } + auto m = (cast(char[]) _src).matchFirst(rgx.heading_a); header_and_content ~= m.pre; header_and_content ~= m.hit ~ m.post; assert(header_and_content.length == 2, |
