Ebooks and documents
Ebooks and documents have no backend. SoundStorm reads library/ebooks and
library/documents itself, with localbooks as the source and two small readers,
internal/epub and internal/pdf, that read just enough of each file to describe
it. An optional OPDS source can still point at a Calibre server running elsewhere.
Why no backend
Calibre-Web was the original ebook backend, and was removed: it needed a database rather than a folder, its setup had no API (CSRF form-scraping), and its password could not be rotated. Reading the folder directly is justified by one property: an EPUB is self-describing - its title, author, language and cover are in documented XML inside it - and a book needs no transcoding. A PDF needs none either. That is the whole precedent, and it should not be cited for anything else: video is never self-describing, and photos fail on HEIC.
localbooks
internal/source/localbooks is one source type, instantiated twice:
localbooks.Config.Kind picks EPUB-and-PDF for ebooks or PDF-only for
documents. A document is opened exactly as a PDF book is; the difference is which shelf it
is on, and so which permission and rescan it belongs to.
- Scanning.
Startruns the first scan before provisioning health-checks it; after that a ticker rescans every two minutes (DefaultRescanInterval), and an upload asks for one at once. It keeps what it learnt in memory - there is nothing to provision and nothing on disk to go stale. - Search and browse are a match over that list. Browsing sorts by title and then path before cutting a page, so paging through the merge is exact.
- Serving. File targets come from SoundStorm's own disk (
Target.FilePath) and covers extracted from inside an EPUB from memory (Target.Bytes).BookOpenerserves a book's inner files to the reader. - Deletion is the question that cannot arise: the folder is read on every scan, so a deleted book is simply gone from the next one.
readSafely, which turns a panic into an ordinary "could not read",
and the whole walk runs inside scanRecovered.internal/epub
An EPUB is a zip with a META-INF/container.xml naming an OPF package file, whose Dublin
Core metadata gives title, author, language, subjects and the cover. internal/epub reads
exactly that. Calibre writes a metadata.opf beside every book in the same format, so
ParseOPF reads both - which is how an existing Calibre library works with no SQLite driver,
series and series index included. The sidecar wins where it exists, since it holds curated metadata.
It is written defensively because the input is a stranger's file:
| Guard | Why |
|---|---|
| Central directory measured first, over 2 MB refused, any zip64 trigger refused | zip.OpenReader parses the directory from its start to the end of the file whatever the entry count claims, at about four times its size; archive/zip also switches to zip64 when a 16-bit field reads 0xFFFF. |
16 MB per entry (maxResourceBytes) | Zip-bomb protection for anything the reader fetches. |
| 2 MB for the OPF and the sidecar | The sidecar was once an unbounded read on every scan. |
| A few books open at once | Bounded memory under concurrent readers. |
| Covers must be declared images, never SVG | A cover declared as x.js was once served as a script. |
internal/pdf
internal/pdf deliberately does not parse PDF structure. Doing it properly means the
cross-reference table, then cross-reference streams, then object streams - hundreds of lines before the
first title. Measured against five real PDFs from five producers (pdfTeX, Gutenberg, Adobe Designer,
Ghostscript, PDFsam), all of it would have returned nothing: every Info dictionary was absent or literally
/Title ().
Two of the five carried an XMP packet - Dublin Core, plain XML, uncompressed - which a byte scan and
encoding/xml reach for free, in the vocabulary internal/epub already speaks. So
the order is: XMP, then the Info dictionary if it happens to be readable, then the file name. For PDFs the
file name is a primary source - "Title - Author (2017).pdf" is often the only information anyone wrote
down - and it is always the dropped name that is parsed, never the staged upload's
part-123456789.
- XMP nests every value inside
rdf:Altorrdf:Seqcontainingrdf:li, so readingdc:titleas a plain string yields an empty one. That is exactly how the five files first looked as if they had no metadata at all. - Packets over 1 MB are not decoded (
maxXMP): a 16 MB PDF of nothing but empty<Description/>s took 2.6 GB to decode, on upload and every scan. - No covers: extracting one means rendering page one, which needs a PDF renderer this project does not carry. No position either: PDFs open in the browser's own viewer, which does not expose one.
Which shelf a PDF belongs on
Deciding book or document is decided at upload, with evidence first and a question only without it: a
PDF beside a metadata.opf or an EPUB is a book (Calibre), one under Books or
Calibre is a book, one under Papers, Manuals or Taxes
is a document - whole folder names only, so "Paperback Writer" is not a paper. A loose PDF is asked about
once per dropped folder. Ebooks are filed Author/Title/; documents keep the shape they were
dropped in. See the library.
The optional OPDS source
internal/source/opds and internal/provision/calibreweb.go remain as an opt-in
for a Calibre server running somewhere else, via SOUNDSTORM_CALIBREWEB_URL. SoundStorm no
longer runs that server, so it cannot vouch for it: a log line used to call its published default password
safe "because it is unreachable except through SoundStorm", which stopped being true when the bundled
container went, and now says the password is safe only if the operator's own server is. OPDS ids are the
client's, fetched with the server's credential, so an id must stay under the server's opds/
tree with no query of its own. A remote catalog cannot serve a book's insides, so the in-browser reader is
for local books only.