Ebooks and documents

Ebooks and documents have no backend. SoundStorm reads library/ebooks and library/documents itself, with localbooks as the source and two small readers, internal/epub and internal/pdf, that read just enough of each file to describe it. An optional OPDS source can still point at a Calibre server running elsewhere.

Why no backend

Calibre-Web was the original ebook backend, and was removed: it needed a database rather than a folder, its setup had no API (CSRF form-scraping), and its password could not be rotated. Reading the folder directly is justified by one property: an EPUB is self-describing - its title, author, language and cover are in documented XML inside it - and a book needs no transcoding. A PDF needs none either. That is the whole precedent, and it should not be cited for anything else: video is never self-describing, and photos fail on HEIC.

localbooks

internal/source/localbooks is one source type, instantiated twice: localbooks.Config.Kind picks EPUB-and-PDF for ebooks or PDF-only for documents. A document is opened exactly as a PDF book is; the difference is which shelf it is on, and so which permission and rescan it belongs to.

This is the one parser that ordinary members can feed directly - anyone with upload access drops a file, and the scan parses it on a ticker with nobody watching. A panic there, unrecovered, would crash the server on that file and again on every restart, because the file is still on disk. So each file is read through readSafely, which turns a panic into an ordinary "could not read", and the whole walk runs inside scanRecovered.

internal/epub

An EPUB is a zip with a META-INF/container.xml naming an OPF package file, whose Dublin Core metadata gives title, author, language, subjects and the cover. internal/epub reads exactly that. Calibre writes a metadata.opf beside every book in the same format, so ParseOPF reads both - which is how an existing Calibre library works with no SQLite driver, series and series index included. The sidecar wins where it exists, since it holds curated metadata.

It is written defensively because the input is a stranger's file:

GuardWhy
Central directory measured first, over 2 MB refused, any zip64 trigger refusedzip.OpenReader parses the directory from its start to the end of the file whatever the entry count claims, at about four times its size; archive/zip also switches to zip64 when a 16-bit field reads 0xFFFF.
16 MB per entry (maxResourceBytes)Zip-bomb protection for anything the reader fetches.
2 MB for the OPF and the sidecarThe sidecar was once an unbounded read on every scan.
A few books open at onceBounded memory under concurrent readers.
Covers must be declared images, never SVGA cover declared as x.js was once served as a script.

internal/pdf

internal/pdf deliberately does not parse PDF structure. Doing it properly means the cross-reference table, then cross-reference streams, then object streams - hundreds of lines before the first title. Measured against five real PDFs from five producers (pdfTeX, Gutenberg, Adobe Designer, Ghostscript, PDFsam), all of it would have returned nothing: every Info dictionary was absent or literally /Title ().

Two of the five carried an XMP packet - Dublin Core, plain XML, uncompressed - which a byte scan and encoding/xml reach for free, in the vocabulary internal/epub already speaks. So the order is: XMP, then the Info dictionary if it happens to be readable, then the file name. For PDFs the file name is a primary source - "Title - Author (2017).pdf" is often the only information anyone wrote down - and it is always the dropped name that is parsed, never the staged upload's part-123456789.

Which shelf a PDF belongs on

Deciding book or document is decided at upload, with evidence first and a question only without it: a PDF beside a metadata.opf or an EPUB is a book (Calibre), one under Books or Calibre is a book, one under Papers, Manuals or Taxes is a document - whole folder names only, so "Paperback Writer" is not a paper. A loose PDF is asked about once per dropped folder. Ebooks are filed Author/Title/; documents keep the shape they were dropped in. See the library.

The optional OPDS source

internal/source/opds and internal/provision/calibreweb.go remain as an opt-in for a Calibre server running somewhere else, via SOUNDSTORM_CALIBREWEB_URL. SoundStorm no longer runs that server, so it cannot vouch for it: a log line used to call its published default password safe "because it is unreachable except through SoundStorm", which stopped being true when the bundled container went, and now says the password is safe only if the operator's own server is. OPDS ids are the client's, fetched with the server's credential, so an id must stay under the server's opds/ tree with no query of its own. A remote catalog cannot serve a book's insides, so the in-browser reader is for local books only.