Books and read-along

Books are where SoundStorm does the most itself. Ebooks and documents have no backend at all - the server reads the folders - while audiobooks belong to Audiobookshelf. Read Along joins the two: an ebook whose page turns with its audiobook, the timing worked out by Storyteller.

The media types SoundStorm owns

Calibre-Web was the original ebook backend and was removed: it needed a database rather than a folder, its setup had no API (it meant scraping forms past CSRF tokens), and its password could not be rotated. SoundStorm reads the folder itself instead.

The justification is narrow and should stay narrow: an EPUB is self-describing. The file holds its own title, author, language and cover in documented XML, and a book needs no transcoding. Dune.2021.mkv holds none of that, which is precisely the work Jellyfin exists to do. So the rule is: SoundStorm can own a media type when it is self-describing and needs no transcoding. EPUB qualifies; video never will; photos fail it on HEIC.

Documents are a second shelf served the same way (media.KindDocument, documents/): a PDF is a book or a gas bill, and the ebook shelf used to hold both. Documents get everything a kind gets - their own chip, access checkbox and rescan - because a half-kind sharing ebooks' permission would be a shelf you could see into from the wrong account. Both shelves are instances of one source type, localbooks; see the ebooks and documents page for the scanner and parsers.

PDFs without a PDF parser

internal/pdf deliberately does not parse PDF structure. Doing it properly means the cross-reference table, then cross-reference streams, then object streams - hundreds of lines before the first title. Measured on five real PDFs from five producers, all that machinery would have returned nothing: every one had an Info dictionary that was absent or literally /Title (). Two carried an XMP packet - plain uncompressed XML - which a byte scan and encoding/xml reach for free. So the order is XMP, then the Info dictionary if readable, then the file name, which for most PDFs is the only real information anybody wrote down.

Calibre libraries

Existing Calibre libraries work with no SQLite driver: Calibre writes a metadata.opf beside every book in exactly the format an EPUB carries inside, so one parser reads both - including series and series index, and a curated cover.jpg. That is also why ebooks are filed Author/Title/: it is Calibre's shape, and flattening it would part books from their sidecars.

Reading position

Reflowable text means a page number is meaningless - change the font size and "page 47" is different words - so position is a content-anchored locator, an EPUB CFI, plus how far through the book. It is kept in SoundStorm's own state, keyed userID/sourceID/itemID, so it follows a person across devices (the Apple TV's native reader reads and writes the same CFI). Because state.json is rewritten whole by every login, saving a position is guarded hard:

PDFs open in the browser's own viewer, which exposes no position, so the web app keeps none for them.

The reader, server side

Rendering is foliate-js, vendored - the one piece of third-party code in the project, pinned by content and embedded. What is SoundStorm's is that the server unzips: the reader fetches each chapter, stylesheet and image from /api/book/resource, so no zip library runs in the browser.

A book is somebody else's file, opened in the owner's signed-in session, and three review passes each found a way for a crafted EPUB to run script. The guards that resulted:

Audiobooks

Listening position belongs to Audiobookshelf, per person, written with its own PATCH /api/me/progress/{id}. A private copy would quietly fork from the one every other client reads; this way a chapter finished in Audiobookshelf's phone app is where the browser picks up. That requires each member to have their own Audiobookshelf account, which provision.TokenFor creates lazily on first use (the owner uses the administrator's). Details of the API's sharp edges - progress not derived from currentTime, isFinished: false resetting a book to zero - are on the Audiobookshelf page.

A book is its files on one clock: source.TrackLister reports each file with its StartSeconds, so "two hours in" converts to a file and an offset. source.ChapterLister sends Audiobookshelf's own chapter list, which names the marks inside a single m4b as well as one file per chapter - so a single-file book has navigation too. Without a chapter list, the files stand in. Books in progress appear in the Continue row through source.InProgressLister.

Read Along: finding pairs

GET /api/books/pairs lists what somebody has as both an ebook and an audiobook. It lists both shelves through the registry, so access applies, and matches on a key with edition noise removed - anything bracketed (ASIN, Unabridged), a subtitle after a colon, a trailing "Book 1", curly quotes, punctuation, a leading article - plus a shared author surname, reading "Rowling, J.K." and "J.K. Rowling" alike. Nothing is stored. On a real library, 1,663 ebooks and 113 audiobooks gave four pairs, and a scan for near-misses found none.

The owner can correct it either way: Not the same book (kept as notPairs, and whatever Storyteller made of it deleted) and Pair with its audiobook (kept as manualPairs). These are decisions in state.json, like the starter-library flag - not facts about the media.

Read Along: syncing with Storyteller

Matching a recording to its text sentence by sentence is forced alignment - transcribe the audio, find it in the book - the ML-shaped layer this project does not own. So it is Storyteller. A cheaper chapter-proportional guess was offered and declined: a page off by one reads as broken.

SoundStorm keeps playing the audiobook through its own player. Storyteller's output is an EPUB 3 with media overlays and its own copy of the audio; playing that would lose the lock screen, the mini-player and Audiobookshelf's listening position. So only its text is used - every sentence wrapped in a span with an id - plus its SMIL timings, converted to the audiobook's whole-book clock (storyteller.Timeline).

The conversion rests on how Storyteller cuts audio, read from its source and checked end to end: files numbered in name order, each cut at its chapter marks (the maximum track length is set huge at provisioning so chapters are the only cuts). Piece C of file F therefore starts at that file's C-th chapter, which is Audiobookshelf's chapter list (source.AudioLayouter). Where pieces and chapters do not line up there is no timeline, and the book opens saying it cannot follow rather than following the wrong sentence.

Books sync by themselves: a look two minutes after start, every half hour, and three minutes after a book scan. A pair that fails is not retried automatically - a failure retried every half hour would be the better part of an hour of CPU, repeated - but a failure never unmatches either, since it is usually Storyteller, not the match. In the reader, four times a second the sentence at the player's time is found, turned to and lit; turning the page by hand pauses the following for twelve seconds. Measured: a 10-hour book takes about 40 minutes to sync.

POST /api/readalong accepts only a pair Read Along lists, and a Storyteller book needs audiobook access as well as ebook access to read, since it holds the audiobook's audio.

Downloads

Downloading is the web app's (the Cache API, with an index in local storage), but the server shapes it. Audiobooks, ebooks, documents and Read Along pairs download - both halves, plus a synced book's text without its copy of the audio, and its timeline. Everything is keyed by the address the app would ask the server for, so one index and one Remove serve every kind. Listening and reading positions are kept on the device too, marked unsynced until a save reaches the server, and an unsynced one wins on the next open. The reader reads a downloaded book from the cache before asking the server at all.