Storyteller (read-along)

Storyteller transcribes an audiobook and finds every sentence of it in the matching ebook. SoundStorm hands it pairs from Read Along, takes back the sentence timings, and uses them to turn the reader's page as the audiobook plays - through SoundStorm's own player, not Storyteller's.

Why a backend

Pages that turn by themselves need to know, sentence by sentence, where the narrator is. That is forced alignment - transcribe the audio, find it in the text - which is machine learning over audio, the expensive layer SoundStorm does not own. A cheap alternative, guessing in proportion through each chapter, was offered and declined: a page off by one reads as broken. Storyteller is MIT-licensed and runs in Docker. Its image is 2.9 GB, idle RAM about 300 MB and 0.7 GB while syncing.

Pinned by digest

Its API moves between releases, so it is pinned by digest (to web-v2.14.21), and a new version is a deliberate change.

Provisioning

All checked live:

The audiobook folder is mounted read-only, so Storyteller can never change a recording.

Two steps to add a pair

The one-call form is broken: naming the EPUB and the audio together in POST /api/v2/books fails when they are in different folders - each tries to create the book and the second hits a duplicate id. So the recording is added by reference (never copied), and the ebook is uploaded into that book over tus (/api/v2/books/upload). That also keeps the shelf's file untouched when an EPUB 2 - most of a Calibre library - has to be upgraded: the upgrade happens to Storyteller's copy. A synced book is found again by its audiobook's folder, so SoundStorm stores nothing about it.

What is used, and the timing conversion

Storyteller's output is an EPUB 3 with media overlays and its own copy of the audio inside. Playing that would lose the lock screen, the mini-player and Audiobookshelf's listening position. So only its text is used - every sentence wrapped in a span with an id - and its SMIL timings, converted by storyteller.Timeline to the audiobook's whole-book clock.

The conversion rests on how Storyteller cuts audio, read from its source and checked end to end: input files are numbered in name order (Audio/0000F-0000C), and each is cut at its chapter marks. Piece C of file F therefore starts at that file's C-th chapter - which is Audiobookshelf's chapter list, exposed as source.AudioLayouter. Where the pieces and chapters do not line up there is no timeline, and the book opens saying it cannot follow rather than following the wrong sentence.

Syncing

Pairs sync by themselves: a look two minutes after start, every half hour, and three minutes after a book scan. One sync runs at a time. Measured on the development machine: 8 minutes of speech in 30 s with base.en, so a 10-hour book takes about 40 minutes; disk use is roughly the audiobook's size again, since the synced EPUB embeds the audio. A failed pair is not retried automatically - a failure retried every half hour would be the better part of an hour of CPU, repeated - and is never unmatched, because failure is usually Storyteller, not the match. "Not the same book" deletes Storyteller's copies only; its delete touches nothing outside its own storage.

Quirks and open items

Verified end to end on an isolated stack with synthetic speech: an EPUB 3 with two MP3s and an EPUB 2 with a chaptered M4B both synced through SoundStorm's API, each timeline ended exactly where its recording did, and in Chrome the lit sentence moved with the narration and the page turned on its own. See Books and read-along.