Provisioning with nobody logging in
A fresh install starts six media servers, every one of which normally wants a
human to click through a setup wizard. Nobody does. internal/provision walks each
backend's first run through its own API, keeps the credentials it made, and reconnects with them
after every restart - without ever mistaking "not up yet" for "wrong password".
A backend is two halves
The third rule the design rests on: a backend is search and provisioning, both, or it is not done. A backend a human must configure by hand defeats the point of the product - and would mean showing them its admin screens, which is exactly the seam SoundStorm hides. That rule is also why Plex was rejected outright: it requires a plex.tv account, so it cannot be set up without a human logging into a cloud service.
Compose decides the stack: a backend is managed if its URL is set
(SOUNDSTORM_NAVIDROME_URL, SOUNDSTORM_JELLYFIN_URL and so on), and
targetsFromEnv turns those into provisioning targets. Each gets its own goroutine in
Manager.Start, so the server opens its port immediately and shows "Films and TV:
Getting ready" while Jellyfin boots. Sources join the registry as each becomes ready.
Each backend's first run
Every endpoint below was checked against a running server rather than taken from documentation, and several only work while the backend has no users - which makes driving them safe: they cannot be used to hijack an install that is already set up.
Navidrome (music)
A /ping first, so the retry loop reports "waiting for navidrome" rather than a
confusing failure, then POST /auth/createAdmin with the account name
soundstorm and a generated password. That endpoint works only while no user exists.
The music folder needs no call: compose sets ND_MUSICFOLDER, and the same compose
file sets ND_SCANINTERVAL (not ND_SCANSCHEDULE, which Navidrome ignores
silently), real file paths (ND_SUBSONIC_DEFAULTREPORTREALPATH) and a transcoding
setting the beat analysis needs. Navidrome's Subsonic API takes credentials in each request, so
the username and password are what is stored.
Jellyfin (films and TV)
Jellyfin's startup wizard is a plain REST API that stops accepting calls once setup completes:
/Startup/Configuration, /Startup/User, /Startup/RemoteAccess,
/Startup/Complete. Then /Users/AuthenticateByName for a token, and
/Library/VirtualFolders to create two libraries - films and series are separate
Jellyfin libraries with different collection types and scrapers. So one Jellyfin is provisioned
once and registered as two sources sharing one token, jellyfin
(Movie) and jellyfin-tv (Series, Episode); buildSources returns a slice
for exactly this reason. Jellyfin accepts its listening socket while still loading and answers 503
"server is loading" for several seconds, which matters below.
Audiobookshelf
/status says whether it is initialized; POST /init creates the root
account; /login returns two tokens. user.accessToken expires in an hour;
user.token is a legacy JWT with no expiry, and SoundStorm stores that one on purpose -
the other would strand the backend an hour after provisioning without a refresh flow. Then a
library on the audiobooks folder, and a scan. Members get their own Audiobookshelf accounts later,
on first use (TokenFor); creating one needs isActive sent explicitly, or
the account is created inactive and its token, returned in an ordinary 200, answers Unauthorized
to everything.
Immich (photos)
/api/auth/admin-sign-up with soundstorm@soundstorm.invalid (RFC 2606's
never-existing domain - Immich validates the shape and sends nothing to it), a login, then an API
key with permissions: ["all"]. The session token expires and the key does not, so
only the key is kept and the password is thrown away. Then an external library
on /pictures, mounted read-only, so Immich can never move, rename or delete a photo.
Folder watching is switched on by reading /api/system-config, changing
library.watch.enabled and sending the whole document back - a test holds that it
changes nothing else. Members get an Immich account each whose one library is their own folder
(PhotoAccountFor), made the first time they look.
Storyteller (read-along)
The first account is a Next.js server action on /init, which works as a plain
multipart form; the action id changes per build, so it is scraped off the page each time.
/api/v2/token gives a thirty-day session, and /api/v2/token/app trades it
for a 35-year one - and reads only the st_token cookie, not a bearer header, which
cost a test cycle. Settings then keep synced books inside Storyteller's own storage (the default
writes beside the source, into the library), choose the transcription model, and set the maximum
track length huge so chapters are the only cuts. Books are later uploaded to it over tus.
AudioMuse-AI (moods)
Its /api/setup answers without a login until the first save and never after - a
second unauthenticated save is a 401 - the same shape as Jellyfin's wizard. SoundStorm generates
its API token and admin password, switches off its third-party lyrics lookup and lyrics
transcription, and gives it a non-admin Navidrome account of its own
(audiomuse, made through Navidrome's /api/user), since it reads songs
through Navidrome and mounts no folder. A nightly analysis is scheduled through
/api/cron. If AudioMuse's volumes were wiped and Navidrome's were not, that account
is reset rather than refused.
Ebooks and documents
No backend: SoundStorm reads those folders itself (source/localbooks), so
"provisioning" is the first scan, run synchronously, with a ticker for the rest.
What is kept
Credentials go into state.json beside the accounts: base URL, username and
password or token, library ids, and when it was provisioned. Passwords are generated with
crypto/rand, 192 bits. Nothing about the media itself is ever stored there - that is
the line that keeps the state from becoming an index that goes stale.
A setup that fails half way
A first run is several steps - make the admin account, sign in, make a key, make a library -
and a failure can come after the first: a timeout, the backend restarting, AudioMuse taking a
minute to come back after saving its settings. The credentials used to be saved only once every
step had worked, so such a failure left an account whose password nobody held, and every retry
was refused as "already set up". Now every password, key and token a setup makes comes from
Manager.secretsFor: written to state.json (SetupSecrets)
before the backend is told it, and the same one on every attempt. A retry that finds its own
account already made signs in with the kept secret and carries on; SetBackend clears
what was kept once the backend is saved. A backend that is already set up with nothing kept
is still refused - carrying on is only for SoundStorm's own unfinished setup. Each person's own
Audiobookshelf and Immich account works the same way, under its own name.
soundstorm backup and restore are that decision's tools, and the
uninstaller takes a backup before down -v, the one moment the credentials would stop
existing anywhere.Reconnecting is not provisioning
On a restart, SoundStorm is listening seconds after its container starts and Jellyfin is not. A health check then fails with "connection refused" - which is emphatically not the same as "this token is wrong". Treating them alike meant throwing away good credentials and falling through to provisioning, which can never succeed on a backend that is already set up. One unlucky restart, and a backend was permanently broken with its working token still on disk.
So run checks for stored credentials first and calls reconnect, which
retries until the backend answers. Only a definite answer from the backend - in practice a 401 or
403 refusing the credentials - sends it to provision (a backend whose volume really was reset can be set up again).
transient() owns the distinction, at two layers:
- Transport: connection refused, DNS failure, timeout, any
net.OpError- nothing is listening yet. - HTTP: something is listening but not ready - 5xx (Jellyfin's 503 while loading) and 429.
provision_test.go guards it.
Backoff, and giving up
Both loops start at a two-second wait and add two seconds per attempt up to about fifteen,
showing "waiting for jellyfin (attempt 6)" to the owner meanwhile - the setup box shows shelves,
never server names, to everybody else. After ten minutes (giveUpAfter) a backend is
marked failed: one can be genuinely absent - an image that failed to pull, a wrong URL - and
SoundStorm must stay useful for the ones that worked. Calls made while somebody waits, like
creating a member's account, get twenty seconds; a search gets five.
Recovering panics
provision.run is started once with go and answers to no request, so
net/http's per-request panic recovery does not cover it. It parses responses from
software outside the process - and a localbooks scan parses EPUBs a member uploaded before the
backend came up. An unrecovered panic would take the whole server down. So run is
wrapped in recover(): a panic marks that one backend failed, with the panic in the
log, and SoundStorm keeps serving the others. It was verified with a real panic (a nil state store
inside provisioning), not a synthetic one.
Telling the backends to look
Every backend indexes on its own timer - Navidrome every minute, the ebook scanner every two,
Jellyfin and Audiobookshelf when their watchers notice - so an uploaded file could sit on disk,
unsearchable, for up to two minutes after the progress bar finished. source.Rescanner
is the optional "look now" call, and an upload schedules one:
| Backend | Call |
|---|---|
| Navidrome | /rest/startScan.view - the quick kind, looking only at what changed |
| Jellyfin | POST /Library/Refresh (204; a made-up path is 404, so the 204 means something). It refreshes every library. |
| Audiobookshelf | POST /api/libraries/{id}/scan |
| Immich | a scan of every library, members' included |
| Ebooks, documents | the scan the ticker would have run |
The debounce is the load-bearing part. A dropped folder arrives as one upload
per file, and triggering per file would ask Navidrome to scan thirty times for one album.
scheduleRescan keeps one timer per kind and pushes it back on each upload, so a long
upload produces one scan, two seconds after the last file. Measured after: an uploaded ebook was
searchable in five seconds.
The manual "Check for new files" button is open to any signed-in person, and each real scan is
expensive. The debounce does nothing against separate triggers spaced further apart, so
minRescanInterval (30s) floors the gap between two real scans of one kind - and a
trigger arriving while a floor-extended timer is pending recomputes the same floor rather than
resetting to two seconds, or a steady stream could keep a scan two seconds away for ever.
Before any scan, library.EnsurePlaceholders puts back each shelf's
README.txt. Jellyfin refuses to remove anything when a library folder comes back empty
- it cannot tell "everything was deleted" from "the drive did not mount" - so a folder emptied by
hand left every deleted film searchable for ever. The placeholder keeps the folder non-empty. A
failed scan is logged and dropped: the upload already succeeded, and the backend's own timer will
find the file regardless.