Provisioning with nobody logging in

A fresh install starts six media servers, every one of which normally wants a human to click through a setup wizard. Nobody does. internal/provision walks each backend's first run through its own API, keeps the credentials it made, and reconnects with them after every restart - without ever mistaking "not up yet" for "wrong password".

A backend is two halves

The third rule the design rests on: a backend is search and provisioning, both, or it is not done. A backend a human must configure by hand defeats the point of the product - and would mean showing them its admin screens, which is exactly the seam SoundStorm hides. That rule is also why Plex was rejected outright: it requires a plex.tv account, so it cannot be set up without a human logging into a cloud service.

Compose decides the stack: a backend is managed if its URL is set (SOUNDSTORM_NAVIDROME_URL, SOUNDSTORM_JELLYFIN_URL and so on), and targetsFromEnv turns those into provisioning targets. Each gets its own goroutine in Manager.Start, so the server opens its port immediately and shows "Films and TV: Getting ready" while Jellyfin boots. Sources join the registry as each becomes ready.

Each backend's first run

Every endpoint below was checked against a running server rather than taken from documentation, and several only work while the backend has no users - which makes driving them safe: they cannot be used to hijack an install that is already set up.

Navidrome (music)

A /ping first, so the retry loop reports "waiting for navidrome" rather than a confusing failure, then POST /auth/createAdmin with the account name soundstorm and a generated password. That endpoint works only while no user exists. The music folder needs no call: compose sets ND_MUSICFOLDER, and the same compose file sets ND_SCANINTERVAL (not ND_SCANSCHEDULE, which Navidrome ignores silently), real file paths (ND_SUBSONIC_DEFAULTREPORTREALPATH) and a transcoding setting the beat analysis needs. Navidrome's Subsonic API takes credentials in each request, so the username and password are what is stored.

Jellyfin (films and TV)

Jellyfin's startup wizard is a plain REST API that stops accepting calls once setup completes: /Startup/Configuration, /Startup/User, /Startup/RemoteAccess, /Startup/Complete. Then /Users/AuthenticateByName for a token, and /Library/VirtualFolders to create two libraries - films and series are separate Jellyfin libraries with different collection types and scrapers. So one Jellyfin is provisioned once and registered as two sources sharing one token, jellyfin (Movie) and jellyfin-tv (Series, Episode); buildSources returns a slice for exactly this reason. Jellyfin accepts its listening socket while still loading and answers 503 "server is loading" for several seconds, which matters below.

Audiobookshelf

/status says whether it is initialized; POST /init creates the root account; /login returns two tokens. user.accessToken expires in an hour; user.token is a legacy JWT with no expiry, and SoundStorm stores that one on purpose - the other would strand the backend an hour after provisioning without a refresh flow. Then a library on the audiobooks folder, and a scan. Members get their own Audiobookshelf accounts later, on first use (TokenFor); creating one needs isActive sent explicitly, or the account is created inactive and its token, returned in an ordinary 200, answers Unauthorized to everything.

Immich (photos)

/api/auth/admin-sign-up with soundstorm@soundstorm.invalid (RFC 2606's never-existing domain - Immich validates the shape and sends nothing to it), a login, then an API key with permissions: ["all"]. The session token expires and the key does not, so only the key is kept and the password is thrown away. Then an external library on /pictures, mounted read-only, so Immich can never move, rename or delete a photo. Folder watching is switched on by reading /api/system-config, changing library.watch.enabled and sending the whole document back - a test holds that it changes nothing else. Members get an Immich account each whose one library is their own folder (PhotoAccountFor), made the first time they look.

Storyteller (read-along)

The first account is a Next.js server action on /init, which works as a plain multipart form; the action id changes per build, so it is scraped off the page each time. /api/v2/token gives a thirty-day session, and /api/v2/token/app trades it for a 35-year one - and reads only the st_token cookie, not a bearer header, which cost a test cycle. Settings then keep synced books inside Storyteller's own storage (the default writes beside the source, into the library), choose the transcription model, and set the maximum track length huge so chapters are the only cuts. Books are later uploaded to it over tus.

AudioMuse-AI (moods)

Its /api/setup answers without a login until the first save and never after - a second unauthenticated save is a 401 - the same shape as Jellyfin's wizard. SoundStorm generates its API token and admin password, switches off its third-party lyrics lookup and lyrics transcription, and gives it a non-admin Navidrome account of its own (audiomuse, made through Navidrome's /api/user), since it reads songs through Navidrome and mounts no folder. A nightly analysis is scheduled through /api/cron. If AudioMuse's volumes were wiped and Navidrome's were not, that account is reset rather than refused.

Ebooks and documents

No backend: SoundStorm reads those folders itself (source/localbooks), so "provisioning" is the first scan, run synchronously, with a ticker for the rest.

What is kept

Credentials go into state.json beside the accounts: base URL, username and password or token, library ids, and when it was provisioned. Passwords are generated with crypto/rand, 192 bits. Nothing about the media itself is ever stored there - that is the line that keeps the state from becoming an index that goes stale.

A setup that fails half way

A first run is several steps - make the admin account, sign in, make a key, make a library - and a failure can come after the first: a timeout, the backend restarting, AudioMuse taking a minute to come back after saving its settings. The credentials used to be saved only once every step had worked, so such a failure left an account whose password nobody held, and every retry was refused as "already set up". Now every password, key and token a setup makes comes from Manager.secretsFor: written to state.json (SetupSecrets) before the backend is told it, and the same one on every attempt. A retry that finds its own account already made signs in with the kept secret and carries on; SetBackend clears what was kept once the backend is saved. A backend that is already set up with nothing kept is still refused - carrying on is only for SoundStorm's own unfinished setup. Each person's own Audiobookshelf and Immich account works the same way, under its own name.

Provisioning is not idempotent across a volume reset. If a backend's volume is wiped but SoundStorm's state survives, or the other way round, you get a backend with an account whose password nobody holds. The provisioners detect it and say so precisely - "navidrome already has an admin account but SoundStorm has no stored credentials for it; either restore SoundStorm's state file or reset the navidrome volume" - rather than retrying for ever. The fix is a human decision. soundstorm backup and restore are that decision's tools, and the uninstaller takes a backup before down -v, the one moment the credentials would stop existing anywhere.

Reconnecting is not provisioning

On a restart, SoundStorm is listening seconds after its container starts and Jellyfin is not. A health check then fails with "connection refused" - which is emphatically not the same as "this token is wrong". Treating them alike meant throwing away good credentials and falling through to provisioning, which can never succeed on a backend that is already set up. One unlucky restart, and a backend was permanently broken with its working token still on disk.

So run checks for stored credentials first and calls reconnect, which retries until the backend answers. Only a definite answer from the backend - in practice a 401 or 403 refusing the credentials - sends it to provision (a backend whose volume really was reset can be set up again). transient() owns the distinction, at two layers:

provision_test.go guards it.

Backoff, and giving up

Both loops start at a two-second wait and add two seconds per attempt up to about fifteen, showing "waiting for jellyfin (attempt 6)" to the owner meanwhile - the setup box shows shelves, never server names, to everybody else. After ten minutes (giveUpAfter) a backend is marked failed: one can be genuinely absent - an image that failed to pull, a wrong URL - and SoundStorm must stay useful for the ones that worked. Calls made while somebody waits, like creating a member's account, get twenty seconds; a search gets five.

Recovering panics

provision.run is started once with go and answers to no request, so net/http's per-request panic recovery does not cover it. It parses responses from software outside the process - and a localbooks scan parses EPUBs a member uploaded before the backend came up. An unrecovered panic would take the whole server down. So run is wrapped in recover(): a panic marks that one backend failed, with the panic in the log, and SoundStorm keeps serving the others. It was verified with a real panic (a nil state store inside provisioning), not a synthetic one.

Telling the backends to look

Every backend indexes on its own timer - Navidrome every minute, the ebook scanner every two, Jellyfin and Audiobookshelf when their watchers notice - so an uploaded file could sit on disk, unsearchable, for up to two minutes after the progress bar finished. source.Rescanner is the optional "look now" call, and an upload schedules one:

BackendCall
Navidrome/rest/startScan.view - the quick kind, looking only at what changed
JellyfinPOST /Library/Refresh (204; a made-up path is 404, so the 204 means something). It refreshes every library.
AudiobookshelfPOST /api/libraries/{id}/scan
Immicha scan of every library, members' included
Ebooks, documentsthe scan the ticker would have run

The debounce is the load-bearing part. A dropped folder arrives as one upload per file, and triggering per file would ask Navidrome to scan thirty times for one album. scheduleRescan keeps one timer per kind and pushes it back on each upload, so a long upload produces one scan, two seconds after the last file. Measured after: an uploaded ebook was searchable in five seconds.

The manual "Check for new files" button is open to any signed-in person, and each real scan is expensive. The debounce does nothing against separate triggers spaced further apart, so minRescanInterval (30s) floors the gap between two real scans of one kind - and a trigger arriving while a floor-extended timer is pending recomputes the same floor rather than resetting to two seconds, or a steady stream could keep a scan two seconds away for ever.

Before any scan, library.EnsurePlaceholders puts back each shelf's README.txt. Jellyfin refuses to remove anything when a library folder comes back empty - it cannot tell "everything was deleted" from "the drive did not mount" - so a folder emptied by hand left every deleted film searchable for ever. The placeholder keeps the folder non-empty. A failed scan is logged and dropped: the upload already succeeded, and the backend's own timer will find the file regardless.