Accounts and sign-in

One login is half of what SoundStorm promises, so the server keeps its own accounts - two roles, a password hash, sessions - and makes the backends' accounts itself. internal/auth owns signing in; internal/state keeps what outlives a restart.

Two roles, deliberately thin

The owner is whoever installed the server; they can add and remove people, choose which shelves each may see, change server settings and delete media. Everybody else is a member. That is the entire difference. A media server for a household does not need a permission matrix, and every role beyond these two is a decision somebody has to make about their family.

The owner is always unrestricted, whatever is stored for them. They are the only account that can change restrictions, so an owner who locked themselves out of a shelf would have no way back in.

There are no invite links and no open registration. Sign-up is a first-boot action that closes the moment an account exists - a second sign-up would be a stranger who found the port claiming somebody else's server. After that, only the owner adds accounts. Even the role of the first account is decided by the server: AddUser makes the first account the owner and everyone after a member under one lock, so two sign-ups racing on a fresh install cannot both become owner.

The setup code

Until the first account exists, whoever reaches the port owns the server. "Only from the home network" was the first answer and cannot be checked: under Docker Desktop every connection - localhost, the LAN address, anything a router forwards from the internet - arrives from Docker's own bridge address. That was measured with a failed sign-in from each, not assumed.

So the first sign-up needs a setup code, eighty random bits. It costs the person installing nothing: the installers generate it into .env (SOUNDSTORM_SETUP_CODE) and open the browser at /?setup=<code>, and the page fills it in. The installer also shows it in full at the end, grouped in fours; case, spaces and dashes are ignored when typed (NormalizeSetupCode). With no installer, SoundStorm makes one up at each start and logs it until an account exists.

The page removes ?setup= from the address once it knows it is staying, so the code is not left in history or a bookmark. That once ran too early: the page stripped the code, then moved itself to the install's secure name carrying an already-empty query, and the first screen asked for a code the person had never seen. forgetSetupCodeInAddress now runs only on the paths that stay.

Passwords

Passwords are hashed with PBKDF2-HMAC-SHA256 at 600,000 iterations (OWASP's recommendation), a 16-byte salt and a 32-byte key, from the standard library's crypto/pbkdf2. It costs a few hundred milliseconds per sign-in, which is the point.

Sessions and cookies

A session is a 256-bit random token in a cookie, valid for thirty days. The state file keys sessions by a SHA-256 hash of the token, never the token itself. That file is what a backup copies and what an old volume leaves behind with nobody running the server around it; with raw tokens as keys it was itself enough to sign in as anybody with a live session. A slow hash is unnecessary - the input is 256 random bits, not a password. The migration was free: since the old key was the raw token, hashing it in place computes exactly what a real cookie looks up next, and nobody was signed out.

Over TLS the cookie is called __Host-soundstorm_session. Every install's real name is under one shared domain not yet on the Public Suffix List, so any other install - one an attacker set up included - is the same site and could plant a cookie for the whole domain, sent ahead of the real one, signing the owner out over and over. A browser refuses a __Host- cookie that names a Domain, so no other host can plant one. Over plain HTTP on a LAN address the cookie keeps the plain name; a domain-scoped cookie never reaches a bare IP anyway. Every session cookie a request carries is tried, not only the first.

Each account keeps at most fifty sessions; the oldest go first. SOUNDSTORM_TRUST_PROXY lets X-Forwarded-Proto decide whether the cookie is Secure behind a TLS-terminating proxy - off unless asked for, because any client can send that header.

The sign-in throttle

Signing in is the one unauthenticated endpoint that costs real CPU: 600,000 rounds is enough to pin every core of a Raspberry Pi with a handful of parallel requests. internal/auth/throttle.go answers that, and guessing, separately.

LimitRuleWhy
Hashing slotsAt most two hashes at once, server-wide; a sign-in waits up to 10s for a slotProtects the CPU whoever is asking.
Per address5 free wrong passwords, then 1s, doubling, capped at 5 minutes; forgotten after an hour without failuresSlows a single guesser.
Per account10 free, then 2s, doubling, capped at 1 minute; counted for any name, existing or notPer address is five guesses per address an attacker can borrow. The cap is short because it reaches the real owner too.
One guess in flight per accountthrottle.reserveThe wait was checked before hashing and a failure recorded after, so a burst sent at once all passed the check. Now a burst against one name waits on itself.

Three details make it honest. A throttled request is refused before hashing, and refused even with the right password - otherwise the 429 would fire only for wrong answers and a guesser would learn something for free. Names are counted whether or not the account exists, so the 429 does not reveal who has one, and a spray of made-up names is pruned first so it cannot flush the real one out of the table. And it is capped backoff, never a lockout: behind Docker Desktop everybody can share one address, and a household locked out of its own server is far worse than a script slowing it down for a few minutes. IPv6 clients are keyed by their /64.

Device tokens

Capped backoff still let a stranger guessing at a known name once a minute hold that account's sign-ins in backoff indefinitely, the owner's included, and let one stranger's wrong guesses on Docker Desktop's single shared address slow the whole household. So every successful sign-in, sign-up and password change leaves a device cookie (soundstorm_device, __Host- over TLS) holding id.issued.nonce.HMAC, keyed by the state's device key and covering the account's current salt. A sign-in carrying a valid one for the account being signed in to skips the address's and account's backoff and is judged on that device's own record alone - the same five free failures and doubling wait, one guess in flight at a time (guardedTrusted).

A stolen device token is no faster a guessing channel than one address, and the stranger, who has never signed in, has no token at all. Any password change or reset changes the salt and retires every token issued before it. What remains exposed is only a brand-new device signing in while an attack is actually running. Sign-out keeps the device cookie: the device is still one the account uses.

The end-to-end test runs every client from one loopback address - exactly the Docker Desktop case - and fails with the exemption disabled.

Asking for new passwords

The owner can ask everybody, themselves included, to choose a new password (POST /api/users/new-passwords). Each account is marked (User.MustChangePassword), and from then on every guarded route answers that account with a 403 saying so - only /api/account/password works - so it is more than a screen; the page shows one, asking for the current password and a new one under the rules. Choosing one clears the mark and signs out the person's other devices. The same mark is set on signing in with a password that would be refused if chosen today: sign-in is the one moment the server sees it, so passwords set before the rules tightened are replaced as people next sign in.

New devices need approval

An owner setting, off unless turned on (People, New devices). With it on, the right password is not enough on a device an account has never signed in on: the sign-in is held (devices.go) and answers 202 with a random id the waiting device polls, while every device already signed in to that account - and the owner's - is asked "Allow a new device?", naming the browser. No session exists while it waits (one made at once would count towards the account's fifty and push real ones out). Allowed, a session is made and handed to the waiting device with a device token, so it is not asked again; refused, or ten minutes gone, the request is forgotten. A guessed or leaked password then opens nothing. "New" is the device token, not where the request came from: behind Docker Desktop every connection looks alike, and a request's account of its own address can be faked. A password change retires every device token, so after one every other device asks again. Somebody with nothing else signed in - the owner on a new laptop - can approve with the setup code from the server's .env, since whoever can read the server's files owns it anyway; a code made up at start (no .env) cannot.

Which shelves somebody can see

User.Libraries is a list of media kinds; nil means all of them, which is what every account made before the field existed has, and the default for a new one. The granularity is a whole media kind, which is the same thing as a source, because each source serves exactly one kind. "No films for the seven-year-old" is answerable; "only these films" is not, and would mean per-person Jellyfin accounts and parental ratings - not worth a second provisioning path without deciding so.

auth.Access(user) turns the account into a source.Access, the middleware puts it in the request context once for every guarded route, and the registry refuses anything outside it. A few kinds need more than one permission: a Storyteller book is an ebook holding an audiobook's audio, so reading one needs audiobook access too (AlsoNeeds).

An empty list must never read back as nil. User.Libraries has no omitempty for exactly this reason: with it, an account allowed no libraries would serialize to nothing, read back as nil, and silently mean every library. A test writes it, reads it back and checks. Similarly, {} sent to the owner's libraries endpoint once made a member unrestricted, because a *[]string cannot tell an absent key from null; it is now refused.

AccessFrom defaults to unrestricted, because provisioning, health checks and the library counter run with no account and must see everything. That is a fail-open default, which is why internal/httpapi/libraries_test.go walks every endpoint that hands over bytes, metadata or a playable address - /api/stream included, since hiding search results is no permission when the address is guessable.

The few things that are per person

Search returns the same results whoever asks, and a film is the same bytes. What differs: reading and video positions (kept by SoundStorm, keyed userID/sourceID/itemID), favorites, playlists and history (in collections), audiobook listening positions, and photos.

Typing a password with a remote is the worst part of a TV app, so a TV can be signed in from a phone instead (tvlink.go). The TV asks for a code (POST /api/link, limited by address) and gets three things: a random 128-bit id that only it knows, a short code of six characters from an alphabet without 0, O, 1, I or L (shown ABC-DEF), and the address a phone should open, /?link=ABCDEF on the install's secure name when it has one. It shows the code and that address as a QR code (GET /api/link/{id}/qr.png, drawn by internal/qr - byte mode, level M, versions 1 to 10, no dependency) and polls GET /api/link/{id} every three seconds.

A phone signed in opens the address (or has the code typed into Settings), looks the code up (GET /api/link/code/{code}, any case, dash or not) and is shown what is asking and that it will be signed in as the phone's person; allowing it (POST /api/link/code/{code}) marks the code, and the TV's next poll is handed a fresh session for that person (auth.SessionFor) and a device cookie. A code lives ten minutes and works once; looking codes up is limited per person.

The short code is what a person types, so it is not what the TV polls with: the id is, as a waiting sign-in's is. Guessing codes from a phone only ever signs a stranger's TV in as yourself. What no limit can stop is somebody being talked into allowing a stranger's code, which is why the phone says plainly what is asking and as whom. A TV signed in this way skips approval of new devices: the phone that allowed it is a device the account already uses.

Who's listening? Several people on one device

A living-room TV or a family tablet is used by several people, and signing out and in again with a password each time is the wrong shape for it. So a device can keep people (profiles.go, auth/profiles.go), after the account-then-people pattern of the streaming services. The device gets a profile cookie of its own (soundstorm_profiles, __Host- over TLS): a random id, not a sign-in. Who may be switched to on it is the server's record (state.Kept, keyed by a hash of that id), each person with a fingerprint of their password's salt, so a new password takes them off every device.

Per-member backend accounts

Two backends get an account per member, because what they remember is personal:

Navidrome and Jellyfin keep one shared account on purpose: favorites and watch positions are SoundStorm's own, so nothing surfaced from them differs per person. source.WithUserID carries the account id to adapters, and lives in internal/source rather than internal/auth so an adapter can learn who is asking without depending on how signing in works.

Changing and resetting passwords

Removing somebody

Removing a person deletes their sessions, positions, collections file, custom covers, listen log and backend accounts. The backend accounts go first, because deleting the SoundStorm account drops the record of which Audiobookshelf or Immich user belonged to it, and after that nothing knows what to clean up. It is best effort: a backend that is down must not stop somebody being removed. Their photo folder is kept, for the owner to decide about.

Getting back in: reset-password

Because sign-up closes forever, a forgotten owner password used to mean hand-editing state.json inside a container. soundstorm reset-password is the supported path, built into the main binary so it is wherever the server is:

docker compose stop soundstorm
docker compose run --rm soundstorm reset-password
docker compose start soundstorm

The server must be stopped first: a running SoundStorm holds the state in memory and rewrites the whole file on its next change, silently undoing the reset. And because state.Open writes the file every time it opens it, a reset run as root against a volume owned by uid 10001 would hand ownership of the file to root, and the server would then crash-loop on "permission denied". So the command stats the file before opening it and puts the ownership back afterwards. Sessions are left alone: recovering your own password is not evidence of a compromise, and signing every device in the house out would be its own small disaster.