The decisions
The code can say what SoundStorm does; it cannot say what was weighed and turned down. These are the choices that shape everything else - including two early instincts that looked right and were reversed.
Own the layer above
The request was for "an all-encompassing server that can do movies, audiobooks, ebooks, music all together". That sounds like build a media server. It is not, and the difference is the whole project.
Two options were rejected:
- Build a media server from scratch. That means owning the expensive part: scanning disks, identifying media, storing metadata and artwork, and transcoding video. Years of work, to arrive somewhere worse than what already exists.
- Fork Jellyfin. All the same ownership, plus a permanent merge tax against a large C# codebase - in exchange for the ability to change internals that none of the requirements touch.
What was chosen instead: SoundStorm owns login, search, playback and the bytes. Each backend owns exactly one media type and one folder. From the user's seat it is indistinguishable from a single server - they never learn Jellyfin exists - but Jellyfin still transcodes, and Navidrome still scans the music.
Why Navidrome for music
Jellyfin can play music, so why run a second server? Because its music support is structurally second-class next to its video. Navidrome brings, for free, the things a music library is judged on: multi-value artist tags, album artist versus track artist, compilations, ReplayGain (which the player uses for leveling), smart playlists and a fast scanner. That music engine is exactly what a Jellyfin fork would have had to rebuild by hand.
It also speaks OpenSubsonic, a documented protocol, which later gave synced lyrics and ReplayGain values through the same adapter.
Why Jellyfin for video
Asked and answered rather than assumed:
| Plex | Requires a plex.tv account, so it cannot be set up without a human logging into a cloud service. That is fatal to the "nobody configures anything" promise, not merely inconvenient. |
|---|---|
| Emby | Went closed-source. |
| Others | Too thin to bet a stack on. |
| Jellyfin | Open source, provisionable through its startup-wizard REST API, and - crucially - a transcoder. |
The honest cost is API churn (Jellyfin 12 broke two documented ways of authenticating) and a 2.5 GB image beside Navidrome's 348 MB. It pays for itself because SoundStorm uses the transcoding: video a browser cannot decode is converted on the fly and served as HLS, and the per-device film quality setting lets Jellyfin copy the picture untouched where it can.
Films and series are separate Jellyfin libraries with different scrapers, so Jellyfin is
provisioned once and registered as two sources sharing one token: jellyfin and
jellyfin-tv.
The ownership rule
There is one place SoundStorm reads a folder itself: ebooks, and later documents. That needed a rule so it would not become a slippery slope back toward building a media server.
| Kind | Self-describing? | Transcoding? | Owner |
|---|---|---|---|
| EPUB | Yes - title, author, language and cover in documented XML inside the file | No | SoundStorm |
| PDF (books and documents) | Enough - an XMP packet or the file name | No | SoundStorm |
| Video | No - Dune.2021.mkv says nothing about itself | Yes | Jellyfin |
| Photos | Partly, but iPhones save HEIC, which Go's standard library cannot decode | Phone clips, yes | Immich |
The rule cuts both ways. It is what said drop Calibre-Web (which needed a database rather than a folder, had no setup API and could not rotate its password), and it is what kept photos delegated: without a HEIC decoder SoundStorm could not show a thumbnail of most phone photos, and a photo library of any size is precisely the scanning-and-indexing job the project exists not to rebuild. Immich was then chosen over lighter options for search by what is in a picture, at the price of four containers and several gigabytes of RAM.
Calibre libraries still work without a SQLite driver, because Calibre writes a
metadata.opf beside every book in the same format an EPUB carries inside it - one
parser reads both.
Two reversals
The first version was an API-only federating gateway. Its instincts were good, and two of them were wrong for this product.
1. "Never proxy media bytes"
The gateway returned absolute upstream URLs and let the browser fetch media straight from Jellyfin or Navidrome. But an upstream URL only works if the browser can reach the upstream, which means publishing each backend on its own port - and then their login screens are one URL away.
internal/stream). That turned out to buy far more than
privacy: picture-free song headers for fast starts away from home, pacing on slow links, resized
covers, guarded content types, and transcodes stopped when a player closes.2. "Stateless, restartable, no config writes"
Zero-keys provisioning means SoundStorm generates the backends' credentials, so it must
remember them. One login means users and sessions. Both have to outlive a restart, so there is now
internal/state.
What SoundStorm deliberately does not do
Each of these has killed a project like this before:
- Transcoding - Jellyfin owns it.
- Metadata scraping - each backend already owns its metadata.
- Library scanning - the same.
internal/libraryreads the folders only to create them and to count files; it must not grow into an indexer. - Client apps for TVs - "the graveyard". The TV apps were later built knowingly, kept to one page (Android TV is the Android app) and one native app (Apple TV, which has no web view).
- Rebuilding an app store - installation is solved and crowded.
- Machine learning over audio - mood analysis and read-along alignment are backends (AudioMuse-AI, Storyteller), run unmodified, not code SoundStorm owns.
SoundStorm also cannot update itself and should never learn how: it has no access to the Docker socket, which is the whole reason a compromise of it cannot reach the host. Updating is running the installer again.
Tailscale is opt-in
Reaching SoundStorm away from home can be done with a Tailscale sidecar
(docker compose --profile tailscale up -d). It is never the default, and that is the
Plex decision applied again: Tailscale needs an account, an auth key and an app on every device,
none of which can be automated. The difference from Plex is that it is remote access, not a
backend - SoundStorm works fully without it, and nobody is walked through a signup they did not
ask for.
It is offered where it is needed: when the router's own WAN address shows carrier-grade NAT
(100.64.0.0/10) or double NAT, where a port forward can never work, and as a standing
line for anybody who would rather not be on the internet.
The names service: the one central piece
A browser trusts only a certificate that somebody on the internet vouches for, and that needs a
name on the internet. cmd/soundstorm-names is a small service that gives each install
<id>.home.soundstorm.dev, points it at the install's LAN address, and publishes
the DNS-01 challenge so Let's Encrypt can issue a certificate for a server nothing on the internet
can reach. It is what Plex does with plex.direct, without the account.
It is the first central infrastructure the project has, a real change made deliberately. Its constraints are what keep it from becoming a trap:
- It never carries media. DNS records and a challenge every couple of months cost about the same for one install or a thousand; a relay's cost grows with every film watched.
- It holds no state. A token is an HMAC of the install's id; the DNS provider's records are the only record of anything. Once a name resolves it resolves without the service, and an outage stops registration and renewal, never a lookup - with a month of slack before a certificate runs out.
- It only names private addresses (plus, with remote access, a separate
netname). A trusted certificate on an arbitrary public address would be a phishing kit with the project's name on it. - Every failure falls back. Until a real certificate arrives, or if it never can, the local certificate authority serves exactly as before.
Answering DNS itself was the first design, but the host offers no UDP, and it would have made
every lookup depend on the service's uptime. Writing ordinary records through the registrar's
API is less clever and strictly more robust. The ACME client is the project's own
(internal/acme), to keep the zero-dependency rule, and is believed because it is
rehearsed against Let's Encrypt's own test server in CI.