Streaming

Every song, film, cover, page and photo reaches a device through SoundStorm. internal/stream is the proxy that carries them - honoring ranges, taking the pictures out of song headers, pacing audio over slow links, shrinking covers, and making sure nothing it serves can run as a script on SoundStorm's origin.

Why the bytes go through SoundStorm

The first design was an API-only gateway: "never proxy media bytes; results carry absolute upstream URLs". That was one of the two decisions later reversed. An upstream URL only works if the browser can reach the upstream, which means publishing Jellyfin and Navidrome on their own ports - which puts their login screens one URL away. You cannot have "one login" and "never touch the bytes" at once.

So SoundStorm is in the data path. The compose file publishes exactly one port, SoundStorm's; the backends are reachable only on the internal Docker network. An adapter turns an item into a source.Target - a backend URL plus whatever credential headers it needs (Jellyfin 12.1.0 accepts only an Authorization: MediaBrowser ... Token="..." header, which is why a target carries headers and not just a URL), or a local file path, or bytes in memory - and the proxy serves it to the person who asked, signed in as them.

The cost is CPU and a hop for every byte. The alternative is a second login and a visibly different app the moment somebody follows a link, and at that point SoundStorm would be a bookmark folder. A home server's own network is rarely the bottleneck; the work that is expensive - transcoding - still happens in Jellyfin.

The proxy

Proxy.ServeMedia and ServeArt look the source up through the registry (so access applies), ask the adapter for a target, and hand it to pipe:

Range support is the difference between seeking and a file that appears corrupt. Navidrome's converted streams get a length and ranges through estimateContentLength=true (checked: a request from byte 100,000 answered 206), Immich's video playback honors Range, and the slim path below implements ranges itself over a virtual file.

Slim headers: the picture inside the song

Away from home, songs took 15-40 seconds to start. Caching, pacing and a lower bitrate each helped a little; the real cause was measured in Chrome at 0.6 Mbps on a real iTunes M4A:

The same songTime to sound
As it is34s
Tags and artwork taken out of its header, audio byte for byte the same2.5s
A plain generated song1.7s

A browser's media engine treats an embedded picture as a second stream and reads about 2.3MB on before playing. An MP3 with a large picture took 75 seconds. At home those megabytes take a fraction of a second, which is why it never showed.

The files are left exactly as they are - the pictures are most songs' only cover, and other players read them. Instead an original stream is sent slim (internal/stream/slim.go): a header built in memory, followed by the original's audio fetched from Navidrome by byte range.

Any range of the slim file maps onto the original, so seeking works. A layout is cached per song for an hour, at most 4MB a header and 64MB in all, with one fetch per song shared among concurrent requests. On a real song all 397 chunk offsets pointed at the same bytes, 525KB smaller; through the preview at 0.6 Mbps a generated M4A and MP3 started in 1.1s (from 28s and 75s), and a jump to 2:30 played in 1.7s. Malformed files - an ID3 chain cut short, a 64-bit box size that overflows - are refused rather than allowed to panic.

Pacing away from home

The server used to hand every song to the network in full the moment it was asked. Over a slow phone link those megabytes queue up in the connection - HTTP/2 lets a browser take several MB per stream before pushing back, and routers and Docker Desktop's port forwarding buffer more - and every request after them waits behind bytes nobody plays: the fallback copy, the next song after a skip, the covers. A report from a phone showed a first byte arriving 67 seconds after play while the server sat idle.

So audio sent to a device that reached the server by an away-from-home name (*.net.soundstorm.dev or a Tailscale *.ts.net) is paced (internal/stream/pace.go):

It was 2.5 times a guessed 320 kbps at first. The phone got about 0.6 Mbps, so every second played added to a backlog a skip had to drain: skips took 13 and 30 seconds. At 1.5 times the real bitrate, the buffer still grows by half a second every second, and nothing piles up.

At home nothing is paced, nor are films (Jellyfin's HLS pieces are asked for one at a time), nor the small copy the app asks for to hear a song's beats (?listen=1), which paced arrived minutes into the song.

The client half lives in the app: it times 128KB from /api/probe once a session on an away-from-home address, and under 1.2 Mbps starts songs at 128 kbps MP3. A song that cannot play six seconds after starting is restarted at 128 kbps too. While slow, it times the link again every fifteen minutes in the background, and over 3 Mbps the next song is at full quality and the slow note is forgotten; a playing song is never changed. "Always original" in the device's settings turns all of that off.

Card-sized covers

A fresh app on a slow link asked for dozens of full-size covers, 150-200KB of iTunes art each, ahead of the song. Now covers are card-sized unless asked: ?size=, 400px by default, full for the original. Navidrome, Audiobookshelf and Jellyfin resize their own through source.ArtSize; SoundStorm resizes a book's cover itself with the standard library (shrink.go). Only Now Playing's big cover and its swipe neighbours (800px) and the lock screen (600px) ask for more.

Shrinking is something a member can make the server do, so it is bounded: 12 megapixels at most, two decodes at a time, results kept by content; a JPEG with over 100 start-of-scan markers (a progressive file built to take minutes to decode) is sent as it is; a shrink that waits ten seconds for a slot streams the original instead. Covers are cacheable for a week unless the backend says otherwise.

Audio is never cached

Songs were briefly cacheable for a day, and that was a mistake. With audio cacheable, a song that ended by itself waited 20 seconds for the next: the app's preload was fetching the next song while the player asked for the same URL, and Chromium lets only one request at a time write a URL to its cache - the other waits, up to a 20-second timeout. The phone's two requests reached the server exactly 20.0s apart. Audio is now no-store (audioNoStore), and the preload fetches with cache: 'no-store'. Gapless playback keeps the next song in memory, and downloads live in the Cache API, so nothing needed the HTTP cache.

Films: quality, HLS and subtitles

Playback is negotiated, never assumed. GET /api/playback posts a device profile to Jellyfin's /Items/{id}/PlaybackInfo and does what it is told. Whether a file needs transcoding depends on container, video codec, audio codec, profile and level; Jellyfin knows all five. The profile errs conservative - h264/aac in mp4, VPx/AV1 in webm - because claiming a codec the browser cannot decode is the silent failure: Jellyfin hands over the original and the video shows nothing, with no error anywhere. Jellyfin's direct-play answer is checked a second time against the same container list, because 12.1.0 says an MKV can direct-play even when the profile offers only mp4.

Transcodes are served as HLS, for seeking. A progressive transcode has no length until it finishes, so a 20-second file reported 9.9s, growing, and seeking to 15s snapped back to 3s. Jellyfin's HLS is a playlist listing every segment up front: the same file reports 20.0s and seeks. /api/hls/{source}/{path...} mirrors Jellyfin's own /videos/ namespace, so the playlists' relative references resolve back onto SoundStorm with no rewriting.

Quality is per device (?vq=, filmquality.go):

SettingWhat it asks for
Smart (default)The original at home; 20 Mbit by an away-from-home name
Always originalThe original
Standard20 Mbit
Data saver4 Mbit, 720p, stereo

Films used to be capped at 20 Mbit, under a Blu-ray rip's ~40, so Jellyfin re-encoded every frame of a picture any browser could play. Under the cap Jellyfin now copies the picture untouched and converts only sound a browser cannot play (TrueHD, DTS) to AAC with up to 5.1 channels. Each play carries its own playSessionId: Jellyfin names a conversion's files after the file, the device and the play session, and SoundStorm is always one device, so without it every play of a film shared one conversion - another quality or language could have got the old pieces. Each person may have four conversions going.

Another audio language is always Jellyfin's HLS with AudioStreamIndex (?audio=), since a browser plays only a file's first audio track. Subtitles are built, not taken from PlaybackInfo (12.1.0 leaves the delivery URL empty): /Videos/{item}/{source}/Subtitles/{index}/Stream.vtt works for embedded and sidecar tracks and converts SRT on the fly. Only text subtitles are offered; Blu-ray's PGS are pictures and would render nothing.

The HLS query reaches Jellyfin with SoundStorm's admin token, and several Jellyfin options write URLs carrying that token into the playlist that is piped back - HLS subtitle playlists, subtitles in the manifest, trickplay tiles. So every key containing "subtitle" is stripped from the query and enableTrickplay=false is forced. An allowlist would be better and waits on a full list of the keys Jellyfin's own playlists use.

Active-content guards

Everything the proxy serves comes from SoundStorm's own origin, where the session cookie lives. A book, a sidecar or a backend's upload can be HTML, XHTML or SVG with a script in it, and the reader renders EPUB chapters on that origin. Three guards, in stream.go, apply to every stream path:

A local target's type is resolved from its name before the guard runs, because ServeContent would otherwise fill in text/javascript afterwards. These guards exist because the same class of bug - a script reaching the reader's frame by an absolute address on our own origin - was found three times by separate reviews, each time by a different door; see the review passes.

Downloads

The server does not package downloads; the app fetches the same addresses it would play and keeps them in the browser's Cache API. Songs are downloaded as originals (streamPath, never the converted playPath). A film the browser can play is kept as it is; one it cannot is kept as the HLS stream it would play - the playlist and every piece - at Standard quality unless the person chooses otherwise, since an original Blu-ray picture would be tens of gigabytes on a phone. More on the web app page.