Streaming
Every song, film, cover, page and photo reaches a device through SoundStorm.
internal/stream is the proxy that carries them - honoring ranges, taking the
pictures out of song headers, pacing audio over slow links, shrinking covers, and making sure
nothing it serves can run as a script on SoundStorm's origin.
Why the bytes go through SoundStorm
The first design was an API-only gateway: "never proxy media bytes; results carry absolute upstream URLs". That was one of the two decisions later reversed. An upstream URL only works if the browser can reach the upstream, which means publishing Jellyfin and Navidrome on their own ports - which puts their login screens one URL away. You cannot have "one login" and "never touch the bytes" at once.
So SoundStorm is in the data path. The compose file publishes exactly one port, SoundStorm's;
the backends are reachable only on the internal Docker network. An adapter turns an item into a
source.Target - a backend URL plus whatever credential headers it needs (Jellyfin
12.1.0 accepts only an Authorization: MediaBrowser ... Token="..." header, which is
why a target carries headers and not just a URL), or a local file path, or bytes in memory - and
the proxy serves it to the person who asked, signed in as them.
The proxy
Proxy.ServeMedia and ServeArt look the source up through the registry
(so access applies), ask the adapter for a target, and hand it to pipe:
- Local files (ebooks, documents, the starter library) go through
http.ServeContent, which gives Range, ETag and If-Modified-Since handling for free. - Upstream URLs are fetched with the request's own context, forwarding only a short list of request headers (Range among them) and response headers (content type, length, range, caching). When the person stops a film or skips a song, the context ends and the upstream fetch stops with it - a cancelled request is the normal way a video ends and is not logged as an error.
- Anything a target started, such as a Jellyfin transcode, is stopped on
every exit path through
Target.OnDone, which for Jellyfin sendsDELETE /Videos/ActiveEncodings. A transcode is an ffmpeg process that outlives the HTTP request otherwise.
Range support is the difference between seeking and a file that appears corrupt. Navidrome's
converted streams get a length and ranges through estimateContentLength=true
(checked: a request from byte 100,000 answered 206), Immich's video playback honors Range, and
the slim path below implements ranges itself over a virtual file.
Slim headers: the picture inside the song
Away from home, songs took 15-40 seconds to start. Caching, pacing and a lower bitrate each helped a little; the real cause was measured in Chrome at 0.6 Mbps on a real iTunes M4A:
| The same song | Time to sound |
|---|---|
| As it is | 34s |
| Tags and artwork taken out of its header, audio byte for byte the same | 2.5s |
| A plain generated song | 1.7s |
A browser's media engine treats an embedded picture as a second stream and reads about 2.3MB on before playing. An MP3 with a large picture took 75 seconds. At home those megabytes take a fraction of a second, which is why it never showed.
The files are left exactly as they are - the pictures are most songs' only cover, and other
players read them. Instead an original stream is sent slim
(internal/stream/slim.go): a header built in memory, followed by the original's
audio fetched from Navidrome by byte range.
- M4A: the header is the file's
moovwithoutudta/meta, the padding after it dropped, and everystco/co64chunk offset shifted to match (stripMoov,shiftChunkOffsets). Only whenmoovcomes before the audio and the saving is over 16KB; fragmented files are sent as they are. - MP3: the stream starts at the first audio frame, the ID3 tags skipped.
Any range of the slim file maps onto the original, so seeking works. A layout is cached per song for an hour, at most 4MB a header and 64MB in all, with one fetch per song shared among concurrent requests. On a real song all 397 chunk offsets pointed at the same bytes, 525KB smaller; through the preview at 0.6 Mbps a generated M4A and MP3 started in 1.1s (from 28s and 75s), and a jump to 2:30 played in 1.7s. Malformed files - an ID3 chain cut short, a 64-bit box size that overflows - are refused rather than allowed to panic.
Pacing away from home
The server used to hand every song to the network in full the moment it was asked. Over a slow phone link those megabytes queue up in the connection - HTTP/2 lets a browser take several MB per stream before pushing back, and routers and Docker Desktop's port forwarding buffer more - and every request after them waits behind bytes nobody plays: the fallback copy, the next song after a skip, the covers. A report from a phone showed a first byte arriving 67 seconds after play while the server sat idle.
So audio sent to a device that reached the server by an away-from-home name
(*.net.soundstorm.dev or a Tailscale *.ts.net) is paced
(internal/stream/pace.go):
- a burst first - eight seconds of music when the song's bitrate is known, else 1MB, enough to start any song including an iTunes file's ~600KB of index, art and padding;
- then 1.5 times the song's own bitrate (
paceFor). The slim path reads it from the file: themvhdlength for an M4A, the first frame for a constant-bitrate MP3. Otherwise it is the?kbps=asked for, or a guess of 320 kbps compressed or CD rate for lossless.
It was 2.5 times a guessed 320 kbps at first. The phone got about 0.6 Mbps, so every second played added to a backlog a skip had to drain: skips took 13 and 30 seconds. At 1.5 times the real bitrate, the buffer still grows by half a second every second, and nothing piles up.
At home nothing is paced, nor are films (Jellyfin's HLS pieces are asked for one at a time),
nor the small copy the app asks for to hear a song's beats (?listen=1), which paced
arrived minutes into the song.
/api/probe
once a session on an away-from-home address, and under 1.2 Mbps starts songs at 128 kbps MP3. A
song that cannot play six seconds after starting is restarted at 128 kbps too. While slow, it
times the link again every fifteen minutes in the background, and over 3 Mbps the next song is at
full quality and the slow note is forgotten; a playing song is never changed. "Always original"
in the device's settings turns all of that off.Card-sized covers
A fresh app on a slow link asked for dozens of full-size covers, 150-200KB of iTunes art each,
ahead of the song. Now covers are card-sized unless asked: ?size=,
400px by default, full for the original. Navidrome, Audiobookshelf and Jellyfin
resize their own through source.ArtSize; SoundStorm resizes a book's cover itself
with the standard library (shrink.go). Only Now Playing's big cover and its swipe
neighbours (800px) and the lock screen (600px) ask for more.
Shrinking is something a member can make the server do, so it is bounded: 12 megapixels at most, two decodes at a time, results kept by content; a JPEG with over 100 start-of-scan markers (a progressive file built to take minutes to decode) is sent as it is; a shrink that waits ten seconds for a slot streams the original instead. Covers are cacheable for a week unless the backend says otherwise.
Audio is never cached
Songs were briefly cacheable for a day, and that was a mistake. With audio cacheable, a song
that ended by itself waited 20 seconds for the next: the app's preload was fetching the next song
while the player asked for the same URL, and Chromium lets only one request at a time write a URL
to its cache - the other waits, up to a 20-second timeout. The phone's two requests reached the
server exactly 20.0s apart. Audio is now no-store (audioNoStore), and the
preload fetches with cache: 'no-store'. Gapless playback keeps the next song in
memory, and downloads live in the Cache API, so nothing needed the HTTP cache.
Films: quality, HLS and subtitles
Playback is negotiated, never assumed. GET /api/playback posts a
device profile to Jellyfin's /Items/{id}/PlaybackInfo and does what it is told.
Whether a file needs transcoding depends on container, video codec, audio codec, profile and
level; Jellyfin knows all five. The profile errs conservative - h264/aac in mp4, VPx/AV1 in webm -
because claiming a codec the browser cannot decode is the silent failure: Jellyfin hands over the
original and the video shows nothing, with no error anywhere. Jellyfin's direct-play answer is
checked a second time against the same container list, because 12.1.0 says an MKV can direct-play
even when the profile offers only mp4.
Transcodes are served as HLS, for seeking. A progressive transcode has no
length until it finishes, so a 20-second file reported 9.9s, growing, and seeking to 15s snapped
back to 3s. Jellyfin's HLS is a playlist listing every segment up front: the same file reports
20.0s and seeks. /api/hls/{source}/{path...} mirrors Jellyfin's own
/videos/ namespace, so the playlists' relative references resolve back onto
SoundStorm with no rewriting.
Quality is per device (?vq=, filmquality.go):
| Setting | What it asks for |
|---|---|
| Smart (default) | The original at home; 20 Mbit by an away-from-home name |
| Always original | The original |
| Standard | 20 Mbit |
| Data saver | 4 Mbit, 720p, stereo |
Films used to be capped at 20 Mbit, under a Blu-ray rip's ~40, so Jellyfin re-encoded every
frame of a picture any browser could play. Under the cap Jellyfin now copies the picture
untouched and converts only sound a browser cannot play (TrueHD, DTS) to AAC with up to 5.1
channels. Each play carries its own playSessionId: Jellyfin names a conversion's
files after the file, the device and the play session, and SoundStorm is always one device, so
without it every play of a film shared one conversion - another quality or language could have
got the old pieces. Each person may have four conversions going.
Another audio language is always Jellyfin's HLS with
AudioStreamIndex (?audio=), since a browser plays only a file's first
audio track. Subtitles are built, not taken from PlaybackInfo (12.1.0 leaves the
delivery URL empty): /Videos/{item}/{source}/Subtitles/{index}/Stream.vtt works for
embedded and sidecar tracks and converts SRT on the fly. Only text subtitles are offered;
Blu-ray's PGS are pictures and would render nothing.
enableTrickplay=false is forced. An
allowlist would be better and waits on a full list of the keys Jellyfin's own playlists use.Active-content guards
Everything the proxy serves comes from SoundStorm's own origin, where the session cookie
lives. A book, a sidecar or a backend's upload can be HTML, XHTML or SVG with a script in it, and
the reader renders EPUB chapters on that origin. Three guards, in stream.go, apply
to every stream path:
RefusedDestinationrefuses any request whoseSec-Fetch-Destis a script, worker, stylesheet or worklet. Nothing here is ever legitimately loaded that way; an injected<script src>in a book chapter would be.GuardActiveContentdemotes every JavaScript content type totext/plain. A sandbox does nothing for a script pulled in by a script tag, which the shell'sscript-src 'self'would allow; withnosniffa browser refuses to run text/plain as a script.- Documents are sandboxed. HTML, XML, SVG, octet-stream and untyped
responses get
Content-Security-Policy: sandbox; default-src 'none', so a pasted link to a booby-trapped file opens with no script and no origin. Not PDFs - Chrome will not render one in a sandboxed document - and not media, which cannot script.
A local target's type is resolved from its name before the guard runs, because
ServeContent would otherwise fill in text/javascript afterwards. These
guards exist because the same class of bug - a script reaching the reader's frame by an absolute
address on our own origin - was found three times by separate reviews, each time by a different
door; see the review passes.
Downloads
The server does not package downloads; the app fetches the same addresses it would play and
keeps them in the browser's Cache API. Songs are downloaded as originals (streamPath,
never the converted playPath). A film the browser can play is kept as it is; one it
cannot is kept as the HLS stream it would play - the playlist and every piece - at Standard
quality unless the person chooses otherwise, since an original Blu-ray picture would be tens of
gigabytes on a phone. More on the web app page.