Music: radio, moods, beats

Navidrome scans the music and serves the bytes. Everything a music app is judged on past that - albums that match the folders, stations that never run dry, moods, lyrics, leveling and an animation that keeps time with the song - is built in the server on top of it, mostly from what is already there, with the two genuinely expensive parts (hearing mood, and lyrics or bios from the internet) delegated or opt-in.

Artists and albums are the folders

Navidrome groups by tags. On a real 4,408-song library that made 1,467 albums out of 551 album folders: an album whose tracks disagree about the year or album artist splits, and "Artist feat. Someone" becomes an artist. SoundStorm files every upload as Artist/Album/track, so the folders are the tidy version, and the music browser shows them (internal/source/subsonic/folders.go): the first folder of a song's path is the artist, the second the album, a disc folder stays in its album, and a song loose in an artist folder is one of their Singles. It is a grouping of Navidrome's own song list, cached like the shelf - nothing reads the disk.

Names are the folders', spelled as the tags spell them when they are the same name with case and punctuation set aside: a folder cannot be "AC/DC", so AC-DC shows as AC/DC. An artist's picture is the cover of their newest album that has one, since Navidrome's own artist image is a placeholder silhouette without an internet agent.

This depended on real paths, and Navidrome had been making them up from the tags unless told otherwise - which looks right exactly when tags and folders agree, the case in every test library. Deleting a song used that path. The fix was ND_SUBSONIC_DEFAULTREPORTREALPATH plus a new client name, because Navidrome reads that setting only once per client.

Radio

Asked for as "blow Plexamp out of the water". Plexamp's radio rests on analysing every track's audio; these stations rest on the whole music shelf, this person's plays and favorites, the sound analysis where AudioMuse-AI has heard a song, and (with discovery on) which artists people play together.

Library mixes come from source.MixSource (Navidrome's random songs and genres): shuffle, recently added, genres with five or more songs, decades.

Moods, by rank

Measured first on a real library: 42% of songs were tagged just "Pop", none carried a mood or tempo tag, and MusicBrainz names moods only for famous albums. Mood from metadata would have been a guess from "Pop". Hearing mood is machine learning over the audio, so it is a backend: AudioMuse-AI, which gives six classifiers (danceable, aggressive, happy, party, relaxed, sad), energy, tempo, key and style tags.

The classifiers answer in a narrow band - a solo piano piece scored 0.58 to 0.63 on all six - so the raw numbers mean little and their order means a lot. Every feature becomes its rank in this library (ties share one), and a mood is an average of ranks (moods.go): Chill is relaxed, calm and not aggressive; Focus leans on the instrumental tag. A mood station takes songs scoring over 0.6, weighted by how far over.

Lyrics

Local lyrics come through OpenSubsonic's getLyricsBySongId - a .lrc beside a song gives synced lines with millisecond starts. LRCLIB (internal/lyrics) fills in songs that have none, and is the first thing SoundStorm sends outside the house about what somebody plays, so it is an owner setting, off by default. LRCLIB because it needs no key or account and has synced lyrics. It is asked only on play, never in bulk, and matched on artist, title, album and duration - duration is what picks the studio take over a live one. Every answer, "none" included, is cached under the state directory for good (none re-asks after 30 days); never in the music folders, which are the user's and scanned.

Discovery

Also off by default (internal/discover). MusicBrainz names the artist (a confident match only - another artist's bio is worse than none), ListenBrainz gives the artists people play in the same sessions, and Wikipedia, found through MusicBrainz's Wikidata link, gives the bio. None needs a key, which is why not Last.fm. MusicBrainz's one request a second is kept to; answers are cached for months. Only artists in the library are ever offered or looked up. It gives artist pages an About and Similar artists, and "More like" mixes - from what is already known, the rest fetched in the background so the mixes page never waits on the network.

Leveling and gapless

Navidrome passes ReplayGain through as OpenSubsonic replayGain. The app uses album gain when an album plays in order, track gain otherwise, with a -6 dB pre-amp so a quiet track can come up, capped by its peak. Checked with tags of 0, -5 and +8 dB: volumes came out at exactly the computed 0.50, 0.28 and 1.00.

It is set through audio.volume, not Web Audio. Routing a phone's music through Web Audio is what stops it when an iPhone locks; iOS ignoring a page's volume is the cheaper loss.

Gapless: the next song is fetched whole into memory while the current one plays. Measured by polling currentTime every 4 ms - timeupdate fires only every ~250 ms and the first measurement taken on it reported its own floor - the gap is 20-23 ms on a fast network and 180-195 ms before falling to 15-17 ms on a throttled one. Streaming quality, slim headers without embedded pictures, and pacing on slow links are on the streaming page.

The server hears every song

The visualizers keep time with the song's actual beats, loudness and hits. The phone used to hear each song itself, which meant downloading every song twice and the first seconds of a hand-started song moving to the tempo only. Now the server hears every song once, ahead of time (internal/beats, GET /api/music/beats, about 11 KB a song), and the app only hears a song itself when the server has not.

Lightning, and how its rule was found

Every look reads the same few signals, and the most visible is "strike": the moment Storm draws lightning and Fireworks bursts. The rule went through several readings of the owner's asks and measurements on real songs:

Analyses carry a version (now 8), so a change of rule hears downloads again, and the Apple TV asks the server for the same answer rather than computing its own.

Training the looks (developer only)

A tool for SoundStorm's developer, not a feature: it exists only with SOUNDSTORM_TRAINING=true in one install's .env, and the routes are owner-only and answer 404 elsewhere. In the Looks sheet the developer taps big moments or slides intensity while a song plays; each song's taps are kept under the state directory's training/ with the device's timing offset, to separate the device's lateness from the hand's.

soundstorm train-looks (internal/training), run inside the container, hears each recorded song again with the same transcoding and tempo prior, finds candidate frames where something rises, gives each 15 measurements (rises and levels per band, loudness, distance from the beat, beat 1-4 and more), lines taps up by the song's median lateness, and fits a logistic regression scored leave-one-song-out against today's rule. On generated click tracks tapped on each bar's first beat, the model caught 98% where the rule caught 13%. train-looks rules scores versions of the hand-written rule against the taps - which is how the 60% lightning threshold was chosen, over 2,270 taps on ten songs. What is learnt ships in SoundStorm for every install; nobody else can train.