AudioMuse-AI (moods)
AudioMuse-AI listens to every song in the library and says how it sounds: danceable, aggressive, happy, party, relaxed, sad, its energy, tempo and key, and which songs sound alike. It is the backend behind mood stations and "sounds like" radio, and nothing it does leaves the house.
Why a backend
Asked for as "we're still missing mood". Measured first, on a real library of 4,413 songs: 42% were tagged just "Pop", none carried a tempo or mood tag, and MusicBrainz's tags name moods only for famous albums. Mood from metadata would have been a guess from "Pop". Hearing mood is machine learning over the audio - the layer SoundStorm does not own - so it is delegated. The owner chose this over tags-only, knowing the cost: three containers (web, worker, Postgres) and a 1.9 GB image. AGPL-3.0, run unmodified.
Pinned by digest
AudioMuse-AI releases weekly, so it is pinned by digest (to 3.6.3).
Provisioned with nobody logging in
Checked live against 3.6.3 in a throwaway stack:
/api/setupanswers without a login until the first save and never after: a second unauthenticated save is a 401 and a made-up token is refused - Jellyfin's startup wizard again. SoundStorm generates its API token and admin password.- It gets a non-admin Navidrome account of its own,
audiomuse, made through Navidrome's native/api/user. It sees every song and is refused a scan. If AudioMuse's volumes were wiped and Navidrome's were not, the account is reset rather than refused. It reads songs through Navidrome, so it mounts no folder at all. - Nothing leaves the house: its models ship in the image (nothing is downloaded), and its third-party lyrics lookup is switched off at setup. Lyrics transcription is off too - hours of CPU for nothing SoundStorm shows.
- A nightly listen is scheduled through its
/api/cron, and SoundStorm's rescan hook starts one (/api/analysis/start) three minutes after a music scan, debounced, once Navidrome has indexed the arrivals.
The calls SoundStorm makes
| Call | Used for |
|---|---|
/api/sync | The whole analysis, 500 a page, include_embeddings=false |
/api/similar_tracks | Song, album and artist radio from what sounds like the seed (a few seeds interleaved) |
/api/last_task | How far the listening has got, shown while it is under way |
/api/analysis/start, /api/cron | Starting and scheduling listens |
Song ids are Navidrome's own, so the analysis joins to SoundStorm's songs with no matching at all. It is registered as a music source that finds nothing in search, so per-account access applies to it like any other source.
Ranks, not raw scores
moods.go): Chill is
relaxed, calm and not aggressive; Focus leans on the instrumental tag. A station takes songs scoring over
0.6, weighted by how far over. Energy, likewise as a rank, also tells the visualizers how hard a song
should hit, and its tempo is the prior the beat tracker leans on.Timing and cost
Measured: about 9 seconds a song on the development machine, so a large library's first listen is roughly 11 hours in the background; worker memory peaked around 700 MB. Until a song is heard, radio from it falls back to similar artists and genres.
Verified end to end on the preview: SoundStorm provisioned a fresh AudioMuse-AI on its first attempt, the 14 test songs were heard, all six mood stations appeared, and song radio came back "Songs that sound like it". Its song-path and text-to-sound search ("rainy day piano") are there for later. Its Postgres password, like Immich's, has a fixed default reachable only on the compose network.