AudioMuse-AI (moods)

AudioMuse-AI listens to every song in the library and says how it sounds: danceable, aggressive, happy, party, relaxed, sad, its energy, tempo and key, and which songs sound alike. It is the backend behind mood stations and "sounds like" radio, and nothing it does leaves the house.

Why a backend

Asked for as "we're still missing mood". Measured first, on a real library of 4,413 songs: 42% were tagged just "Pop", none carried a tempo or mood tag, and MusicBrainz's tags name moods only for famous albums. Mood from metadata would have been a guess from "Pop". Hearing mood is machine learning over the audio - the layer SoundStorm does not own - so it is delegated. The owner chose this over tags-only, knowing the cost: three containers (web, worker, Postgres) and a 1.9 GB image. AGPL-3.0, run unmodified.

Pinned by digest

AudioMuse-AI releases weekly, so it is pinned by digest (to 3.6.3).

Provisioned with nobody logging in

Checked live against 3.6.3 in a throwaway stack:

The calls SoundStorm makes

CallUsed for
/api/syncThe whole analysis, 500 a page, include_embeddings=false
/api/similar_tracksSong, album and artist radio from what sounds like the seed (a few seeds interleaved)
/api/last_taskHow far the listening has got, shown while it is under way
/api/analysis/start, /api/cronStarting and scheduling listens

Song ids are Navidrome's own, so the analysis joins to SoundStorm's songs with no matching at all. It is registered as a music source that finds nothing in search, so per-account access applies to it like any other source.

Ranks, not raw scores

The classifiers answer in a narrow band - a solo piano piece scored 0.58 to 0.63 on all six - so the raw numbers mean little and their order means a lot. Every feature becomes its rank in this library (ties share one), and a mood is an average of ranks (moods.go): Chill is relaxed, calm and not aggressive; Focus leans on the instrumental tag. A station takes songs scoring over 0.6, weighted by how far over. Energy, likewise as a rank, also tells the visualizers how hard a song should hit, and its tempo is the prior the beat tracker leans on.

Timing and cost

Measured: about 9 seconds a song on the development machine, so a large library's first listen is roughly 11 hours in the background; worker memory peaked around 700 MB. Until a song is heard, radio from it falls back to similar artists and genres.

Verified end to end on the preview: SoundStorm provisioned a fresh AudioMuse-AI on its first attempt, the 14 test songs were heard, all six mood stations appeared, and song radio came back "Songs that sound like it". Its song-path and text-to-sound search ("rainy day piano") are there for later. Its Postgres password, like Immich's, has a fixed default reachable only on the compose network.