Certificates and names
A home server has no domain name, and no route on which a certificate authority
could reach it. Yet phones refuse to install an app from an address with a certificate warning,
and a session cookie on shared Wi-Fi deserves encryption. internal/servetls,
internal/acme, internal/names and internal/portmap are how
SoundStorm gets a browser-trusted certificate with no account, no domain and no openssl.
Four modes
SOUNDSTORM_TLS picks one. An unknown value is fatal: quietly serving plain HTTP to
somebody who asked for encryption is the worst available way to be wrong.
| Mode | What it serves | For |
|---|---|---|
off | Plain HTTP | The compose file's own default: on localhost there is nothing on the wire to protect. |
self-signed | Certificates from a local authority SoundStorm makes | Encryption without the internet, at the cost of trusting /ca.crt on each device. |
file | A PEM pair given in SOUNDSTORM_TLS_CERT and _KEY | Somebody with a real certificate of their own. |
auto | Plain HTTP and the local authority at once, plus a real Let's Encrypt certificate for <id>.home.soundstorm.dev once it arrives | The installers' default since real certificates became free to get. |
soundstorm.dev took the cost away, so both installers
now write SOUNDSTORM_TLS=auto for new installs and for existing ones that never
chose. A choice somebody made is left alone.The local authority
Self-signed mode (and auto mode's fallback) makes a local authority rather than a bare
self-signed certificate. Install /ca.crt once on a device and everything SoundStorm
issues afterwards is trusted, including certificates for addresses it has never seen. A
self-signed leaf would need re-trusting at every renewal and every new address.
A hostname arrives in the TLS handshake (SNI), so certificates for names are minted on demand.
A bare IP does not - browsers send no SNI for one - and the obvious fallback,
Conn.LocalAddr(), is wrong in exactly the way the software ships: Docker NATs the
published port, so inside the container the local address is the container's own, never the LAN
address the phone dialed. That passed every unit test and every run of the binary directly on the
host. It was caught by a client that trusted only the published authority fetching the LAN address
- the thing the feature exists for. So SOUNDSTORM_TLS_HOSTS is required to reach the
server by IP, and the installer, which runs on the host and can find the LAN address, writes it.
The certificate's names are exactly what was configured plus localhost - not every address the machine has. An early version listed a machine's VPN address, eight IPv6 prefixes and its hostname to anybody who opened a connection.
Name constraints
A device that trusts the authority trusts anything it signs, so a leaked
ca-key.pem would be a way to impersonate any site to that device. The authority
carries critical name constraints: it may sign only localhost, .local,
.lan, .home, home.arpa and configured names; loopback; and
for each configured address its /16 (/48 for IPv6), or exactly itself if public. It used to permit
all of soundstorm.dev, corporate-style names and every private range - and a laptop
that trusts it goes to the office, where 10.0.0.0/8 is somebody else's intranet.
loadOrMakeCA replaces an authority whose constraints are not exactly the current
policy, logs why, and reissues the server certificate if the current authority did not sign it.
The authority itself has MaxPathLenZero, so it cannot mint another authority, and
every private key file is written 0600.
The server certificate is kept
Most devices never install the authority; people click through the browser's warning once,
and a browser pins that exception to the exact certificate it saw. Minting a fresh one on every
start revoked it on every restart, and with a service worker holding the page in cache the
symptom was not a warning but a spinner that never stopped. So server.pem is
persisted beside the authority and reissued only near expiry or when the configured hosts change.
The service worker also never answers a page load, so when a certificate does change - a
reinstall, a yearly renewal - the browser's own warning comes through instead of a cached page
that cannot work.
One port, both protocols
The address everybody already has is http://localhost:8099; the real certificate
is for https://<name>:8099; the published port lives in compose, so both must be
the same port. A TLS connection opens with byte 0x16 and no HTTP request does, so
servetls.Listener peeks at one byte per connection and hands net/http either the plain
connection or a real *tls.Conn. A real *tls.Conn, not a wrapper, because
that type check is what fills in r.TLS, and r.TLS is what makes the
session cookie Secure and __Host-. HTTP/2 still negotiates, which a test asserts.
The listener backs off and retries on accept errors rather than closing - running out of file
descriptors once restarted the server.
Auto mode serves plain HTTP beside TLS even when it has no LAN address to name; falling back to
https-only would have broken the http://localhost address the installer prints.
The name service
The only thing that removes the warning is a certificate somebody on the internet vouches for,
which needs a name on the internet. cmd/soundstorm-names is a small service that gives
each install <id>.home.soundstorm.dev, points it at the install's LAN address,
and publishes the DNS-01 challenge that lets Let's Encrypt issue a certificate for a server
nothing on the internet can reach. It is what Plex does with plex.direct, without the
account. It is the first piece of central infrastructure the project has, and its shape is what
keeps it from becoming a trap:
- It never carries media. A relay's cost grows with every film watched; DNS records and a challenge every couple of months do not. Roughly $5 a month on Railway plus the domain, however many installs.
- It holds no state. An install's token is an HMAC of its id; the DNS provider (Porkbun) is the only record of anything. Once a name resolves it resolves without the service, and an outage stops registration and renewal, never a lookup - with a month of slack, since renewal starts with a third of the certificate's life left.
- It only names private addresses for the home name. A trusted certificate on a public address is a phishing kit with the project's name on it. It also only publishes values shaped like an ACME challenge, and caps challenges per install, per client network and overall, because every certificate under the zone spends the domain's weekly Let's Encrypt allowance until the zone is on the Public Suffix List.
- Every failure falls back to how things were. Until the certificate arrives, or if it never can, the local authority serves exactly as in self-signed mode.
192-168-1-20.<id>...). Railway offers no UDP and no port 53, and it
would have made every lookup depend on the service's uptime. Writing ordinary records through the
registrar's API is less clever and strictly more robust. netip.ParseAddr does the
address check, which rejects the leading-zero and IPv4-mapped-IPv6 tricks that would otherwise
smuggle a public address past IsPrivate().Limits and the sweep
Porkbun allows 2,500 records per domain, and each install holds one for good - roughly 2,400
installs in use, which is a lower ceiling than Let's Encrypt's. So abandoned installs are swept:
the date an install was last heard from is kept in the record's notes field, refreshed monthly by
the install's twice-daily re-announce, and records silent for 180 days are deleted daily. Nothing
is lost by it - the registration is an id and an HMAC, valid for ever, so a server switched back on
after a year recreates its record under the same name on its first announce. The sweep matches only
install-shaped names, skips undated records, and refuses outright to delete more than a quarter of
installs at once: a wrong clock is likelier than a mass exodus. Sign-ups are limited per network
(30 a day per /24 or /48) and overall (300 a day); every limit keys IPv6 by /64. The client
address comes from Railway's X-Real-IP, measured to be unspoofable, where the last
X-Forwarded-For entry would have been Railway's own edge and put every install behind
one limit.
An ACME client of its own
internal/acme exists for the zero-dependency rule: RFC 8555 for one account, a
couple of names, dns-01 and ES256 is a few hundred lines. JWS signatures use fixed-width R and S,
not the DER encoding crypto/ecdsa produces by default. It is believed because of
scripts/acme-rehearsal.sh, which runs it in CI against Pebble - Let's
Encrypt's own test authority - with pebble-challtestsrv standing in for the DNS provider. A fake
authority written beside the client would share its mistakes; Pebble does not. It found one: at
Pebble's 50% nonce-refusal rate, five retries failed a run in six, so it is ten.
The first live run, against real Railway, Porkbun and Let's Encrypt, got a staging certificate in 13 seconds and found three things Pebble could not:
- Compose passes settings by name, and two new variables were missing from
docker-compose.yml. Any newSOUNDSTORM_setting must be added there. - "Service busy; retry later" is typed
rateLimitedbut sent as a 503 when Let's Encrypt sheds load. Taken for a real limit, it silenced the install for a day. Only a 429 is a limit. - A push to main redeploys the service, and a request in flight gets a bare 502. The install retries in five minutes, which is correct.
The issuer is recorded beside the certificate, so changing the ACME directory forces a new one rather than keeping a staging certificate until renewal. The certificate loop checks every twelve hours, and an unanswered check is retried after five minutes.
The page moves itself, after checking
GET /api/session offers secureName to a page not already on it. The
page then fetches https://<name>:<port>/healthz in no-cors
mode - which resolves only if the name resolved, the connection opened and the certificate
verified - and only then location.replaces itself there. Plenty of routers refuse to
resolve a public name that points at a private address (DNS rebinding protection), and redirecting
into that failure would be worse than staying on http. The shell's CSP admits the secure name in
connect-src for exactly this check. HSTS is then sent on that name for a week, and
only while the real certificate is loaded (see the HTTP API).
Remote access: the *.net name
Reaching the server away from home needs a public address, so remote access adds a second name,
<id>.net.soundstorm.dev, which the owner switches on in Settings. Two names, not
one, because many routers cannot "hairpin" - loop a LAN client back in through the public address -
and a single public name would break access at home on those routers. Both names go on one
certificate, one order and one renewal, so remote access does not double an install's draw on the
shared Let's Encrypt allowance.
The name service lifts its private-addresses-only rule for this name only behind a reachability challenge, built to be safe against being turned into a probe of somebody else:
- It only ever probes the caller's own source address - the install supplies a port, never a target - so it cannot be pointed at a victim or at cloud metadata endpoints.
- That address must be public; private, loopback, link-local and carrier-grade NAT addresses are refused.
- It fetches
/api/remote-reachable?nonce=at that address and port and expects an HMAC of the nonce under the install's own token, so a stranger on the same address cannot claim the name.
IPv4 and IPv6 are published from separate calls pinned to each family, so an AAAA
record appears only when the install actually reached the service over IPv6. Audio to a device on
a *.net name is paced (see streaming). After a restart the
remote name is read back from the certificate on disk, so a DNS-provider hiccup at start-up cannot
drop it from the reissued certificate.
Opening the port, and knowing when you cannot
internal/portmap opens the port on the router where it allows, in pure standard
library: PCP first, then NAT-PMP, then UPnP-IGD.
PCP carries a client-address field that a strict router checks against the packet's source, and
behind Docker's NAT those differ, so NAT-PMP, which has no such field, is the fallback. UPnP's SSDP
discovery is multicast and does not cross the Docker bridge, so the installer discovers the
router's description URL on the host and passes it in (SOUNDSTORM_UPNP_URL), as it does
the gateway (SOUNDSTORM_GATEWAY) - inside the container the default route is the Docker
bridge, not the router. UPnP is used only to open this one port at the owner's request, follows no
redirects, and with a gateway configured talks only to a device at that address. The name service's
probe, not any protocol's reply, is the final word on whether the port opened.
The router's own WAN address says when remote access cannot work at all, and
Maintainer.ExternalAddress asks for it even when the mapping succeeded -
because a router behind carrier-grade NAT opens the port without complaint, on an address the
internet cannot reach. ClassifyWAN judges it:
| Router's WAN address | Meaning | What Settings says |
|---|---|---|
In 100.64.0.0/10 | Carrier-grade NAT: no forward can ever be reached | Use Tailscale |
| In a private range | Double NAT: a forward is needed on both routers | Explains, and points at Tailscale |
| Public, or no answer | Ordinary | Forward the port if it did not open by itself |
Tailscale, as a profile
docker compose --profile tailscale up -d runs a Tailscale sidecar that puts the
server on a tailnet at https://<hostname>.<tailnet>.ts.net, with
Tailscale's own certificate. It is for away from home, especially behind carrier-grade NAT where
there is no port to forward, and for anybody who would rather not put a server on the internet.
It is opt-in for the same reason Plex was rejected: it needs an account, an auth key and an app on every device, none of which can be automated. SoundStorm works fully without it, and nobody is walked through a sign-up they did not ask for. It is offered where it is needed - when the WAN check says remote access cannot work - and set up from a Start menu shortcut that takes the key.
Three things were learned by running it. Userspace networking is the default, so the sidecar
needs no NET_ADMIN and no /dev/net/tun. The proxy target must not be
called soundstorm, because that is the sidecar's own tailnet hostname, and Docker writes
a container's hostname into its /etc/hosts: the proxy looped back to itself and
answered 502 while every container reported healthy, so SoundStorm has a second network alias,
soundstorm-app. And the scheme in the serve config must match what SoundStorm speaks on
the compose network - http:// or https+insecure:// - or it is another
silent 502.