Make PulseAudio readiness mean a live connection, clear stale socket state
The pulse_socket named volume survives container recreation, so after a
compose recreate the previous container's /run/pulse/pid and native
socket were still present. PulseAudio read the stale pid file, decided a
daemon was already running, and refused to start:
E: [pulseaudio] pid.c: Daemon already running.
The entrypoint still reported "PulseAudio socket ready" because it only
checked that the socket file existed - and a stale one did. Capture then
failed in a restart loop against a dead daemon.
Readiness in both the op25 entrypoint and drb-edge-node now means a
pactl probe actually succeeds. Stale pid/socket are removed only when
that probe fails, so a live daemon's socket is never deleted.
pulseaudio-utils was missing from the edge-node image (only libpulse0
was installed), so no pactl binary existed there at all - added.
Capture exits are now classified: a missing source logs at ERROR and
names the configured PULSE_SOURCE, rather than looking identical to
"daemon not up yet". Retrying forever against a wrong source name is how
the April PulseAudio failure stayed hidden.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,32 +1,65 @@
|
||||
#!/bin/bash
|
||||
|
||||
# --- Start PulseAudio Daemon ---
|
||||
# -n: skip default config (load modules inline — avoids system.pa parsing issues)
|
||||
# --system: run as system-wide daemon
|
||||
# --log-target=stderr: makes errors visible in Docker logs
|
||||
# &: background so this script continues; output still captured by Docker
|
||||
echo "Starting PulseAudio daemon..."
|
||||
mkdir -p /run/pulse
|
||||
chmod 777 /run/pulse
|
||||
pulseaudio --exit-idle-time=-1 -n --system \
|
||||
--load="module-native-protocol-unix socket=/run/pulse/native auth-anonymous=1" \
|
||||
--load="module-null-sink sink_name=drb_sink sink_properties=device.description=DRB-Sink" \
|
||||
--log-target=stderr &
|
||||
|
||||
# Wait for the socket to actually exist before continuing
|
||||
echo "Waiting for PulseAudio socket..."
|
||||
for i in $(seq 1 20); do
|
||||
if [ -S /run/pulse/native ]; then
|
||||
echo "PulseAudio socket ready."
|
||||
break
|
||||
fi
|
||||
sleep 0.5
|
||||
done
|
||||
if [ ! -S /run/pulse/native ]; then
|
||||
echo "WARNING: PulseAudio socket not found after 10s — edge-node audio will fail."
|
||||
fi
|
||||
ls -la /run/pulse/
|
||||
|
||||
# --- Execute the main command (uvicorn) ---
|
||||
echo "Starting FastAPI application..."
|
||||
exec "$@"
|
||||
#!/bin/bash
|
||||
|
||||
PULSE_SOCKET=/run/pulse/native
|
||||
PULSE_PIDFILE=/run/pulse/pid
|
||||
|
||||
mkdir -p /run/pulse
|
||||
chmod 777 /run/pulse
|
||||
|
||||
# Returns 0 (true) only when a PulseAudio daemon actually answers on
|
||||
# $PULSE_SOCKET. A socket/pid FILE existing proves nothing by itself — that
|
||||
# is exactly the bug this script works around (see stale-state check below).
|
||||
pulse_daemon_alive() {
|
||||
PULSE_SERVER="unix:${PULSE_SOCKET}" timeout 2 pactl info >/dev/null 2>&1
|
||||
}
|
||||
|
||||
# --- Clear stale PulseAudio state left behind by a killed daemon ---
|
||||
# The `pulse_socket` named volume survives container recreation, but the
|
||||
# PulseAudio process that owned it does not. If the previous container was
|
||||
# recreated (not gracefully stopped), its pid file and native socket are
|
||||
# still sitting in the volume; pulseaudio's pid.c sees the pid file and
|
||||
# refuses to start ("Daemon already running") even though nothing is
|
||||
# listening. Only remove these when nothing actually answers on the socket —
|
||||
# never delete a socket a live daemon is using.
|
||||
if [ -S "$PULSE_SOCKET" ] || [ -f "$PULSE_PIDFILE" ]; then
|
||||
if pulse_daemon_alive; then
|
||||
echo "PulseAudio daemon already alive and responding at ${PULSE_SOCKET} — leaving state as-is."
|
||||
else
|
||||
echo "STALE STATE: found ${PULSE_PIDFILE} / ${PULSE_SOCKET} from a previous container, but no daemon answers — clearing before start."
|
||||
rm -f "$PULSE_SOCKET" "$PULSE_PIDFILE"
|
||||
fi
|
||||
fi
|
||||
|
||||
# --- Start PulseAudio Daemon ---
|
||||
# -n: skip default config (load modules inline — avoids system.pa parsing issues)
|
||||
# --system: run as system-wide daemon
|
||||
# --log-target=stderr: makes errors visible in Docker logs
|
||||
# &: background so this script continues; output still captured by Docker
|
||||
echo "Starting PulseAudio daemon..."
|
||||
pulseaudio --exit-idle-time=-1 -n --system \
|
||||
--load="module-native-protocol-unix socket=${PULSE_SOCKET} auth-anonymous=1" \
|
||||
--load="module-null-sink sink_name=drb_sink sink_properties=device.description=DRB-Sink" \
|
||||
--log-target=stderr &
|
||||
|
||||
# Wait for the daemon to actually answer — NOT just for the socket file to
|
||||
# exist. A stale socket file from a killed daemon exists but nothing is
|
||||
# listening on it; a file-existence check reports "ready" against a dead
|
||||
# daemon, which is exactly how this class of bug slipped through before.
|
||||
echo "Waiting for PulseAudio to become live..."
|
||||
PULSE_LIVE=0
|
||||
for i in $(seq 1 20); do
|
||||
if pulse_daemon_alive; then
|
||||
echo "PulseAudio daemon is live (pactl info succeeded)."
|
||||
PULSE_LIVE=1
|
||||
break
|
||||
fi
|
||||
sleep 0.5
|
||||
done
|
||||
if [ "$PULSE_LIVE" -ne 1 ]; then
|
||||
echo "WARNING: PulseAudio daemon not responding after 10s — edge-node audio will fail until it recovers."
|
||||
fi
|
||||
ls -la /run/pulse/
|
||||
|
||||
# --- Execute the main command (uvicorn) ---
|
||||
echo "Starting FastAPI application..."
|
||||
exec "$@"
|
||||
|
||||
Reference in New Issue
Block a user