36 Commits
Author SHA1 Message Date
Logan CusanoandClaude Opus 5.5 9b50ba8114 Pin each SDR service to a dongle by serial; OP25 always opens its SDR by serial
CI / lint (push) Successful in 6s
CI / test (push) Successful in 41s
Fixes node-26#11. OP25's generated config said "rtl" (= whichever dongle
enumerates first), so on 2-SDR nodes a decoder could take OP25's dongle
and stop recording. Now:

- sdr_pins: {op25|adsb|ais: serial}, absent = automatic. Every OP25
  config generation (op25_client.generate_config) rewrites the device to
  rtl=<serial>: the pin, else the first dongle's serial, which is what
  "rtl" always opened. Left as "rtl" only for unknown/shared serials.
- secondary-sdr: /secondary/devices lists dongles + serials via librtlsdr
  (works while claimed). apply(priority, pins, reserved) never touches
  OP25's dongle, gives a pinned service only its own dongle, lets a
  higher-priority service take a spare from a lower one, and still runs a
  pinned lower-priority service when the top pick has no dongle.
- sdr_settings.py replaces secondary_priority.py: one apply path for the
  local dashboard, the new set_sdr_config C2 command (set_secondary_priority
  kept as an alias) and config pushes. OP25 restarts only when its own
  dongle changes. Checkin reports sdr_devices, sdr_pins, op25_sdr_serial.
- Local dashboard: 'SDRs' card with an OP25 SDR dropdown and a per-service
  dongle dropdown, duplicate-serial and double-pin warnings.

Verified: edge-node pytest 194 passed; secondary-sdr tests 7 passed;
flake8 clean; page JS passes node --check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 14:35:22 -04:00
Logan CusanoandClaude Opus 5.5 3ecc7eff1a checkin: take sdr_count from secondary-sdr first; op25's 0 is never real
CI / lint (push) Successful in 7s
Build edge-node / build (push) Successful in 44s
CI / test (push) Successful in 48s
op25's :stable image has no lsusb, and /op25/devices answers count 0
instead of unknown, so the dashboard said radio-box 'reports 0 SDRs'
with 2 plugged in. Prefer the secondary-sdr container's lsusb count;
treat 0 from op25 as unknown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 14:23:30 -04:00
Logan CusanoandClaude Opus 5.5 358766d88b Merge feat/secondary-sdr-priority: every spare SDR runs the next decoder in priority order
CI / lint (push) Successful in 11s
CI / test (push) Successful in 55s
Build edge-node / build (push) Successful in 1m12s
Build secondary-sdr / build (push) Successful in 3m29s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 14:14:53 -04:00
Logan CusanoandClaude Opus 5.5 c82fc61910 Local Secondary SDRs card: 'Unsaved' outranks 'Running'; unknown when unreachable
QA blocker: a reordered but unsaved list still showed the old order's
decoder as 'Running'. Also shows 'Unknown' rather than 'Waiting for SDR'
when the secondary-sdr service doesn't answer (server-26#187).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 14:14:49 -04:00
Logan CusanoandClaude Opus 5.5 9b6c64fbbf Secondary SDR priority: every spare SDR runs the next decoder in an ordered list
CI / lint (push) Successful in 6s
CI / test (push) Successful in 39s
Replaces the single secondary_sdr_mode with secondary_sdr_priority, e.g.
["adsb", "ais"]. OP25 always keeps its own dongle; the secondary-sdr
container starts decoders top-down until it runs out of free SDRs, so a
3-SDR node runs ADS-B and AIS at once and a 2-SDR node runs the top pick.

- secondary-sdr: one decoder per mode, POST /secondary/apply(priority)
  (no-op when the right prefix is already running), orphaned decoders
  from a uvicorn reload are reaped on start, /status reports sdr_count
  via lsusb (op25's :stable image has none).
- edge-node: one apply path (set_secondary_priority) for the local
  dashboard, a new C2 'set_secondary_priority' MQTT command, and config
  pushes; it never restarts op25. Legacy mode migrates on load. Checkin
  reports priority, what's running, and sdr_count. Uplink forwards
  aircraft and vessels whenever either is present.
- Local dashboard: 'Secondary SDRs' card to enable/reorder/save.

Verified: edge-node pytest 190 passed, flake8 clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 14:08:07 -04:00
Logan CusanoandClaude Opus 5.5 1c74edffea secondary-sdr: AIS-catcher output flag is -o 5, not -o JSON
CI / lint (push) Successful in 7s
CI / test (push) Successful in 48s
Build secondary-sdr / build (push) Successful in 47m56s
AIS-catcher rejects '-o JSON' (Unknown message format) and exited on
every index. -o 5 is JSON Full, whose field names the reader already
expects. Verified on radio-box: 16 AIS transmitters (Hudson AtoN buoys)
decoded on a 9cm whip. (node-26#9)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 13:32:23 -04:00
Logan CusanoandClaude Opus 5.5 c0a05a8ad6 Merge feat/second-sdr-adsb-ais: second SDR (ADS-B verified on hardware, AIS untested) (node-26#9)
CI / lint (push) Successful in 7s
Build op25 / build (push) Failing after 31s
CI / test (push) Successful in 48s
Build edge-node / build (push) Successful in 8m40s
Build secondary-sdr / build (push) Successful in 47m9s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 13:17:00 -04:00
Logan CusanoandClaude Opus 5.5 52d11c1546 ci: build and publish the secondary-sdr image
Without it, compose references secondary-sdr:latest that no registry
has, and 'make pull' on every node fails once this branch lands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 13:16:58 -04:00
Logan CusanoandClaude Opus 5.5 b54624e176 secondary-sdr: fix ADS-B on real hardware (tested on radio-box)
CI / lint (push) Successful in 6s
CI / lint (pull_request) Successful in 7s
CI / test (push) Successful in 39s
CI / test (pull_request) Successful in 40s
First hardware test of node-26#9/#10 found three blockers:
- antirez/dump1090 has no --write-json, so ADS-B mode exited on start.
  Swapped to wiedehopf/readsb; map alt_baro/gs (old names as fallback).
- op25 is not always on RTL-SDR index 0 (radio-box: op25 on 1, 0 free).
  start() now tries each index and keeps the first decoder that stays up.
  AIS-catcher index flag fixed to -d:N ("-d N" selects by serial).
- status() reported "running" for a decoder that died on startup: the
  zombie still answered killpg(pgid, 0). Liveness now via Popen.poll().

install.sh blacklists dvb_usb_rtl28xxu, which claimed the second dongle.

Verified on radio-box: 831 msgs/min, 6 aircraft (3 with position) on a
9cm whip; op25 unaffected. AIS still untested.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 12:34:18 -04:00
Logan CusanoandClaude Sonnet 5 52f31bbcc0 Wire AIS end to end: AIS-catcher support in secondary-sdr container
CI / lint (push) Successful in 5s
CI / lint (pull_request) Successful in 5s
CI / test (push) Successful in 38s
CI / test (pull_request) Successful in 38s
node-26#9. decoder_control.py's AIS mode now actually launches AIS-catcher
instead of 400ing: a background thread reads its stdout JSON stream
(one message per line) and keeps a live in-memory snapshot keyed by mmsi,
since AIS-catcher streams rather than writing a periodic file the way
dump1090's --write-json does. Messages are merged onto the existing entry
per mmsi rather than replacing it, since AIS-catcher emits static data
(name) and position reports (lat/lon) as separate message types — a naive
overwrite would blank the name back to null on every position-only update
(caught by a standalone unit check against a fake stdout stream before
this fix, not by pytest — this container has no test suite yet).

edge-node's telemetry_uplink_loop now posts non-empty vessel snapshots to
the new /telemetry/ais the same way it already does for aircraft.

UNVERIFIED against real hardware/binary in this session, same caveat as
the ADS-B commit — AIS-catcher's JSON field names are believed correct
from its docs, not confirmed against a real capture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 19:10:30 -04:00
Logan CusanoandClaude Sonnet 5 b2e3804dc3 Wire ADS-B end to end: secondary-sdr container running dump1090
node-26#9. New secondary-sdr-container claims the node's SECOND physical
SDR (RTL-SDR index 1 — op25 always claims index 0; no serial-based binding
yet, same gap op25 itself has). Its control API (start/stop/status/data,
mirroring op25_controller.py) launches dump1090 in adsb mode and exposes
the decoded aircraft.json snapshot; AIS mode 400s until it's wired next.

edge-node: on_config_push starts/stops it when secondary_sdr_mode changes,
lifespan resumes it after a restart if already configured, and a new
telemetry_uplink_loop polls its /secondary/data every 10s and POSTs
non-empty snapshots to C2's new /telemetry/adsb (same bearer-key pattern
call_recorder.py already uses for audio upload).

UNVERIFIED: this container has not been built or run against real hardware
in this session (sandboxed authoring machine, no docker) — dump1090's
--write-json field names are believed correct from its docs but not
confirmed against a real capture. Build + hardware smoke test before this
reaches a real node.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 19:10:24 -04:00
Logan CusanoandClaude Sonnet 5 8f5fca8757 Add second-SDR plumbing: secondary_sdr_mode config + hardware reporting
Adds a per-node secondary_sdr_mode config field (none|adsb|ais|op25_2),
applied the same way as hardware_preset/ppm_override so it survives system
reassignment. op25-container gets a GET /devices endpoint that counts
connected SDRs via lsusb; the edge-node checkin now reports sdr_count and
secondary_sdr_mode up to the server, the first node-initiated hardware
report (everything else was C2 pushing config down). Tracked as node-26#9.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 19:10:17 -04:00
logan a53000a092 ci: publish node images to a public registry on a version tag, DISABLED (#8)
CI / lint (push) Successful in 6s
CI / test (push) Successful in 37s
2026-09-06 23:49:02 -04:00
Logan CusanoandClaude Sonnet 5 bee9e173e6 ci: workflow to publish node images to a public registry on a version tag (DISABLED)
CI / lint (pull_request) Successful in 6s
CI / lint (push) Successful in 7s
CI / test (pull_request) Successful in 37s
CI / test (push) Successful in 39s
git.vpn.cusano.net is now behind REQUIRE_SIGNIN_VIEW (INCIDENT-2026-09-06),
so a fresh Pi can no longer anonymously `docker pull` the node images or
fetch install.sh. Plan (owner, option 3): on a `v*` tag, build the three
node images (edge-node, icecast, op25-client, arm64) and push them to a
PUBLIC registry; source and the private Gitea registry stay walled.

This workflow is wired but INERT — the job is gated on
`vars.NODE_PUBLIC_PUBLISH == 'true'`, which is unset. It triggers on tags
and skips. Enabling is three repo variables + two secrets, documented in the
file header; no code change. Reuses the existing Gitea buildcache refs so
op25 doesn't recompile from scratch.

Not turning it on now — still building the core.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 22:39:42 -04:00
logan 31e1176c45 install.sh: re-run collects the key via pickup_secret, never re-enrolls (#7)
CI / test (push) Successful in 39s
CI / lint (push) Successful in 5s
2026-09-06 22:32:19 -04:00
Logan CusanoandClaude Sonnet 5 4f942bd770 install.sh: re-run collects the key via pickup_secret, never re-enrolls
CI / lint (push) Successful in 12s
CI / test (push) Successful in 51s
CI / test (pull_request) Failing after 11m12s
CI / lint (pull_request) Failing after 11m21s
Confirmed prod bug: §5 keyed idempotency on configs/credentials.json only.
Re-running the installer on a pending-then-approved node (the exact flow the
script's own output tells you to do) found no credentials.json and fell
through to a fresh POST /nodes/enroll — which enrollment.py's CRITICAL GUARD
answers with 403 for an already-approved node_id, rendered as "use Reissue
key". The working path (GET /nodes/{id}/credentials with the saved
pickup_secret — no already-approved guard on that endpoint) was only ever
tried within a single run.

§5 rewritten as a decision tree that runs before any POST /nodes/enroll:
  - credentials.json has api_key            -> skip (unchanged)
  - configs/pickup_secret exists            -> GET /credentials with it:
      200 + api_key   -> write credentials.json, done
      200, no key     -> say "approve it, re-run"; clean exit, start stack
      401 (rotated)   -> re-enroll iff --token, else specific die
      404 (deleted)   -> re-enroll iff --token, else specific die
      000             -> connectivity die
  - no creds, no pickup_secret              -> fresh enroll (do_fresh_enroll)

Also: fresh-enroll 403/401/429 handlers are now specific and point at
pickup-secret recovery, not just "Reissue key"; --wait-approval polling now
applies on the re-run path; §7's "re-run the installer" banner is
conditional on pickup_secret existing.

Response shapes verified against enrollment.py @ v1. approve_node mints
node_keys/{id}.api_key synchronously — no second bug; assigning a system is
independent and not required for a key. bash -n clean.

Recovery for a node stuck by the old behaviour (approved, no credentials.json,
pickup_secret on disk): re-run the patched install.sh (no --token needed), or
  curl -fsS $C2_URL/nodes/$NODE_ID/credentials \
    -H "X-Pickup-Secret: $(cat configs/pickup_secret)" | jq '{api_key}' > configs/credentials.json

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 21:37:14 -04:00
logan 94c3e2a952 Merge pull request 'One-shot install.sh for a fresh Pi; retire setup.sh (node-26#4)' (#5) from feat/one-shot-install into main
CI / lint (push) Successful in 7s
CI / test (push) Successful in 42s
Reviewed-on: #5
2026-09-06 19:29:30 -04:00
Logan CusanoandClaude Sonnet 5 5b0048e8db node: one-shot install.sh for a fresh Pi; retire setup.sh (node-26#4)
CI / lint (pull_request) Successful in 10s
CI / lint (push) Successful in 11s
CI / test (pull_request) Successful in 46s
CI / test (push) Successful in 46s
install.sh is the `curl -fsSL <url> | sudo bash -s -- --token ...` bootstrap:
preflight (root/arch/apt/SDR) → install docker + compose + git + jq → clone
node-26 at a pinned ref (default `v1`, `--track-main` opt-in) → write .env
non-interactively from flags/env with a /dev/tty interactive fallback →
enroll with C2 (POST /nodes/enroll, poll GET /nodes/{id}/credentials, matches
drb-c2-core/app/routers/enrollment.py exactly) and write configs/credentials.json
→ docker compose pull && up -d (prebuilt; --build opts into the ~1h op25 build)
→ print the admin-approval step. Idempotent: re-run picks up the api_key after
approval; existing .env is preserved.

- setup.sh deleted — two scripts writing .env drift. install.sh owns it now.
- Makefile `setup:` no longer calls the removed script (cp .env.example fallback).
- README Setup section rewritten around the one-liner; `make setup`/`make up`
  kept as the local-dev path.

Notes carried in the script header: git.vpn.cusano.net is public (D1, settled);
`v1` must be re-cut at this change's merge commit so the tag actually contains
install.sh. Standing hazard D3: docker-compose.yml bind-mounts the app source
over the image, so a pinned ref and the pulled image tags must not diverge.

Client-side enrollment still belongs in the edge-node app (mqtt_manager.py:73-85);
install.sh doing it is the interim. Tracked for follow-up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 19:14:40 -04:00
Logan CusanoandClaude Opus 5 0c08275482 Stop throwing away the audio Whisper has to read
Build edge-node / build (push) Successful in 35s
CI / lint (push) Successful in 7s
CI / test (push) Successful in 42s
The single encode at save time was %mp3(bitrate=16), and the comment said why:
it matched what Liquidsoap pushes to Icecast. That was the wrong thing to
match. Icecast is the LISTENING path and 16 kbps is a bandwidth budget for a
live stream; this file is the ACCURACY path -- it is what Whisper transcribes,
and CLAUDE.md is explicit that everything downstream is hostage to it. P25 has
already been through a vocoder, so 16 kbps MP3 stacked a second lossy stage on
the one copy that had to stay faithful.

FLAC instead. Lossless, so the bytes Whisper receives are the bytes PulseAudio
captured. ~1.3 MB/min against 120 KB/min, which keeps a 600 s call (the time
cap) around 13 MB -- inside Whisper's 25 MB request cap and well inside
upload_max_bytes. Icecast's own 16 kbps stream is untouched; nothing about
live listening changes.

Capture, buffering, silence detection and the byte-offset trim are all
unchanged: they operate on raw PCM and never saw the encode. The sample rate
stays pinned to pcm.SAMPLE_RATE so the encode remains a straight pass -- the
trim arithmetic depends on that, and Whisper resamples to 16 kHz itself.

encode_mp3 is now encode_recording, the upload sends audio/flac, and the test
that pinned the old contract now pins losslessness instead, including an
assertion that no bitrate constant comes back.

This is the before/after boundary for STT quality. Last night's window is the
16 kbps baseline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 12:48:32 -04:00
Logan Cusano fb13bb8ae3 Pin *.sh to LF so Windows checkouts cannot ship a CRLF shebang
CI / lint (push) Successful in 10s
CI / test (push) Successful in 57s
2026-08-20 03:13:10 -04:00
Logan Cusano e6aab7589a Stop shipping "hackme" as the Icecast password
Build edge-node / build (push) Failing after 6s
CI / lint (push) Successful in 6s
CI / test (push) Successful in 40s
Build op25 / build (push) Successful in 1m48s
Build icecast / build (push) Successful in 3m16s
The source password had a `hackme` fallback in five places -- entrypoint.sh,
docker-compose.yml, setup.sh's prompt default, .env.example, edge-node's
config.py -- plus op25-container's os.getenv default and two README rows. Any
node whose operator pressed Enter through setup.sh is running a credential
that is written down in this repo.

That matters more than the usual default-password case because Icecast binds
all interfaces and the SOURCE password is write access: it does not just let a
LAN neighbour listen, it lets them PUSH audio into the stream the frontend and
mobile clients play as live radio. Injecting fake traffic into a public-safety
feed is the failure worth preventing.

Approach: remove every fallback rather than change them to a better default.
- icecast/entrypoint.sh refuses to start if either password is empty, and says
  how to generate one. This is the single hard gate; everything else is
  defence in depth behind it.
- docker-compose.yml uses ${VAR:?message} so a missing value stops the stack
  at compose time with a readable error instead of becoming an empty string.
- setup.sh GENERATES a random password when the operator presses Enter, via
  openssl rand -base64 24 with a /dev/urandom fallback. Pressing Enter now
  gives you a random password rather than a known one, which is the actual
  behaviour change -- a prompt default nobody types over is not a default, it
  is the value.
- .env.example ships the keys empty with the generation command in a comment,
  and README.md now marks both as required with no default.

Client suite: 185 passed.

Note this does NOT rotate anything already deployed. node-002's .env still has
whatever it was set up with; that is an operational step, tracked in the issue.

Closes logan/node-26#3
2026-08-20 03:12:44 -04:00
Logan Cusano 28266b4441 Pin the op25 container to a Python major version
CI / lint (push) Successful in 6s
Build op25 / build (push) Successful in 21s
CI / test (push) Successful in 38s
`python:slim-trixie` carried no version at all, so a rebuild could move the
interpreter across a major release without anything in the repo changing. That
is not hypothetical here: app/models.py referenced IcecastConfig about 85 lines
before its definition and ran only because trixie currently ships Python 3.14,
where PEP 649 defers annotation evaluation. On 3.13 it was a hard NameError.
The ordering was fixed on 2026-08-16; the unpinned base outlived it.

Pinned to 3.14-slim rather than 3.14-slim-trixie so it matches drb-edge-node,
which was already on 3.14-slim. Patch releases still float, which is what we
want for security updates -- only the major version is nailed down.

Every other Dockerfile in both repos already pinned a major version
(python:3.12-slim, python:3.14-slim, node:20-slim, debian:bookworm-slim), so
this was the only genuinely unpinned base image, despite server-26#11 claiming
none of them were pinned.

Closes logan/server-26#11 (filed against the wrong repo -- the file lives in
the client repo).
2026-08-20 03:01:38 -04:00
Logan CusanoandClaude Opus 5 7f1d09c753 Satisfy flake8 so the lint job stops failing
CI / lint (push) Successful in 6s
CI / test (push) Successful in 39s
Build edge-node / build (push) Successful in 7m44s
Pure formatting, no behaviour change: strip trailing whitespace from blank
lines, give top-level defs in system_cacher.py their two blank lines, wrap
the long discord_radio.join signature, and split the duplicated
active_config ternary in main.py and routers/api.py across lines.

Verified clean with flake8 --max-line-length=120, matching CI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 13:56:22 -04:00
Logan CusanoandClaude Opus 5 9d86304b8a Revert the CI full-fetch workaround and drop unreachable workflows
CI / lint (push) Failing after 17s
CI / test (push) Successful in 49s
Build op25 / build (push) Successful in 1h29m25s
Shallow clones were never a Gitea packing bug. An intruder had set
uploadpack.packObjectsHook in Gitea's HOME gitconfig, pointing at a
non-executable dropper, so every upload-pack died mid-pack. That hook is
gone and --depth=1 clones are verified working, so fetch-depth: 0 buys
nothing but slower CI. See INCIDENT-2026-08-11.md.

op25-container/.gitea/workflows/* never ran: Gitea only executes
workflows under .gitea/workflows at the repo root, and op25-container is
a subdirectory of this repo, not a repo of its own. The live OP25 image
build is .gitea/workflows/build-op25.yml.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 13:00:02 -04:00
Logan CusanoandClaude Opus 5 5ff1089551 Use a full fetch in CI: Gitea fails to pack a shallow clone
CI / test (push) Failing after 30s
CI / lint (push) Failing after 39s
actions/checkout defaults to depth=1, and Gitea aborted generating that pack
with a bad pack header protocol error on every retry, failing the run before
any image was built. Same fix already applied on the server repo; applied here
to all four workflows so the edge-node, op25 and icecast images can build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 10:09:48 -04:00
Logan CusanoandClaude Opus 5 87633ab50d Authenticate the node dashboard, and the broker connection per node
CI / lint (push) Failing after 24s
CI / test (push) Failing after 28s
Build edge-node / build (push) Failing after 43s
Build op25 / build (push) Failing after 47s
Two unauthenticated surfaces closed on the edge node.

Dashboard and API: the local dashboard and every /api/* route were open to
anything on the node's LAN. Adds a login page plus session-cookie auth for
the browser, and cookie-or-Basic for the API so scripted callers stay
possible. Passwords are hashed with stdlib scrypt (no new dependency, this
runs on a Pi) and compared in constant time; the salt and session-signing
secret persist in credentials.json. Startup warns while the default password
is still in place. No non-browser callers of the node API exist today
(C2 talks to nodes over MQTT and nodes call C2 outbound), so nothing breaks.

Adds python-multipart, which FastAPI's Form() needs for the login POST and
which was missing from requirements entirely.

MQTT: nodes authenticated with a shared drb-node password, and the broker
ACL keyed off %c — the client-supplied client id — so any holder of that one
password could claim another node's topic namespace. Nodes now connect as
username=<node_id>, password=<their C2-issued api_key>, which mosquitto's
dynamic-security plugin checks, with the ACL keyed off the authenticated %u.
TLS is gated on MQTT_TLS and uses default CA verification.

The old key_request MQTT path stays in place behind TODO(mqtt-cutover)
markers as the fallback until the cutover is proven; a node with no api_key
on disk logs a clear repeated refusal rather than spinning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 09:34:16 -04:00
Logan CusanoandClaude Opus 5 a61a7b2c31 Bind OP25 control API and terminal to loopback by default
Both :8001 (FastAPI control API) and :8081 (OP25's HTTP terminal) listened on
0.0.0.0 with no authentication, on a container that is privileged with /dev
mounted and network_mode: host. Nodes get deployed to third-party sites, so
that exposed start/stop/retune to anyone on the host's LAN.

All three containers share the host network namespace, so edge-node still
reaches both over 127.0.0.1 unchanged. OP25_DEBUG_EXPOSE=true restores the
old 0.0.0.0 binding and logs a loud warning; it is off by default.

Confirmed against boatbod/op25 gr310 that the terminal's http:<host>:<port>
string is honoured as a real bind address (http_server.py splits it and hands
the host to create_server), so no flag was invented.

Also reorder models.py so IcecastConfig precedes ConfigGenerator, which
annotates a field with it. That only worked because python:slim-trixie is
currently Python 3.14, where PEP 649 defers annotation evaluation; on 3.13 or
earlier the same file is a hard NameError at import.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 09:33:25 -04:00
Logan CusanoandClaude Opus 5 d6dfe5a293 Drive call boundaries from audio, use the console only for the label
CI / lint (push) Failing after 35s
Build edge-node / build (push) Failing after 36s
Build op25 / build (push) Failing after 40s
CI / test (push) Failing after 40s
The control channel was wrong in both directions. Grants fire 0.84-1.62s
before anyone speaks, and srcaddr can drop to 0 while someone is still
talking - one recording came back "-1.61s lead, -0.00s tail", the trim
finding nothing to remove because the window had closed on live speech.
Confirmed by ear: the cut lands at a word boundary on an unfinished word.

Audio is ground truth for WHEN. The console remains the only source of
WHO, so it still supplies talkgroup, alias and rid.

  START  voice onset in the captured audio, with a 0.25s pre-roll that
         now covers only chunk quantisation and threshold ramp-up rather
         than a variable control-channel offset.
  STOP   call_silence_timeout seconds of silence heard in the audio.
  LABEL  resolved AT CLOSE from a bounded rolling history of console
         observations overlapping the window, +4s/-2s, because there is
         no guaranteed ordering between a grant and its audio.
  SPLIT  a console talkgroup change still forces a cut, since two calls
         with no silence between them would otherwise merge into one.

Capture now emits raw PCM instead of MP3. Silence detection becomes
integer arithmetic per chunk with no decode, trimming becomes a byte
offset slice rather than a second ffmpeg pass, and MP3 encoding happens
exactly once at save - uploads are no longer double-encoded.

Audio with no talkgroup anywhere in its window is discarded rather than
uploaded: an untagged call silently poisons incident correlation, which
is worse than losing the audio. Logged at ERROR and counted on
/api/status.

When capture produces no audio at all the old console state machine
still runs, so a node with a broken audio path keeps reporting radio
activity. That is now the only consumer of call_idle_timeout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 18:19:45 -04:00
Logan CusanoandClaude Opus 5 085fcdf1a1 Fix test suite: pad expectations, container pytest config, asyncio markers
CI / lint (push) Failing after 5s
CI / test (push) Successful in 19s
Three unrelated failures from `make test` on the node:

test_metadata_watcher: two tests asserted the old zero-pad tgid_change
contract. Both were off by exactly call_tail_pad_seconds, i.e. the code
was doing what f1de157 intended and the tests encoded the behaviour we
deliberately changed. Updated to expect the pad.

test_pulse: four async tests errored with "async def functions are not
natively supported". Root cause is not the tests - pytest.ini sets
asyncio_mode = auto but the Dockerfile only copies app/ and tests/, so
the container had no pytest config at all and fell back to strict mode.
That is also why these passed in a local venv and failed on the node.
Copy pytest.ini into the image, and add explicit @pytest.mark.asyncio so
the tests hold up regardless of how pytest is configured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 00:14:05 -04:00
Logan CusanoandClaude Opus 5 f1de157d69 Pad the tgid_change closes so a talkgroup switch stops clipping the tail
CI / lint (push) Failing after 4s
CI / test (push) Failing after 22s
The tgid_change and tgid_change_unlogged paths closed the outgoing
segment with no tail pad, on the reasoning that the new grant's
timestamp is an exact, already-known boundary. That is exact only in
control-channel time. The buffered audio lags control timestamps by
~1.5s (measured 0.84-1.62s across seven field calls), so slicing there
cut roughly the outgoing call's last 1.5s of speech - recordings ending
mid-word with ~0s trailing silence.

Both paths now pad, and the recorder's bounded tail wait blocks until
that audio has actually been captured. TAIL_WAIT_TIMEOUT_SECONDS goes
2.0 -> 4.0 so it can satisfy the 3.0s pad instead of giving up and
warning on every talkgroup switch.

The incoming call's pre-roll is served from the ring buffer, so the
delay costs it nothing. The two slices overlapping in the underlying
audio is correct: the stream genuinely contains one call's tail and
then the next call's start.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 23:53:27 -04:00
Logan CusanoandClaude Opus 5 ceb2836371 Raise tail pad to 3s so short transmissions are not clipped
The recording window is anchored to OP25 control-channel timestamps, but
the buffered audio lags those by roughly 1.5s. Trim logs across seven
calls measured the offset at 0.84-1.62s, consistently present.

At a 1.0s pad a short call closed its window before the voice arrived:
a 0.97s control-channel call closed at T+1.97 while voice started around
T+1.5, capturing ~0.4s of speech and cutting mid-word. Confirmed by a
0.57s file whose final 0.10s measured -12.2dB against its own -18.2dB
average - clipped speech, not a tail - and by two short calls that
logged no trim at all because no trailing silence remained.

Being generous is free here: trim_silence already strips trailing
silence back to the guard margin before upload, so long calls are
unaffected while short ones gain the window they need. Over-capture
costs nothing; under-capture loses words permanently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 22:34:17 -04:00
Logan CusanoandClaude Opus 5 7c4a3f2f20 Make PulseAudio readiness mean a live connection, clear stale socket state
The pulse_socket named volume survives container recreation, so after a
compose recreate the previous container's /run/pulse/pid and native
socket were still present. PulseAudio read the stale pid file, decided a
daemon was already running, and refused to start:

    E: [pulseaudio] pid.c: Daemon already running.

The entrypoint still reported "PulseAudio socket ready" because it only
checked that the socket file existed - and a stale one did. Capture then
failed in a restart loop against a dead daemon.

Readiness in both the op25 entrypoint and drb-edge-node now means a
pactl probe actually succeeds. Stale pid/socket are removed only when
that probe fails, so a live daemon's socket is never deleted.

pulseaudio-utils was missing from the edge-node image (only libpulse0
was installed), so no pactl binary existed there at all - added.

Capture exits are now classified: a missing source logs at ERROR and
names the configured PULSE_SOURCE, rather than looking identical to
"daemon not up yet". Retrying forever against a wrong source name is how
the April PulseAudio failure stayed hidden.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 22:10:28 -04:00
Logan CusanoandClaude Opus 5 b0a8ed2a5a Fix tail truncation, decouple call length from buffer, trim silence
CI / lint (push) Failing after 4s
CI / test (push) Successful in 21s
Measured six recordings off a live P25 node and found two independent
defects causing lost audio at the end of calls:

- stop_recording() sliced the ring buffer immediately, so if the MP3
  muxer had not yet delivered the tail the file was silently short.
  Now waits (bounded, 2s) until buffered audio covers the end epoch.
- TGID-change closes used the new grant's epoch as the end with no pad
  at all, guaranteeing truncation on every split. Tail pad is now a
  setting, default raised 0.5s -> 1.0s.

The ring buffer also capped maximum call length: a call longer than the
buffer had its front silently clamped. The ring now serves the pre-roll
only, with a per-call accumulator for the rest, bounded at 4.8MB.
Clamping is loudly warned rather than silent.

Uploads averaged 63% silence, which inflates STT cost and is a known
Whisper hallucination trigger. Leading/trailing silence is now trimmed
conservatively (-40dB, 0.25s guard, internal pauses untouched).
started_at/ended_at still describe the call; new audio_* fields carry
the trimmed audio bounds so playback can map back to wall clock.
All-silence recordings are skipped and logged instead of uploaded.

Also: log measured control-channel idle on idle-timeout closes so
CALL_IDLE_TIMEOUT can be tuned from data, and quiet the httpx logger
which emitted ~170k lines/day of poll noise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 21:39:36 -04:00
Logan Cusano efdbe7d803 Move recording and Discord voice to PulseAudio
CI / lint (push) Failing after 26s
CI / test (push) Successful in 41s
2026-08-04 19:53:45 -04:00
Logan Cusano 9addce7716 chore: Update Makefile targets for local building and git update shorthand 2026-07-13 00:16:50 -04:00
Logan Cusano cc9af6ff26 feat: Buffer offline call_end events and relay them upon MQTT reconnection 2026-07-13 00:16:24 -04:00
62 changed files with 8026 additions and 692 deletions
+105 -5
View File
@@ -5,26 +5,126 @@ NODE_LAT=0.0
NODE_LON=0.0
# MQTT — point to your C2 server
#
# Post-cutover (MQTT-PUBLIC-AUTH-PLAN.md dynsec revision): there is no shared
# node login any more. The node authenticates as username=NODE_ID,
# password=<this node's C2-issued api_key> automatically — nothing to set
# here for that; the api_key is provisioned via MQTT after an admin approves
# the node (see credentials.json) and does not go in this file.
#
# Local/dev, pointed at a plaintext broker on :1883: leave MQTT_TLS unset.
MQTT_BROKER=localhost
MQTT_PORT=1883
# Must match MQTT_NODE_USER/MQTT_NODE_PASS in the server's top-level .env
MQTT_USER=drb-node
MQTT_PASS=change-me-node
MQTT_TLS=false
# Production, pointed at the public broker (real Let's Encrypt cert, default
# CA verification — do not disable it):
# MQTT_BROKER=mqtt.<domain>
# MQTT_PORT=8883
# MQTT_TLS=true
# DEPRECATED / REMOVED post-cutover — the shared "drb-node" login these
# backed no longer exists on the server (dynsec creates only per-node
# clients, keyed by api_key; see Server/drb-c2-core/app/internal/dynsec.py).
# Leave unset for any node pointed at a cut-over broker. Only meaningful as a
# legacy fallback if MQTT_BROKER still points at a pre-cutover broker running
# mosquitto's old password_file auth.
# MQTT_USER=drb-node
# MQTT_PASS=change-me-node
# C2 server for audio upload (leave blank to disable upload)
C2_URL=http://localhost:8888
# API key is provisioned automatically via MQTT after admin approves the node
# Icecast (local container — usually no need to change)
ICECAST_SOURCE_PASSWORD=hackme
ICECAST_ADMIN_PASSWORD=admin
# Live listening only. Call recording and Discord voice use PulseAudio instead.
# REQUIRED, no default — the container refuses to start without them.
# Generate with: openssl rand -base64 24
ICECAST_SOURCE_PASSWORD=
ICECAST_ADMIN_PASSWORD=
ICECAST_HOST=localhost
ICECAST_PORT=8000
ICECAST_MOUNT=/radio
# PulseAudio capture (usually no need to change)
# Monitor of the drb_sink null sink that Liquidsoap writes into.
PULSE_SOURCE=drb_sink.monitor
# Seconds to wait for the shared PulseAudio socket before giving up and retrying.
PULSE_WAIT_TIMEOUT=30
# --- Call segmentation -------------------------------------------------------
# Recording boundaries come from the AUDIO, not the control channel: a recording
# starts at voice onset and ends after this many seconds of silence actually
# heard in the stream. Transmissions on the same talkgroup separated by less
# than this stay in ONE recording, so back-and-forth traffic is one file.
# Tune from the "measured trailing silence" line logged on every close.
CALL_SILENCE_TIMEOUT=3.0
# dBFS (RMS over one ~46ms chunk) below which audio counts as silence and the
# recording is allowed to close.
#
# This does NOT need calibrating against your radio's noise floor. Between
# transmissions the capture is the monitor of a PulseAudio *null sink*, which
# emits DIGITAL silence: measured on a live node it sits at about -91 dBFS —
# one least-significant bit of a 16-bit sample — while speech averages about
# -18 dBFS. Anything from roughly -70 to -40 behaves identically. Only change
# this if you have replaced the audio path with something that has a real
# analog noise floor.
CALL_SILENCE_THRESHOLD_DB=-50
# DEPRECATED as a primary control. Used ONLY when PulseAudio capture is not
# producing audio, where the old control-channel state machine takes over so
# the node still reports radio activity (with no recordings) while its audio
# path is broken.
CALL_IDLE_TIMEOUT=3
# Seconds of audio kept past a CONTROL-CHANNEL-derived boundary — a talkgroup
# change, or a close in the fallback mode above. Buffered audio lags the
# control channel by ~1.5s (grant-to-speech offset measured 0.84-1.62s), so
# cutting at the exact control-channel timestamp clipped the last words of the
# outgoing call. Does NOT apply to the normal end of a call any more; that
# boundary comes from the audio and needs no pad. Safe to be generous — the
# extra is trimmed off again before upload.
CALL_TAIL_PAD_SECONDS=3.0
# Strip leading/trailing dead air before upload. Recordings deliberately
# over-capture at both ends, and silence costs Whisper spend and makes it
# hallucinate text that was never spoken. Trimming is a sample-offset slice of
# the buffered PCM (no re-encode) and only ever touches the head and tail, with
# a guard margin so no syllable is clipped. Set to false to upload raw audio.
TRIM_SILENCE=true
# dBFS (RMS) below which audio counts as silence when trimming the ends. Kept
# stricter than CALL_SILENCE_THRESHOLD_DB on purpose.
TRIM_SILENCE_THRESHOLD_DB=-40
# Seconds of audio kept either side of detected speech.
TRIM_SILENCE_GUARD_SECONDS=0.25
# OP25 container (usually no need to change)
OP25_API_URL=http://localhost:8001
OP25_TERMINAL_URL=http://localhost:8081
# DEBUGGING AID, NOT A DEPLOYMENT OPTION. Both OP25's control API (:8001) and
# its HTTP terminal (:8081) have NO authentication, so they are bound to
# 127.0.0.1 by default — reachable only from other containers on this same
# host (they share its network namespace), not from the site's LAN. Setting
# this to true rebinds both to 0.0.0.0, exposing unauthenticated OP25
# start/stop/config-rewrite and the raw terminal to anyone on that LAN. Only
# for local development off a real node; leave false everywhere else.
OP25_DEBUG_EXPOSE=false
# Secondary SDR container (node-26#9) — only matters if a second physical SDR
# is present and secondary_sdr_mode is set to adsb|ais via the edge dashboard
# or C2. Usually no need to change.
SECONDARY_SDR_API_URL=http://localhost:8002
# Same caveat as OP25_DEBUG_EXPOSE — debugging aid only, leave false.
SECONDARY_SDR_DEBUG_EXPOSE=false
# --- Local dashboard / API login ---------------------------------------------
# Protects the node's local dashboard (port 80) and JSON API. The node is
# reachable by anyone on whatever site's LAN it's deployed to, so this MUST be
# changed before the node leaves the bench — the default below is flagged at
# every startup in the logs until it's changed.
DASHBOARD_USERNAME=admin
DASHBOARD_PASSWORD=CHANGE-ME-drb-default
# Container registry — set these to pull pre-built images instead of building locally.
# Must match the DOCKER_ORG variable and repo name configured in Gitea.
+4
View File
@@ -0,0 +1,4 @@
# Shell scripts run inside Linux containers. A CRLF shebang there fails as
# "bad interpreter: /bin/sh^M", which surfaces only as a container that will
# not start. Windows checkouts have core.autocrlf=true, so pin these to LF.
*.sh text eol=lf
+52
View File
@@ -0,0 +1,52 @@
name: Build secondary-sdr
on:
workflow_dispatch:
push:
branches: [main, master]
paths:
- "secondary-sdr-container/**"
jobs:
build:
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
env:
CONTAINER_NAME: secondary-sdr
steps:
- uses: actions/checkout@v4
- uses: docker/setup-qemu-action@v3
- uses: docker/setup-buildx-action@v3
with:
config-inline: |
[registry."git.vpn.cusano.net"]
http = false
insecure = false
- uses: docker/login-action@v3
with:
registry: git.vpn.cusano.net
username: ${{ gitea.actor }}
password: ${{ secrets.BUILD_TOKEN }}
- name: Get version
id: meta
run: |
echo "REPO_NAME=$(echo ${GITHUB_REPOSITORY} | awk -F'/' '{print $2}')" >> $GITHUB_OUTPUT
echo "VERSION=$(git describe --tags --always | sed 's/^v//')" >> $GITHUB_OUTPUT
- uses: docker/build-push-action@v6
with:
context: ./secondary-sdr-container
file: ./secondary-sdr-container/Dockerfile
platforms: linux/arm64
push: true
tags: |
git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ env.CONTAINER_NAME }}:${{ steps.meta.outputs.VERSION }}
git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ env.CONTAINER_NAME }}:latest
cache-from: type=registry,ref=git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ env.CONTAINER_NAME }}:buildcache
cache-to: type=registry,ref=git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ env.CONTAINER_NAME }}:buildcache,mode=max
+103
View File
@@ -0,0 +1,103 @@
name: Publish public images
# Push the three node images to a PUBLIC registry on a version tag, so a fresh
# Pi can `docker pull` them without a Gitea login (git.vpn.cusano.net is now
# behind REQUIRE_SIGNIN_VIEW — see INCIDENT-2026-09-06). Source + the private
# registry stay walled; only the built node images go public.
#
# ── DISABLED ────────────────────────────────────────────────────────────────
# The job is gated on `vars.NODE_PUBLIC_PUBLISH == 'true'`. Until that repo
# variable is set the workflow triggers on tags but the job is skipped, so
# this file is wired and inert. We're still building the core; flip it on when
# self-serve node install is actually needed.
#
# To enable:
# 1. Repo → Settings → Actions → Variables:
# NODE_PUBLIC_PUBLISH = true
# PUBLIC_REGISTRY = ghcr.io (or docker.io)
# PUBLIC_NAMESPACE = <org-or-user> (images land at <ns>/drb-<name>)
# 2. Repo → Settings → Actions → Secrets:
# PUBLIC_REGISTRY_USER = <push user>
# PUBLIC_REGISTRY_TOKEN = <push token / PAT with write:packages>
# 3. Re-push a tag (or run this workflow via workflow_dispatch).
# ---------------------------------------------------------------------------
on:
workflow_dispatch:
push:
tags:
- "v*"
concurrency:
group: publish-public-${{ github.ref }}
cancel-in-progress: false
jobs:
publish:
# Inert until the repo variable is set. Do NOT convert this to `if: false`
# — the variable is the switch, no code change needed to go live.
if: ${{ vars.NODE_PUBLIC_PUBLISH == 'true' }}
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
strategy:
fail-fast: false
matrix:
include:
- name: edge-node
context: ./drb-edge-node
file: ./drb-edge-node/Dockerfile
cache_name: edge-node
- name: icecast
context: ./icecast
file: ./icecast/Dockerfile
cache_name: icecast
- name: op25-client
context: ./op25-container
file: ./op25-container/Dockerfile
cache_name: op25-client
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # need tags for `git describe`
- uses: docker/setup-qemu-action@v3
- uses: docker/setup-buildx-action@v3
with:
config-inline: |
[registry."git.vpn.cusano.net"]
http = false
insecure = false
# Private Gitea registry — read only, to reuse the existing build cache
# (keeps the op25 image off a ~1h from-scratch compile).
- uses: docker/login-action@v3
with:
registry: git.vpn.cusano.net
username: ${{ gitea.actor }}
password: ${{ secrets.BUILD_TOKEN }}
# Public registry — where the images are pushed.
- uses: docker/login-action@v3
with:
registry: ${{ vars.PUBLIC_REGISTRY }}
username: ${{ secrets.PUBLIC_REGISTRY_USER }}
password: ${{ secrets.PUBLIC_REGISTRY_TOKEN }}
- name: Version
id: meta
run: |
echo "REPO_NAME=$(echo ${GITHUB_REPOSITORY} | awk -F'/' '{print $2}')" >> $GITHUB_OUTPUT
echo "VERSION=$(git describe --tags --always | sed 's/^v//')" >> $GITHUB_OUTPUT
- uses: docker/build-push-action@v6
with:
context: ${{ matrix.context }}
file: ${{ matrix.file }}
platforms: linux/arm64
push: true
tags: |
${{ vars.PUBLIC_REGISTRY }}/${{ vars.PUBLIC_NAMESPACE }}/drb-${{ matrix.name }}:${{ steps.meta.outputs.VERSION }}
${{ vars.PUBLIC_REGISTRY }}/${{ vars.PUBLIC_NAMESPACE }}/drb-${{ matrix.name }}:latest
cache-from: type=registry,ref=git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ matrix.cache_name }}:buildcache
+14 -3
View File
@@ -1,16 +1,21 @@
.PHONY: setup test up up-prebuilt pull down logs
# Local-dev only: seed .env from the example so `make up` has something to read.
# A real edge node is provisioned with install.sh (one-shot bootstrap for a
# clean Pi — clones, enrolls with C2, pulls prebuilt images):
# curl -fsSL https://git.vpn.cusano.net/logan/node-26/raw/tag/v1/install.sh | sudo bash -s -- --help
setup:
@bash setup.sh
@test -f .env || cp .env.example .env
@echo ".env ready — edit it, then 'make up' (local build) or 'make up-prebuilt'"
# Run pytest inside the running edge-node container.
# Requires: docker compose up (or at least the edge-node image built).
test:
docker compose run --no-deps --rm edge-node pytest -v
# Build all images locally and start.
# Build all images locally and start (dev mode with local code mounted).
up:
docker compose up -d
docker compose up -d --build
# Pull pre-built images from the registry and start (no local build).
# Requires IMAGE_REGISTRY, DOCKER_ORG, DOCKER_REPO set in .env.
@@ -21,6 +26,12 @@ up-prebuilt:
pull:
docker compose pull
# Helper to pull the latest git commits and rebuild/restart the dev stack.
update-git:
git pull
docker compose build
docker compose up -d
down:
docker compose down
+25 -14
View File
@@ -135,21 +135,32 @@ Client/
## Setup
### Provision a real node — `install.sh`
One-shot bootstrap for a clean Raspberry Pi OS (arm64). Installs Docker, clones
this repo at the `v1` tag, enrols with C2, pulls the prebuilt images and starts:
```bash
# 1. Copy env template
cp .env.example .env
# 2. Fill in at minimum: NODE_ID and MQTT_BROKER
nano .env
# 3. Build all images (op25 takes ~10-15 minutes first time)
docker compose build
# 4. Start
docker compose up -d
curl -fsSL https://git.vpn.cusano.net/logan/node-26/raw/tag/v1/install.sh \
| sudo bash -s -- --token DRB-xxxx --node-id node-003 \
--c2-url https://api.<domain> --mqtt-broker mqtt.<domain>
```
The node will appear as **pending** in the server admin dashboard. An admin must approve it before it becomes operational. After approval, assign a radio system in the dashboard and the node will start decoding automatically.
Mint the `--token` at **Settings → Nodes** in the web app (the panel prints the
whole command). Run `install.sh --help` for every flag; each also has a
`DRB_*` env var. `--build` compiles op25 on the Pi (~1h) instead of pulling.
### Local dev / manual
```bash
make setup # seeds .env from .env.example
nano .env # at minimum: NODE_ID, MQTT_BROKER, C2_URL
make up # build locally (op25 ~10-15 min first time)
# or: make up-prebuilt # pull images, no local build
```
The node appears as **pending** in the admin dashboard. An admin approves it,
then assigns a radio system, and the node starts decoding automatically.
## Environment Variables (`.env`)
@@ -167,8 +178,8 @@ The node will appear as **pending** in the server admin dashboard. An admin must
| `ICECAST_HOST` | No | `localhost` | Icecast hostname (leave as localhost — host network mode) |
| `ICECAST_PORT` | No | `8000` | Icecast HTTP port |
| `ICECAST_MOUNT` | No | `/radio` | Icecast mount point |
| `ICECAST_SOURCE_PASSWORD` | No | `hackme` | Icecast source password — change this |
| `ICECAST_ADMIN_PASSWORD` | No | `hackme` | Icecast admin password — change this |
| `ICECAST_SOURCE_PASSWORD` | **Yes** | none | Icecast source password. No default — the container refuses to start without it. `install.sh` generates one; otherwise `openssl rand -base64 24` |
| `ICECAST_ADMIN_PASSWORD` | **Yes** | none | Icecast admin password. Same rules |
| `OP25_API_URL` | No | `http://localhost:8001` | OP25 container HTTP API |
| `OP25_TERMINAL_URL` | No | `http://localhost:8081` | OP25 HTTP terminal (live talkgroup metadata) |
+24 -2
View File
@@ -5,9 +5,16 @@ services:
restart: unless-stopped
network_mode: host
environment:
ICECAST_SOURCE_PASSWORD: ${ICECAST_SOURCE_PASSWORD:-hackme}
ICECAST_ADMIN_PASSWORD: ${ICECAST_ADMIN_PASSWORD:-admin}
# :? not :- — a missing password must stop the stack, not silently
# become a credential that is published in this file.
ICECAST_SOURCE_PASSWORD: ${ICECAST_SOURCE_PASSWORD:?set ICECAST_SOURCE_PASSWORD in .env (run setup.sh, or openssl rand -base64 24)}
ICECAST_ADMIN_PASSWORD: ${ICECAST_ADMIN_PASSWORD:?set ICECAST_ADMIN_PASSWORD in .env (run setup.sh, or openssl rand -base64 24)}
# No `ports:` here — network_mode: host makes it a no-op either way. The
# control API (:8001) and OP25's HTTP terminal (:8081) are unauthenticated,
# so they bind 127.0.0.1 by default (see OP25_DEBUG_EXPOSE in .env.example)
# rather than being exposed. edge-node still reaches both over localhost
# because it shares this host network namespace.
op25:
image: ${IMAGE_REGISTRY:-git.vpn.cusano.net}/${DOCKER_ORG:-logan}/${DOCKER_REPO:-node-26}/op25-client:stable
build: ./op25-container
@@ -25,6 +32,21 @@ services:
depends_on:
- icecast
# Claims the node's SECOND physical SDR (op25 always claims the first).
# Only useful if secondary_sdr_mode is set to adsb|ais via the edge-node
# config; otherwise it just sits idle answering /secondary/status. See
# node-26#9. Same network/device access as op25 for the same reason: it
# needs the raw USB device, not a virtualized one.
secondary-sdr:
image: ${IMAGE_REGISTRY:-git.vpn.cusano.net}/${DOCKER_ORG:-logan}/${DOCKER_REPO:-node-26}/secondary-sdr:latest
build: ./secondary-sdr-container
restart: unless-stopped
privileged: true
network_mode: host
env_file: .env
volumes:
- /dev:/dev
edge-node:
image: ${IMAGE_REGISTRY:-git.vpn.cusano.net}/${DOCKER_ORG:-logan}/${DOCKER_REPO:-node-26}/edge-node:latest
build: ./drb-edge-node
+4
View File
@@ -5,6 +5,7 @@ RUN apt-get update && apt-get install -y \
libopus0 \
libopus-dev \
libpulse0 \
pulseaudio-utils \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
@@ -14,5 +15,8 @@ RUN pip install uv && uv pip install --system --no-cache-dir -r requirements.txt
COPY app/ ./app/
COPY tests/ ./tests/
# Without this the container runs pytest with asyncio_mode defaulting to strict,
# so unmarked async tests error out even though they pass locally.
COPY pytest.ini .
CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "80", "--reload"]
+132 -2
View File
@@ -10,28 +10,158 @@ class Settings(BaseSettings):
node_lon: float = 0.0
# MQTT
#
# Broker cutover (MQTT-PUBLIC-AUTH-PLAN.md, dynsec revision): the server no
# longer has a shared node login. Each node authenticates as
# username=NODE_ID, password=<its C2-issued api_key> (the same credential
# /upload already trusts via node_keys) — see mqtt_manager._build_client().
# For local dev against the old-style broker (localhost:1883, no TLS) set
# MQTT_BROKER=localhost and leave MQTT_TLS unset/false.
mqtt_broker: str
mqtt_port: int = 1883
# Set true for the public broker (mqtt.<domain>:8883, real Let's Encrypt
# cert) so client.tls_set() runs with default system-CA verification.
# False by default so local/dev against a plaintext :1883 broker still
# works unchanged. Do NOT pair with a self-signed/insecure cert setup —
# verification is never disabled (no tls_insecure_set(True) anywhere).
mqtt_tls: bool = False
# DEPRECATED / effectively dead post-cutover: the shared node login these
# backed no longer exists on the server (dynsec has no such client — see
# dynsec.py). Left in only as a legacy fallback for a pre-cutover broker
# that still uses mosquitto's old password_file auth; mqtt_manager only
# falls back to these when no api_key is on disk yet. Do not provision new
# nodes with these — see MQTT_USER/MQTT_PASS removal note in .env.example.
mqtt_user: Optional[str] = None
mqtt_pass: Optional[str] = None
# C2 server (audio upload destination); None disables upload
c2_url: Optional[str] = None
# Local Icecast
# Local Icecast — live listening only (frontend / mobile).
# NOT used for call recording or Discord voice: it lags 1s and drifts to 100s+.
icecast_host: str = "localhost"
icecast_port: int = 8000
icecast_mount: str = "/radio"
icecast_source_password: str = "hackme"
# No default: see icecast/entrypoint.sh, which refuses to start without one.
icecast_source_password: str = ""
# PulseAudio — the low-latency path used for call recording and Discord voice.
# Liquidsoap (op25 container) writes into the `drb_sink` null sink; we capture
# its monitor. Addressed explicitly rather than via "default" because the op25
# entrypoint starts pulseaudio with -n and never applies system.pa's
# `set-default-source` line.
pulse_source: str = "drb_sink.monitor"
# Bounded wait for the shared PulseAudio socket before launching FFmpeg.
pulse_wait_timeout: float = 30.0
# ------------------------------------------------------------------
# Call segmentation
#
# Boundaries come from the AUDIO, not the control channel. A recording
# starts at voice onset and ends after call_silence_timeout seconds of
# silence actually heard in the stream. See internal/metadata_watcher.py
# for why the control channel is no longer trusted for either edge.
# ------------------------------------------------------------------
# Seconds of continuous silence IN THE AUDIO before the current recording is
# closed. This is the primary segmentation control. Consecutive
# transmissions on the SAME talkgroup separated by less than this stay in
# one recording, so back-and-forth traffic is one file.
#
# Defaults to 3.0 to match the behaviour of the control-channel idle timer
# it replaces, but it is NOT the same clock: this one measures real silence
# in the audio, with no grant->speech delay mixed in. metadata_watcher logs
# the measured trailing silence on every close — tune from that number.
call_silence_timeout: float = 3.0
# dBFS (RMS, measured over one ~46ms capture chunk) below which audio counts
# as silence for the purpose of ending a recording.
#
# This does NOT need field calibration against radio noise. Between
# transmissions the capture is the monitor of a PulseAudio *null sink*,
# which emits digital silence, not an analog noise floor: measured on a live
# node the gap sits at about -91 dBFS, i.e. one least-significant bit of a
# 16-bit sample. Speech on the same node averages about -18 dBFS. Anything
# between roughly -70 and -40 therefore behaves identically; -50 is chosen
# to sit far below even quiet speech while staying far above the floor.
call_silence_threshold_db: float = -50.0
# DEPRECATED as a primary control — used ONLY in console fallback mode, i.e.
# when PulseAudio capture is not producing audio and there is nothing to
# segment on. Then, and only then, the old control-channel state machine
# runs and closes a segment this many seconds after the last observed
# transmission. Those segments carry no audio; they exist so the node keeps
# reporting real radio activity to C2 while its audio path is broken.
#
# Do NOT tune this against measured *audio* silence — use
# call_silence_timeout for that.
call_idle_timeout: float = 3.0
# Audio kept past a CONSOLE-DERIVED segment boundary, covering the fact that
# buffered audio lags control-channel timestamps by ~1.5s (grant->speech
# offset measured 0.84-1.62s across 7 field calls).
#
# Still needed, with a narrower job than before. It no longer pads the
# normal end of a call — that boundary now comes from the audio itself and
# needs no pad at all. It applies to the three boundaries that are still
# control-channel timestamps:
#
# tgid_change close the outgoing call at the new grant + pad
# tgid_change_unlogged close at the observing poll + pad
# idle_timeout console fallback mode only
#
# Safe to be generous: trim_silence strips trailing silence back to
# trim_silence_guard_seconds before upload, so a larger pad costs long calls
# nothing. Over-capture is free; under-capture loses words permanently. If
# the outgoing and incoming recordings overlap in the underlying audio
# because of this pad, that is correct — the audio contains both.
call_tail_pad_seconds: float = 3.0
# Strip leading/trailing dead air before upload. A recording deliberately
# over-captures at both ends (pre-roll at the head, the whole measured
# silence run at the tail), which inflates Whisper cost and is a
# well-documented trigger for hallucinated transcript text. Trimming is a
# sample-offset slice of the buffered PCM — no re-encode — and only ever
# touches the head and tail. See internal/audio_trim.py.
trim_silence: bool = True
# dBFS (RMS) below which audio counts as silence when trimming the ends.
# Kept above call_silence_threshold_db on purpose: the closer must not miss
# speech (permissive), the trimmer must not leave dead air (stricter), and
# trim_silence_guard_seconds protects the syllable either way.
trim_silence_threshold_db: float = -40.0
# Guard margin kept around detected speech so no syllable is clipped.
trim_silence_guard_seconds: float = 0.25
# OP25 container
op25_api_url: str = "http://localhost:8001"
op25_terminal_url: str = "http://localhost:8081"
# Secondary SDR container (node-26#9) — ADS-B / AIS on a second SDR
secondary_sdr_api_url: str = "http://localhost:8002"
# Paths (volume mounts)
config_path: str = "/configs"
recordings_path: str = "/recordings"
# Offline call buffer — how many call_end events to keep while disconnected
offline_call_buffer_size: int = 35
# ------------------------------------------------------------------
# Local dashboard / API authentication
#
# These nodes are deployed at arbitrary third-party locations, reachable by
# anyone on that site's LAN — there is no auth on this HTTP surface without
# these. The password below is a FIRST-BOOT DEFAULT ONLY: change it via
# DASHBOARD_PASSWORD in .env before a node leaves the bench. main.py logs a
# startup warning every boot the default is still active.
#
# See app/internal/auth.py — the password is never compared or stored in
# plaintext (scrypt-hashed, constant-time compare); this setting just holds
# the operator-facing plaintext the same way MQTT_PASS/ICECAST_* already do.
# ------------------------------------------------------------------
dashboard_username: str = "admin"
dashboard_password: str = "CHANGE-ME-drb-default"
class Config:
env_file = ".env"
+212
View File
@@ -0,0 +1,212 @@
"""
Leading/trailing silence removal, as a slice of raw PCM.
WHY: P25 grants the channel, radios tune, and only then does a human start
talking; the recorder also deliberately over-captures at the tail (it closes a
call only after N seconds of silence have actually been HEARD). Both ends
therefore carry dead air. That is not just wasted Whisper spend: silence is a
well-documented trigger for Whisper hallucinating text that was never spoken,
and a hallucinated sentence poisons entity extraction and then incident
correlation downstream.
WHY IT IS SAFE: only the head and tail are touched, never the middle, and a
guard margin is kept around the detected speech so no syllable can be clipped.
If detection says the whole buffer is silent we do NOT emit a zero-length
recording — the caller is told and decides (see call_recorder: it skips the
upload and logs).
TIMING: trimming changes the audio's duration relative to the call's wall-clock
start/end, so every trim reports exactly how much was removed from each end.
Callers must carry those offsets forward — `started_at`/`ended_at` keep meaning
the CALL's bounds, and the trimmed audio's own bounds are reported separately.
HISTORY — THIS USED TO BE TWO FFMPEG PASSES. Detection was `silencedetect`
parsed out of FFmpeg's stderr, and the cut was a second FFmpeg re-encode. Both
are gone: the recorder now buffers PCM, so detection is arithmetic over the
samples and the cut is a byte-offset slice. Consequences worth keeping in mind:
* The recording is encoded to MP3 exactly ONCE, after this runs, instead of
being captured as MP3 and then re-encoded. One less generation of lossy
encoding on every upload, and one less subprocess per call.
* The threshold is now RMS over a short window (see pcm.rms_dbfs), where
FFmpeg's silencedetect compared |sample| per sample. Same units (dBFS),
slightly different meaning — do not port an old threshold across without
re-reading the field logs.
* There is no "is it worth re-encoding" minimum any more. A slice is free, so
even a 0.05 s trim is applied.
"""
from dataclasses import dataclass
from typing import Optional, Tuple
from app.config import settings
from app.internal import pcm
from app.internal.logger import logger
# Window the head/tail scan works in. 20 ms is short enough that the guard
# margin below dwarfs the quantisation error, and long enough that RMS means
# something.
ANALYSIS_WINDOW_SECONDS = 0.02
# How far in from each end the scan is willing to look before giving up.
#
# Bounds the only unbounded cost in this module: the per-sample RMS loop. A
# normal recording resolves within a window or two at the head (the recorder
# starts on voice onset) and within the silence run at the tail, so this cap is
# never reached in practice. If it IS reached, we leave the audio untrimmed and
# say so — shipping an untrimmed recording is always better than shipping none.
MAX_SCAN_SECONDS = 30.0
@dataclass(frozen=True)
class TrimResult:
"""Outcome of a trim attempt. `lead`/`tail` are seconds actually removed."""
lead: float = 0.0
tail: float = 0.0
duration_before: float = 0.0
duration_after: float = 0.0
all_silence: bool = False
applied: bool = False
# True when the scan hit MAX_SCAN_SECONDS without finding speech, so
# `all_silence` could not be determined and nothing was trimmed.
scan_truncated: bool = False
@property
def trimmed_seconds(self) -> float:
return self.lead + self.tail
def _window_bytes() -> int:
return max(pcm.FRAME_BYTES, pcm.byte_offset(ANALYSIS_WINDOW_SECONDS))
def first_signal_offset(
audio: bytes,
threshold_db: float,
limit_seconds: float = MAX_SCAN_SECONDS,
) -> Optional[int]:
"""
Byte offset of the first window carrying signal, scanning forward.
None means "no signal found" — either the buffer really is all silence or
the scan hit `limit_seconds` first; the caller distinguishes the two by
comparing the scanned span against the buffer length.
"""
window = _window_bytes()
limit = min(len(audio), pcm.byte_offset(limit_seconds) or len(audio))
offset = 0
while offset < limit:
chunk = audio[offset:offset + window]
if not pcm.is_silent(chunk, threshold_db):
return offset
offset += window
return None
def last_signal_offset(
audio: bytes,
threshold_db: float,
limit_seconds: float = MAX_SCAN_SECONDS,
) -> Optional[int]:
"""
Byte offset of the END of the last window carrying signal, scanning back.
Returns the offset one past the last signal-bearing window, so it can be
used directly as a slice bound.
"""
window = _window_bytes()
total = pcm.align(len(audio))
floor = max(0, total - (pcm.byte_offset(limit_seconds) or total))
offset = total
while offset > floor:
start = max(floor, offset - window)
if not pcm.is_silent(audio[start:offset], threshold_db):
return offset
offset = start
return None
def keep_window(
first_signal: Optional[int],
last_signal: Optional[int],
total_bytes: int,
guard_bytes: int,
) -> Tuple[int, int]:
"""
Turn detected signal bounds into the byte range to keep.
Pure and side-effect free so the decision that can destroy a transmission
stays unit-testable without any audio. Offsets are sample-aligned and
clamped to the buffer.
"""
total = pcm.align(total_bytes)
start = 0 if first_signal is None else max(0, first_signal - guard_bytes)
end = total if last_signal is None else min(total, last_signal + guard_bytes)
start = pcm.align(start)
end = pcm.align(end)
if end <= start:
return 0, total
return start, end
def trim_pcm(
audio: bytes,
threshold_db: Optional[float] = None,
guard: Optional[float] = None,
) -> Tuple[bytes, TrimResult]:
"""
Return (kept_audio, result). Never raises and never returns empty audio.
An all-silence buffer is returned UNCHANGED with `all_silence=True`: the
caller decides what to do with a recording that contains no speech at all —
that is itself a signal (squelch misconfigured, wrong sink, dead audio
path), not something to silently truncate to nothing.
"""
threshold = settings.trim_silence_threshold_db if threshold_db is None else threshold_db
margin = settings.trim_silence_guard_seconds if guard is None else guard
total = pcm.align(len(audio))
duration = pcm.seconds(total)
if total <= 0:
return audio, TrimResult()
first = first_signal_offset(audio, threshold)
if first is None:
scanned = min(total, pcm.byte_offset(MAX_SCAN_SECONDS) or total)
if scanned < total:
# Could not prove it is all silence; refuse to guess.
logger.warning(
f"Silence scan gave up after {MAX_SCAN_SECONDS:.0f}s without finding speech in a "
f"{duration:.1f}s recording — leaving it untrimmed."
)
return audio, TrimResult(
duration_before=duration, duration_after=duration, scan_truncated=True
)
logger.warning(
f"Recording is entirely silence ({duration:.2f}s, threshold {threshold:.1f}dBFS RMS) — "
"no speech detected."
)
return audio, TrimResult(duration_before=duration, duration_after=duration, all_silence=True)
last = last_signal_offset(audio, threshold)
guard_bytes = pcm.byte_offset(margin)
keep_start, keep_end = keep_window(first, last, total, guard_bytes)
lead = pcm.seconds(keep_start)
tail = pcm.seconds(total - keep_end)
if keep_start <= 0 and keep_end >= total:
return audio[:total], TrimResult(duration_before=duration, duration_after=duration)
kept = audio[keep_start:keep_end]
after = pcm.seconds(len(kept))
logger.info(
f"Trimmed recording: -{lead:.2f}s lead, -{tail:.2f}s tail "
f"({duration:.2f}s -> {after:.2f}s, threshold {threshold:.1f}dBFS RMS)"
)
return kept, TrimResult(
lead=lead,
tail=tail,
duration_before=duration,
duration_after=after,
applied=True,
)
+178
View File
@@ -0,0 +1,178 @@
"""
Local dashboard / API authentication.
These edge nodes are deployed at arbitrary third-party locations and serve
both an HTML dashboard and a JSON API on the same FastAPI app (port 80,
network_mode: host) — anyone on that site's LAN can otherwise reach every
control endpoint. This module adds username/password auth in front of it.
Design:
- Username + password come from app/config.py (DASHBOARD_USERNAME /
DASHBOARD_PASSWORD env vars), with a first-boot default that MUST be
changed — see is_using_default_password() and its call site in main.py.
- The password is never compared or stored in plaintext. It's hashed with
stdlib hashlib.scrypt (no new dependency — this image runs on a Raspberry
Pi) using a salt generated once on first boot and persisted via
app/internal/credentials.py, then compared with hmac.compare_digest.
- Two auth paths, both accepted on every protected route:
* Browser dashboard: a signed session cookie set by POST /login
(HMAC-SHA256 over "username:expiry", no server-side session store —
the signing key is the persisted session secret from credentials.py).
* Machine callers: HTTP Basic with the same username/password. As of
this writing no non-browser caller of this node's own API was found
anywhere in Client/ or Server/ (nodes are only ever reached over MQTT
+ node-initiated outbound HTTP to C2, never the other way around) —
Basic is kept anyway as a stateless fallback for curl/scripts in the
field, since it needs no login flow and costs little to support.
Caveat worth knowing: this node's dashboard is plain HTTP (no TLS
termination on :80), so both the session cookie and Basic credentials travel
unencrypted on the local network either way. Auth here stops a passerby from
opening the dashboard and pressing buttons; it does not stop a LAN-level
sniffer. That would need TLS in front of the node, which is out of scope here.
"""
import base64
import hashlib
import hmac
import time
from typing import Optional
from fastapi import Cookie, Header, HTTPException, status
from app.config import settings
from app.internal import credentials
from app.internal.logger import logger
SESSION_COOKIE_NAME = "drb_node_session"
# 12h: long enough that the dashboard doesn't demand a daily re-login on a
# device left open on someone's desk, short enough that a stolen cookie isn't
# valid forever.
SESSION_TTL_SECONDS = 12 * 60 * 60
# scrypt cost parameters. n=2**14 (16384) keeps the derivation well under the
# ~1s ballpark on a Raspberry Pi's memory/CPU budget — this only runs on
# login attempts (rare), never on the hot path.
_SCRYPT_N = 2 ** 14
_SCRYPT_R = 8
_SCRYPT_P = 1
_SCRYPT_DKLEN = 32
# Kept in sync with app/config.py's Settings.dashboard_password default.
DEFAULT_PASSWORD = "CHANGE-ME-drb-default"
def _hash_password(password: str, salt: bytes) -> bytes:
return hashlib.scrypt(
password.encode("utf-8"),
salt=salt,
n=_SCRYPT_N,
r=_SCRYPT_R,
p=_SCRYPT_P,
dklen=_SCRYPT_DKLEN,
)
def is_using_default_password() -> bool:
return settings.dashboard_password == DEFAULT_PASSWORD
def verify_credentials(username: str, password: str) -> bool:
"""Constant-time check of a submitted username/password against config."""
salt = credentials.get_auth_salt()
expected_hash = _hash_password(settings.dashboard_password, salt)
submitted_hash = _hash_password(password, salt)
user_ok = hmac.compare_digest(
username.encode("utf-8"), settings.dashboard_username.encode("utf-8")
)
pass_ok = hmac.compare_digest(submitted_hash, expected_hash)
return user_ok and pass_ok
def _sign(payload: str) -> str:
secret = credentials.get_session_secret()
return hmac.new(secret, payload.encode("utf-8"), hashlib.sha256).hexdigest()
def create_session_token(username: str) -> str:
"""Build a signed, expiring, opaque session token (no server-side state)."""
expiry = int(time.time()) + SESSION_TTL_SECONDS
payload = f"{username}:{expiry}"
sig = _sign(payload)
raw = f"{payload}:{sig}"
return base64.urlsafe_b64encode(raw.encode("utf-8")).decode("utf-8")
def _verify_session_token(token: str) -> Optional[str]:
try:
raw = base64.urlsafe_b64decode(token.encode("utf-8")).decode("utf-8")
username, expiry_s, sig = raw.rsplit(":", 2)
expiry = int(expiry_s)
except Exception:
return None
expected_sig = _sign(f"{username}:{expiry_s}")
if not hmac.compare_digest(sig, expected_sig):
return None
if time.time() > expiry:
return None
if not hmac.compare_digest(
username.encode("utf-8"), settings.dashboard_username.encode("utf-8")
):
return None
return username
def _verify_basic_auth(header_value: str) -> bool:
try:
scheme, _, encoded = header_value.partition(" ")
if scheme.lower() != "basic":
return False
decoded = base64.b64decode(encoded).decode("utf-8")
username, _, password = decoded.partition(":")
except Exception:
return False
return verify_credentials(username, password)
def is_authenticated(
session_cookie: Optional[str], authorization: Optional[str]
) -> bool:
if session_cookie and _verify_session_token(session_cookie):
return True
if authorization and _verify_basic_auth(authorization):
return True
return False
async def require_session(
drb_node_session: Optional[str] = Cookie(default=None, alias=SESSION_COOKIE_NAME),
) -> bool:
"""Dependency for dashboard HTML pages. Returns False rather than raising
so the route can redirect to /login instead of showing a bare 401."""
return bool(drb_node_session and _verify_session_token(drb_node_session))
async def require_auth(
drb_node_session: Optional[str] = Cookie(default=None, alias=SESSION_COOKIE_NAME),
authorization: Optional[str] = Header(default=None),
) -> None:
"""Dependency for /api/* routes — session cookie (dashboard's own fetch
calls) or HTTP Basic (machine callers) both satisfy it."""
if is_authenticated(drb_node_session, authorization):
return
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED,
detail="Authentication required",
headers={"WWW-Authenticate": "Basic"},
)
def warn_if_default_password() -> None:
if is_using_default_password():
logger.warning(
"DASHBOARD_PASSWORD is still the first-boot default — "
"set DASHBOARD_USERNAME/DASHBOARD_PASSWORD in .env before this "
"node leaves the bench. Anyone on the node's LAN can currently "
"log in with the default credentials."
)
+702 -94
View File
@@ -1,55 +1,335 @@
"""
Continuous PulseAudio capture: a ring buffer for PRE-ROLL, a per-call
accumulator for the call itself, and the voice-activity signal that decides
where calls begin and end.
A persistent capture process runs for the lifetime of the node. Spawning FFmpeg
per call used to lose the first 1-2 s to process startup, which meant short
transmissions produced empty files, so capture never stops.
RAW PCM, NOT A COMPRESSED STREAM — this is the change everything else hangs
off. FFmpeg is asked for s16le/22050/mono on stdout instead of an encoded
stream, so:
* silence detection is integer arithmetic over each chunk as it arrives, with
no decode, which is what makes AUDIO-DRIVEN call boundaries possible;
* trimming is a byte-offset slice, not a second FFmpeg pass;
* encoding happens exactly ONCE, at save time, so uploads are no longer
double-encoded. That encode is now FLAC (lossless) rather than 16 kbps MP3
— see the AUDIO_* constants below for why.
TWO BUFFERS, TWO JOBS — this split is load-bearing:
RING BUFFER holds the last RING_BUFFER_SECONDS of audio at all times. Its
only job is PRE-ROLL: however late the segmenter notices voice
onset, we can still seek back before it. It is sized for
detection latency, nothing else.
ACCUMULATOR opened by start_recording(), fed by every subsequent chunk, and
closed by stop_recording(). Call length is therefore bounded by
MAX_RECORDING_SECONDS alone — NOT by the ring buffer size. The
old design sliced the finished call back out of the ring buffer,
which silently clamped the front of any call longer than
RING_BUFFER_SECONDS.
VOICE ACTIVITY is tracked CONTINUOUSLY, not only while recording, because the
segmenter starts a call from audio onset. `_last_voice_epoch` and
`_voice_onset_epoch` are the whole interface: metadata_watcher polls them via
audio_activity() and owns every decision about segment boundaries. This module
deliberately does not open or close calls by itself — attribution (which
talkgroup this audio belongs to) lives in the watcher, and audio alone cannot
answer it.
Why not Icecast: it lags ~1 s at connect and drifts progressively to 100 s+, so
slice timestamps and audio content diverge without bound. Icecast stays in the
stack for frontend/mobile live listening; it is not an accuracy path.
CLOCK DOMAIN: chunks are stamped with time.time(), the host wall clock. OP25's
call_log timestamps are time.time() from inside the op25 container. All client
containers run network_mode: host and share the host kernel clock, so the two are
the same clock and the pre-roll arithmetic below is a direct subtraction with no
offset mapping. (A wall-clock STEP — e.g. a large NTP correction — would corrupt
at most the calls in flight at that instant; the buffer self-heals within
RING_BUFFER_SECONDS.)
NOTE on what a chunk timestamp means: it is the ARRIVAL time of that audio at
this process, which lags the moment the words were spoken by the PulseAudio →
FFmpeg → pipe latency. Every timestamp this module produces is in that same
arrival clock, so differences between them are exact; only comparisons against
OP25's control-channel timestamps carry the lag, and those are padded for it.
"""
import asyncio
import time
from collections import deque
from dataclasses import dataclass, field
from datetime import datetime, timezone
from pathlib import Path
from typing import Optional
from typing import List, Optional, Tuple
import httpx
from app.config import settings
from app.internal import credentials
from app.internal import audio_trim, credentials, pcm, pulse
from app.internal.logger import logger
MAX_RECORDING_SECONDS = 600 # safety cap; drop call if it runs this long
PRE_BUFFER_SECONDS = 1.0 # seconds of audio to include before call_start
RING_BUFFER_SECONDS = 60 # how much history to keep when no call is active
READ_CHUNK_BYTES = 4096 # bytes per httpx read
# Safety cap on a single recording; mirrors MAX_SEGMENT_SECONDS in metadata_watcher.
MAX_RECORDING_SECONDS = 600
# Audio included ahead of the detected voice onset.
#
# Under audio-driven segmentation this is no longer covering a variable
# control-channel offset — it covers exactly two things: the analysis chunk
# quantisation (~46 ms) and the possibility that the first syllable ramps up
# through the silence threshold rather than crossing it instantly. 0.25 s is
# generous for both, and anything it drags in that is genuinely silence gets
# trimmed off again before upload.
PRE_ROLL_SECONDS = 0.25
# Rolling history kept for PRE-ROLL ONLY. Budget for the worst realistic
# detection latency: the segmenter polls every 0.5 s and can stall for up to a
# 3 s httpx timeout on a bad OP25 poll, so ~4 s from onset to start_recording().
# 30 s is ~7x that margin. At 44.1 KB/s of PCM it costs ~1.3 MB of RAM. This
# value does NOT bound call length — the accumulator does.
RING_BUFFER_SECONDS = 30
# ~46 ms of audio per chunk. Chunk size is BOTH the timestamp resolution of the
# ring buffer AND the window silence detection runs over, so it has to stay well
# under PRE_ROLL_SECONDS and well under the shortest utterance we care about.
READ_CHUNK_BYTES = 2048
# Encoder settings for the single encode at save time.
#
# This used to be %mp3(bitrate=16) chosen to match what Liquidsoap pushes to
# Icecast. That was the wrong thing to match: Icecast is the LISTENING path and
# 16 kbps is a bandwidth budget for a live stream, while this file is the
# ACCURACY path — it is what Whisper transcribes, and the transcript is what
# every downstream stage is hostage to. P25 audio has already been through a
# vocoder; 16 kbps MP3 stacked a second lossy stage on top of that, on the one
# copy that had to stay faithful.
#
# FLAC instead: lossless, so the bytes Whisper receives are the bytes PulseAudio
# captured. Roughly 1.3 MB/min against 120 KB/min for 16k MP3 — larger, but a
# 600 s call is still ~13 MB, inside both Whisper's 25 MB request cap and the
# C2 upload_max_bytes (100 MB). Icecast's own 16 kbps stream is untouched;
# nothing about the listening path changes.
#
# AUDIO_SAMPLE_RATE MUST equal pcm.SAMPLE_RATE: the encode is a straight pass
# with no resampling. Whisper resamples to 16 kHz itself, so handing it 22050
# unresampled keeps the one resample in the pipeline inside the model.
AUDIO_SAMPLE_RATE = str(pcm.SAMPLE_RATE)
AUDIO_FORMAT = "flac"
AUDIO_SUFFIX = ".flac"
AUDIO_MIME = "audio/flac"
# -compression_level 5 is ffmpeg's default: near-best ratio, and the encode is
# off the hot path anyway (once per call, at save).
FLAC_COMPRESSION_LEVEL = "5"
# Bounded so a wedged encoder can never stall the upload path.
ENCODE_TIMEOUT_SECONDS = 60.0
# Hard memory ceiling for one call's accumulator.
#
# PCM costs 44.1 KB/s where the old MP3 buffer cost 2 KB/s, so this had to be
# re-derived rather than carried over. 600 s (the time cap) of PCM is 26.5 MB;
# 32 MiB is ~761 s, which guarantees the TIME cap always bites first and a legal
# call is never truncated by the byte cap. Peak resident audio is therefore
# ~33.5 MB for the accumulator plus ~1.3 MB for the ring buffer.
#
# Rejected alternatives, for the record: spilling to disk (SD-card wear on a Pi,
# and I/O in the close path); encoding incrementally into MP3 as chunks arrive
# (puts a subprocess back in the hot path and makes the sample-accurate post-hoc
# trim impossible); a lower capture sample rate (changes what Whisper receives).
MAX_RECORDING_BYTES = 32 * 1024 * 1024
# How long stop_recording() will wait for captured audio to actually reach the
# call's end timestamp. PulseAudio → FFmpeg → our pipe read is a pipeline with
# latency, so at the instant a CONTROL-CHANNEL derived end is computed the
# newest buffered chunk is typically a few hundred ms OLDER than that epoch.
# Slicing immediately therefore cuts the tail short — which costs the last word
# of the transmission, usually the disposition or the address.
#
# An audio-driven close never needs this (its end epoch is derived from audio
# that is already buffered, by construction), but a tgid_change close still
# pads past a control-channel timestamp that is ~now, so the wait must exceed
# settings.call_tail_pad_seconds (default 3.0) or it would warn on every
# talkgroup switch. Bounded so a dead capture can never hang the upload path.
TAIL_WAIT_TIMEOUT_SECONDS = 4.0
TAIL_WAIT_POLL_SECONDS = 0.05
# Backoff bounds for restarting a dead capture process.
RESTART_BACKOFF_MIN = 1.0
RESTART_BACKOFF_MAX = 15.0
# Substrings looked for in FFmpeg's stderr to classify why capture exited.
# "No such process" is what FFmpeg's pulse input prints when the daemon is up
# but the named source does not exist — a real PULSE_SOURCE misconfiguration,
# not a startup race. This is exactly the failure mode that hid unnoticed
# behind a generic "restarting" log in the April outage: silently retrying
# forever against a wrong source name looks identical to a normal startup
# wait unless it is logged differently.
_SOURCE_MISSING_MARKERS = ("no such process", "no such device")
# "Connection refused"/"Connection failure" is what the pulse client library
# prints when nothing is listening on the socket at all — expected while the
# op25 container's daemon is still coming up.
_NO_DAEMON_MARKERS = ("connection refused", "connection failure")
@dataclass(frozen=True)
class AudioActivity:
"""
What the capture stream is doing right now, as raw facts.
Deliberately carries no decision — no "should a call be open" boolean —
because every threshold comparison belongs to metadata_watcher, which owns
the segment state machine and (in tests) an injectable clock. This is a
snapshot of observations, nothing more.
`voice_onset_epoch` is the arrival timestamp of the first chunk of the most
recent run of voice. It is NOT cleared when that run ends, so a caller must
check `last_voice_epoch` against its own clock before treating the run as
live. That is intentional: the segmenter needs the onset of the run it just
finished recording in order to avoid re-opening on the same run.
"""
capturing: bool
recording: bool
last_voice_epoch: Optional[float] = None
voice_onset_epoch: Optional[float] = None
# Convenience for /api/status only; the segmenter recomputes this against
# its own clock.
silence_seconds: float = 0.0
@dataclass
class _ActiveRecording:
"""Audio accumulating for the call currently being recorded."""
call_id: str
call_start: float # detected voice onset (or a caller-supplied epoch)
slice_start: float # call_start - PRE_ROLL_SECONDS
chunks: List[Tuple[float, bytes]] = field(default_factory=list)
total_bytes: int = 0
# Seconds of requested pre-roll that were not in the buffer at open time.
clamped_seconds: float = 0.0
# True once the byte ceiling was hit and audio started being dropped.
truncated_by_cap: bool = False
@dataclass
class Recording:
"""
A finished recording plus the timing metadata needed to map audio position
back to wall clock.
`started_at`/`ended_at` upstream keep meaning the CALL's bounds. These are
the AUDIO's bounds, which differ once silence is trimmed:
wall_clock_of(audio_offset_t) == audio_start_epoch + t
Bounds are accurate to ±one capture chunk (~46 ms).
"""
call_id: str
path: Optional[Path]
audio_start_epoch: float
audio_end_epoch: float
lead_trimmed: float = 0.0
tail_trimmed: float = 0.0
clamped_seconds: float = 0.0
all_silence: bool = False
async def encode_recording(audio: bytes, path: Path) -> bool:
"""
The one and only encode in the pipeline: raw PCM in, FLAC file out.
Lossless on purpose — see the AUDIO_* constants above. This file is what
Whisper transcribes, so the encode must not throw anything away.
Module-level rather than a method so tests can substitute it without
needing FFmpeg, and so the "exactly one encode per call" property is
trivially observable.
"""
if not audio:
return False
cmd = [
"ffmpeg",
"-hide_banner", "-nostdin", "-nostats",
"-loglevel", "warning", "-y",
"-f", "s16le",
"-ar", AUDIO_SAMPLE_RATE,
"-ac", str(pcm.CHANNELS),
"-i", "pipe:0",
"-ar", AUDIO_SAMPLE_RATE,
"-ac", str(pcm.CHANNELS),
"-compression_level", FLAC_COMPRESSION_LEVEL,
"-f", AUDIO_FORMAT, str(path),
]
try:
proc = await asyncio.create_subprocess_exec(
*cmd,
stdin=asyncio.subprocess.PIPE,
stdout=asyncio.subprocess.DEVNULL,
stderr=asyncio.subprocess.PIPE,
)
except Exception as e:
logger.error(f"Could not launch the MP3 encoder ({e}) — recording not saved.")
return False
try:
_, stderr = await asyncio.wait_for(proc.communicate(audio), timeout=ENCODE_TIMEOUT_SECONDS)
except asyncio.TimeoutError:
try:
proc.kill()
except Exception:
pass
logger.error(f"MP3 encode timed out after {ENCODE_TIMEOUT_SECONDS:.0f}s — recording not saved.")
return False
except Exception as e:
logger.error(f"MP3 encode failed ({e}) — recording not saved.")
return False
if proc.returncode != 0:
logger.error(f"MP3 encode exited {proc.returncode}: {stderr.decode(errors='replace').strip()}")
return False
return True
class CallRecorder:
"""
Maintains a persistent HTTP connection to the Icecast stream and buffers
the raw MP3 bytes in a ring buffer. When a call starts we note the
monotonic clock; when it ends we slice the buffer and write the file.
This approach eliminates per-call FFmpeg startup latency, which was
causing empty recordings for calls shorter than ~1–2 s.
"""
"""Continuous PulseAudio capture: ring buffer for pre-roll, accumulator per call."""
def __init__(self):
self._recordings_dir = Path(settings.recordings_path)
# Ring buffer: deque of (monotonic_time, bytes_chunk)
self._buffer: deque[tuple[float, bytes]] = deque()
# Ring buffer: deque of (wall_clock_epoch_at_arrival, pcm_bytes).
# Pre-roll only — see the module docstring.
self._buffer: deque[Tuple[float, bytes]] = deque()
self._buffer_bytes: int = 0
self._stream_task: Optional[asyncio.Task] = None
self._proc: Optional[asyncio.subprocess.Process] = None
self._capturing: bool = False
# Active call state
self._call_id: Optional[str] = None
self._call_start_mono: Optional[float] = None
# Voice activity, tracked continuously — not only while recording.
self._last_voice_epoch: Optional[float] = None
self._voice_onset_epoch: Optional[float] = None
# Active recording state (None when idle)
self._active: Optional[_ActiveRecording] = None
# Last few lines of the most recent FFmpeg stderr, used to classify
# why a capture process exited (see _classify_capture_exit).
self._last_stderr_lines: deque[str] = deque(maxlen=10)
# ------------------------------------------------------------------
# Lifecycle
# ------------------------------------------------------------------
async def start(self) -> None:
"""Start the persistent stream buffer. Call once from app lifespan."""
self._stream_task = asyncio.create_task(self._stream_loop())
logger.info("Stream ring-buffer started.")
"""Start the persistent capture. Call once from app lifespan."""
self._stream_task = asyncio.create_task(self._capture_loop())
logger.info("PulseAudio ring-buffer starting.")
async def stop(self) -> None:
"""Cancel the stream reader."""
if self._stream_task:
self._stream_task.cancel()
try:
@@ -57,111 +337,415 @@ class CallRecorder:
except asyncio.CancelledError:
pass
self._stream_task = None
await self._terminate_proc()
# ------------------------------------------------------------------
# Stream reader
# Capture
# ------------------------------------------------------------------
async def _stream_loop(self) -> None:
stream_url = (
f"http://{settings.icecast_host}:{settings.icecast_port}"
f"{settings.icecast_mount}"
)
timeout = httpx.Timeout(connect=10.0, read=None, write=10.0, pool=10.0)
def _ffmpeg_command(self) -> List[str]:
return [
"ffmpeg",
"-hide_banner", "-nostdin", "-nostats",
"-loglevel", "warning",
"-f", "pulse", "-i", settings.pulse_source,
"-ac", str(pcm.CHANNELS),
"-ar", AUDIO_SAMPLE_RATE,
# Raw PCM on stdout. No muxer, so no -flush_packets games: s16le is
# a bare byte stream and every byte FFmpeg produces is immediately
# readable, which is what keeps arrival timestamps honest.
"-f", "s16le", "-",
]
async def _capture_loop(self) -> None:
backoff = RESTART_BACKOFF_MIN
while True:
try:
async with httpx.AsyncClient(timeout=timeout) as client:
async with client.stream("GET", stream_url) as response:
response.raise_for_status()
logger.info(f"Stream buffer connected to {stream_url}")
async for chunk in response.aiter_bytes(READ_CHUNK_BYTES):
self._ingest(chunk)
# The original bug this path was abandoned for: FFmpeg was launched
# with -f pulse before the shared socket existed, failed instantly,
# and never recovered. Wait for it, bounded, every time.
if not await pulse.wait_until_ready():
await asyncio.sleep(backoff)
backoff = min(backoff * 2, RESTART_BACKOFF_MAX)
continue
await self._run_capture()
self._log_capture_exit()
except asyncio.CancelledError:
await self._terminate_proc()
raise
except Exception as e:
logger.warning(f"Stream buffer disconnected ({e}) — retrying in 3 s")
await asyncio.sleep(3)
logger.warning(f"PulseAudio capture error ({e}) — restarting.")
self._capturing = False
await asyncio.sleep(backoff)
backoff = min(backoff * 2, RESTART_BACKOFF_MAX)
async def _run_capture(self) -> None:
cmd = self._ffmpeg_command()
logger.info(f"Starting capture: ffmpeg -f pulse -i {settings.pulse_source} (s16le/{AUDIO_SAMPLE_RATE}/mono)")
self._last_stderr_lines.clear()
proc = await asyncio.create_subprocess_exec(
*cmd,
stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.PIPE,
)
self._proc = proc
stderr_task = asyncio.create_task(self._drain_stderr(proc))
try:
assert proc.stdout is not None
while True:
# readexactly, not read: a fixed chunk keeps every analysis
# window the same length AND guarantees sample alignment, so a
# short read can never split a 16-bit sample across chunks.
try:
chunk = await proc.stdout.readexactly(READ_CHUNK_BYTES)
except asyncio.IncompleteReadError as partial:
if partial.partial:
self._ingest(partial.partial)
break # EOF — FFmpeg died or the source went away
if not self._capturing:
self._capturing = True
logger.info("PulseAudio capture is producing audio.")
self._ingest(chunk)
finally:
self._capturing = False
# _drain_stderr swallows its own CancelledError, so it finishes cleanly
# and never needs awaiting here.
stderr_task.cancel()
await self._terminate_proc()
async def _drain_stderr(self, proc: asyncio.subprocess.Process) -> None:
"""Surface FFmpeg's diagnostics instead of letting the pipe fill and block."""
if proc.stderr is None:
return
try:
while True:
line = await proc.stderr.readline()
if not line:
return
text = line.decode(errors="replace").strip()
self._last_stderr_lines.append(text)
logger.warning(f"ffmpeg(pulse): {text}")
except asyncio.CancelledError:
return
except Exception:
return
def _log_capture_exit(self) -> None:
"""
Log why the just-finished capture process exited, distinguishing the
two failure modes that matter operationally instead of one generic
"restarting" line for both:
- no daemon / connection refused: infrastructure isn't up yet. This
is expected during startup/op25 restarts, so it stays at INFO —
the retry loop above already handles it.
- daemon up but the named source is missing: almost always a real
PULSE_SOURCE misconfiguration. This gets a loud, distinct ERROR
naming the configured source, because silently retrying forever
against a wrong source name is exactly how this hid in the past.
"""
text = " ".join(self._last_stderr_lines).lower()
if any(marker in text for marker in _SOURCE_MISSING_MARKERS):
logger.error(
f"PulseAudio capture exited: source '{settings.pulse_source}' does not exist on the "
"daemon (FFmpeg reported 'No such process'). This looks like a real PULSE_SOURCE "
"misconfiguration or a missing drb_sink — retrying will not fix it by itself. "
"Restarting anyway."
)
return
if any(marker in text for marker in _NO_DAEMON_MARKERS):
logger.info(
"PulseAudio capture exited: daemon not accepting connections — "
"infrastructure still coming up, restarting."
)
return
logger.warning("PulseAudio capture process exited — restarting.")
async def _terminate_proc(self) -> None:
proc, self._proc = self._proc, None
if proc is None or proc.returncode is not None:
return
try:
# Synchronous, so the signal lands even if we are being cancelled and
# the reap below never gets to run.
proc.terminate()
except Exception:
return
try:
await asyncio.wait_for(proc.wait(), timeout=5)
except asyncio.TimeoutError:
try:
proc.kill()
except Exception:
pass
except asyncio.CancelledError:
raise
except Exception:
pass
# ------------------------------------------------------------------
# Voice activity
# ------------------------------------------------------------------
def _note_voice(self, chunk: bytes, now: float) -> None:
"""
Update the continuous voice-activity marks from one arriving chunk.
A chunk carrying signal starts a NEW run whenever the gap since the last
one is at least the configured silence timeout — i.e. runs are separated
by exactly the same threshold that ends a recording, so the segmenter's
"did a new run begin" and "did the recording end" questions can never
disagree with each other.
"""
if pcm.is_silent(chunk, settings.call_silence_threshold_db):
return
gap = settings.call_silence_timeout
if self._last_voice_epoch is None or (now - self._last_voice_epoch) >= gap:
self._voice_onset_epoch = now
self._last_voice_epoch = now
def audio_activity(self) -> AudioActivity:
"""Snapshot of the capture stream for metadata_watcher (and /api/status)."""
last = self._last_voice_epoch
silence = (time.time() - last) if last is not None else 0.0
return AudioActivity(
capturing=self._capturing,
recording=self._active is not None,
last_voice_epoch=last,
voice_onset_epoch=self._voice_onset_epoch,
silence_seconds=max(0.0, silence),
)
def _ingest(self, chunk: bytes) -> None:
"""Append a chunk and trim stale data from the front of the buffer."""
now = time.monotonic()
"""Append a chunk to the ring buffer and, if recording, the accumulator."""
now = time.time()
self._note_voice(chunk, now)
self._buffer.append((now, chunk))
self._buffer_bytes += len(chunk)
# During a call, never trim data newer than (call_start - pre_buffer).
# Between calls, keep a rolling RING_BUFFER_SECONDS window.
if self._call_start_mono is not None:
keep_from = self._call_start_mono - PRE_BUFFER_SECONDS
else:
keep_from = now - RING_BUFFER_SECONDS
while self._buffer:
ts, old = self._buffer[0]
if ts >= keep_from:
break
self._buffer.popleft()
# The ring buffer serves pre-roll only, so it is trimmed to a fixed
# window unconditionally — an open recording no longer pins it, because
# the accumulator owns that audio.
keep_from = now - RING_BUFFER_SECONDS
while self._buffer and self._buffer[0][0] < keep_from:
_, old = self._buffer.popleft()
self._buffer_bytes -= len(old)
active = self._active
if active is None:
return
if active.total_bytes + len(chunk) > MAX_RECORDING_BYTES:
if not active.truncated_by_cap:
active.truncated_by_cap = True
logger.warning(
f"Recording {active.call_id} hit the {MAX_RECORDING_BYTES} byte memory ceiling "
f"after {now - active.slice_start:.1f}s — further audio is being dropped."
)
return
active.chunks.append((now, chunk))
active.total_bytes += len(chunk)
# ------------------------------------------------------------------
# Call recording API (same interface as before)
# Recording API
# ------------------------------------------------------------------
async def start_recording(self, call_id: str) -> bool:
if self._call_id:
logger.warning("Recording already active — ignoring start.")
async def start_recording(self, call_id: str, start_epoch: Optional[float] = None) -> bool:
"""
Open a recording. `start_epoch` is the detected voice onset (host wall
clock, same clock as the chunk stamps); the slice begins
PRE_ROLL_SECONDS before it. Omit it only when no onset is available —
then we fall back to "now", losing the pre-roll's precision.
"""
if self._active is not None:
logger.warning(f"Recording already active ({self._active.call_id}) — ignoring start for {call_id}.")
return False
self._call_id = call_id
self._call_start_mono = time.monotonic()
logger.info(f"Recording started (ring-buffer): {call_id}")
call_start = start_epoch if start_epoch else time.time()
slice_start = call_start - PRE_ROLL_SECONDS
if not self._capturing:
logger.warning(f"Recording {call_id} opened while PulseAudio capture is down — audio may be missing.")
# Seed the accumulator with the pre-roll already sitting in the ring
# buffer. No await between reading the buffer and publishing _active, so
# the capture task cannot slip a chunk in between and double-count it.
clamped = 0.0
oldest = self._buffer[0][0] if self._buffer else None
if oldest is not None and slice_start < oldest:
# Pre-roll predates the buffer: node just started, capture restarted,
# or the onset is far in the past. Clamp and say so LOUDLY — this is
# silent audio loss otherwise.
clamped = oldest - slice_start
logger.warning(
f"BUFFER CLAMP: pre-roll for call {call_id} predates buffered audio by "
f"{clamped:.2f}s — recording starts at the buffer head and that audio is lost."
)
seeded = [(ts, chunk) for ts, chunk in self._buffer if ts >= slice_start]
self._active = _ActiveRecording(
call_id=call_id,
call_start=call_start,
slice_start=slice_start,
chunks=seeded,
total_bytes=sum(len(c) for _, c in seeded),
clamped_seconds=clamped,
)
logger.info(f"Recording started: {call_id} (slice from {slice_start:.3f})")
return True
async def stop_recording(self) -> Optional[Path]:
if not self._call_id:
async def discard_recording(self) -> None:
"""
Drop the open recording without writing anything.
Used when the segmenter decides the audio must not be kept — today only
the unattributed/orphan-audio path, where uploading would inject a call
with no talkgroup into correlation.
"""
active, self._active = self._active, None
if active is not None:
logger.info(f"Discarded buffered audio for {active.call_id} ({active.total_bytes} bytes).")
async def stop_recording(self, end_epoch: Optional[float] = None) -> Optional[Recording]:
"""
Close the recording, trim it, encode it once and write the file.
`end_epoch` is host wall clock.
Waits (bounded) for captured audio to actually cover `end_epoch` before
slicing — see TAIL_WAIT_TIMEOUT_SECONDS. An audio-driven close never
needs that wait, because its end epoch is derived from audio that is
already buffered; a control-channel-derived close (tgid_change) does.
Returns None only when there was no recording open or no audio at all.
"""
active = self._active
if active is None:
return None
call_id = self._call_id
call_start = self._call_start_mono
self._call_id = None
self._call_start_mono = None
call_id = active.call_id
slice_start = active.slice_start
# Slice: everything from (call_start - pre_buffer) to now
cutoff = (call_start - PRE_BUFFER_SECONDS) if call_start else 0.0
chunks = [chunk for ts, chunk in self._buffer if ts >= cutoff]
end = end_epoch if end_epoch else time.time()
end = min(end, active.call_start + MAX_RECORDING_SECONDS)
# Safety cap: if the call ran very long, truncate to MAX_RECORDING_SECONDS
if call_start is not None:
cap_cutoff = call_start + MAX_RECORDING_SECONDS
now = time.monotonic()
if now > cap_cutoff:
# Approximate: trim chunks that arrived after the cap
cap_keep_until = call_start + MAX_RECORDING_SECONDS
chunks = [
chunk for ts, chunk in self._buffer
if cutoff <= ts <= cap_keep_until
]
# The accumulator keeps filling during this wait — that is the point.
await self._await_tail(end, call_id)
self._active = None
if not chunks:
parts: List[bytes] = []
total = 0
last_ts = slice_start
for ts, chunk in active.chunks:
if ts < slice_start:
continue
parts.append(chunk)
total += len(chunk)
last_ts = ts
if ts >= end:
# Include the chunk straddling `end` so the tail is never clipped,
# then stop.
break
if not parts:
logger.warning(
f"No buffered audio for call {call_id} — "
"stream may not have been connected yet."
f"No buffered audio for call {call_id} "
f"(window {slice_start:.3f}–{end:.3f}) — PulseAudio capture may be down."
)
return None
if last_ts < end - TAIL_WAIT_POLL_SECONDS:
logger.warning(
f"BUFFER CLAMP: call {call_id} ends {end - last_ts:.2f}s after the newest captured "
"audio — the tail is short. Capture may be stalled or restarting."
)
raw = b"".join(parts)
recording = Recording(
call_id=call_id,
path=None,
audio_start_epoch=slice_start,
audio_end_epoch=slice_start + pcm.seconds(len(raw)),
clamped_seconds=active.clamped_seconds,
)
return await self._finish(recording, raw)
async def _finish(self, recording: Recording, raw: bytes) -> Optional[Recording]:
"""Trim, encode exactly once, and write the MP3."""
audio = raw
if settings.trim_silence:
audio, result = audio_trim.trim_pcm(audio)
if result.all_silence:
logger.warning(
f"Call {recording.call_id} contains no speech at all — skipping upload. "
"Check squelch, the Liquidsoap output and the drb_sink monitor."
)
recording.all_silence = True
return recording
if result.applied:
recording.lead_trimmed = result.lead
recording.tail_trimmed = result.tail
# Wall clock of the trimmed audio's first and last sample.
recording.audio_start_epoch += result.lead
recording.audio_end_epoch -= result.tail
self._recordings_dir.mkdir(parents=True, exist_ok=True)
ts_str = datetime.now(timezone.utc).strftime("%Y%m%d_%H%M%S")
output_path = self._recordings_dir / f"{ts_str}_{call_id}.mp3"
output_path = self._recordings_dir / f"{ts_str}_{recording.call_id}{AUDIO_SUFFIX}"
data = b"".join(chunks)
output_path.write_bytes(data)
if not await encode_recording(audio, output_path):
output_path.unlink(missing_ok=True)
return None
size = output_path.stat().st_size
if size > 0:
logger.info(f"Recording saved: {output_path.name} ({size} bytes)")
return output_path
size = output_path.stat().st_size if output_path.exists() else 0
if size <= 0:
output_path.unlink(missing_ok=True)
logger.warning(f"Recording for call {recording.call_id} produced an empty file.")
return None
output_path.unlink(missing_ok=True)
logger.warning(f"Recording for call {call_id} produced an empty file.")
return None
recording.path = output_path
logger.info(
f"Recording saved: {output_path.name} ({size} bytes, "
f"{pcm.seconds(len(audio)):.2f}s audio)"
)
return recording
async def _await_tail(self, end: float, call_id: str) -> float:
"""
Block until captured audio reaches `end`, or the bounded timeout expires.
Returns seconds waited. Logs whenever a wait was actually needed so the
real pipeline latency is observable in the field.
"""
if not self._buffer:
return 0.0
if self._buffer[-1][0] >= end:
return 0.0
started = time.monotonic()
deadline = started + TAIL_WAIT_TIMEOUT_SECONDS
while time.monotonic() < deadline:
await asyncio.sleep(TAIL_WAIT_POLL_SECONDS)
if self._buffer and self._buffer[-1][0] >= end:
waited = time.monotonic() - started
logger.info(f"Waited {waited:.2f}s for the tail of call {call_id} to reach the buffer.")
return waited
if not self._capturing:
break # capture died mid-wait; nothing more is coming
waited = time.monotonic() - started
newest = self._buffer[-1][0] if self._buffer else end
logger.warning(
f"Tail wait for call {call_id} gave up after {waited:.2f}s — captured audio is still "
f"{max(0.0, end - newest):.2f}s short of the call end. Tail may be clipped."
)
return waited
# ------------------------------------------------------------------
# Upload (unchanged interface)
@@ -174,6 +758,8 @@ class CallRecorder:
talkgroup_id: Optional[int] = None,
talkgroup_name: Optional[str] = None,
system_id: Optional[str] = None,
audio_start_epoch: Optional[float] = None,
audio_end_epoch: Optional[float] = None,
) -> Optional[str]:
if not settings.c2_url:
logger.info("No C2_URL configured — skipping upload.")
@@ -190,13 +776,20 @@ class CallRecorder:
form["talkgroup_name"] = talkgroup_name
if system_id:
form["system_id"] = system_id
# Where this audio really sits on the wall clock once silence is trimmed.
# C2 does not declare these Form fields yet, so FastAPI ignores them —
# they cost nothing and are here for when playback/correlation want them.
if audio_start_epoch is not None:
form["audio_start_epoch"] = f"{audio_start_epoch:.3f}"
if audio_end_epoch is not None:
form["audio_end_epoch"] = f"{audio_end_epoch:.3f}"
try:
async with httpx.AsyncClient(timeout=120) as client:
with open(file_path, "rb") as f:
r = await client.post(
upload_url,
files={"file": (file_path.name, f, "audio/mpeg")},
files={"file": (file_path.name, f, AUDIO_MIME)},
data=form,
headers=headers,
)
@@ -213,9 +806,24 @@ class CallRecorder:
except Exception:
pass
# ------------------------------------------------------------------
# State
# ------------------------------------------------------------------
@property
def is_recording(self) -> bool:
return self._call_id is not None
return self._active is not None
@property
def is_capturing(self) -> bool:
"""True when FFmpeg is alive and audio is actually arriving."""
return self._capturing
@property
def buffered_seconds(self) -> float:
if len(self._buffer) < 2:
return 0.0
return self._buffer[-1][0] - self._buffer[0][0]
call_recorder = CallRecorder()
+57 -5
View File
@@ -1,30 +1,73 @@
"""
Manages the persisted node API key.
Manages the persisted node API key, plus the local-auth signing material used
by app/internal/auth.py.
The key is provisioned by the C2 server after an admin approves the node.
The API key is provisioned by the C2 server after an admin approves the node.
It arrives via MQTT and is saved to /configs/credentials.json so it survives
container restarts.
The scrypt salt and session-signing secret are generated locally on first boot
(never provisioned externally) and persisted the same way, so dashboard
sessions survive a container restart instead of forcing every operator to
re-login whenever the node restarts.
"""
import json
import secrets
from pathlib import Path
from app.config import settings
from app.internal.logger import logger
_CREDS_FILE = Path(settings.config_path) / "credentials.json"
_api_key: str | None = None
_auth_salt: bytes | None = None
_session_secret: bytes | None = None
def load() -> None:
"""Load persisted credentials from disk on startup."""
global _api_key
global _api_key, _auth_salt, _session_secret
if _CREDS_FILE.exists():
try:
data = json.loads(_CREDS_FILE.read_text())
_api_key = data.get("api_key")
if data.get("auth_salt"):
_auth_salt = bytes.fromhex(data["auth_salt"])
if data.get("session_secret"):
_session_secret = bytes.fromhex(data["session_secret"])
if _api_key:
logger.info("Node credentials loaded from disk.")
except Exception as e:
logger.warning(f"Could not read credentials file: {e}")
_ensure_auth_material()
def _ensure_auth_material() -> None:
"""Generate (once) and persist the local-auth salt + session secret."""
global _auth_salt, _session_secret
changed = False
if _auth_salt is None:
_auth_salt = secrets.token_bytes(16)
changed = True
if _session_secret is None:
_session_secret = secrets.token_bytes(32)
changed = True
if changed:
_write()
logger.info("Generated local-auth signing material (first boot).")
def get_auth_salt() -> bytes:
"""Scrypt salt for dashboard password hashing — generated once, persisted."""
if _auth_salt is None:
_ensure_auth_material()
return _auth_salt # type: ignore[return-value]
def get_session_secret() -> bytes:
"""HMAC key used to sign dashboard session cookies — generated once, persisted."""
if _session_secret is None:
_ensure_auth_material()
return _session_secret # type: ignore[return-value]
def get_api_key() -> str | None:
@@ -34,6 +77,15 @@ def get_api_key() -> str | None:
def save_api_key(key: str) -> None:
global _api_key
_api_key = key
_CREDS_FILE.parent.mkdir(parents=True, exist_ok=True)
_CREDS_FILE.write_text(json.dumps({"api_key": key}))
_write()
logger.info("Node API key saved to disk.")
def _write() -> None:
_CREDS_FILE.parent.mkdir(parents=True, exist_ok=True)
data: dict = {"api_key": _api_key}
if _auth_salt is not None:
data["auth_salt"] = _auth_salt.hex()
if _session_secret is not None:
data["session_secret"] = _session_secret.hex()
_CREDS_FILE.write_text(json.dumps(data))
+48 -11
View File
@@ -2,11 +2,14 @@ import asyncio
from typing import Optional
import discord
from discord.ext import commands
from app.config import settings
from app.internal import pulse
from app.internal.logger import logger
BOT_READY_TIMEOUT = 15 # seconds to wait for Discord bot to become ready
WATCHDOG_INTERVAL = 30 # seconds between voice-connection health checks
REJOIN_DELAY = 5 # seconds to wait before attempting a rejoin
STREAM_RETRY_DELAY = 5 # seconds to back off before re-arming the audio source
class RadioBot:
@@ -24,7 +27,8 @@ class RadioBot:
self._channel_id: Optional[int] = None
self._was_streaming: bool = False
async def join(self, guild_id: int, channel_id: int, token: str, call_active: bool = False, system_name: str = None) -> bool:
async def join(self, guild_id: int, channel_id: int, token: str,
call_active: bool = False, system_name: str = None) -> bool:
# (Re)start the bot if the token changed or the bot isn't running
if self._current_token != token or not self._is_bot_running():
if not await self._start_bot(token):
@@ -47,6 +51,10 @@ class RadioBot:
# Remember where we are so the watchdog can rejoin if we drop
self._guild_id = guild_id
self._channel_id = channel_id
# Bounded wait for the shared PulseAudio socket. Historically FFmpeg was
# launched before the op25 container had created it, failed instantly,
# and the bot sat silently connected forever.
await pulse.wait_until_ready()
self._play_stream()
if system_name:
await self._bot.change_presence(
@@ -108,19 +116,48 @@ class RadioBot:
self._ready_event = None
def _play_stream(self):
"""
Feed Discord voice straight from the PulseAudio monitor.
Icecast is NOT used here: it lags ~1 s at connect and drifts to 100 s+,
which is unusable for live listening. Icecast remains the frontend/mobile
listening path only.
"""
if not self._voice_client:
return
from app.config import settings
stream_url = f"http://{settings.icecast_host}:{settings.icecast_port}{settings.icecast_mount}"
if not pulse.is_ready():
logger.error(
f"PulseAudio socket {pulse.socket_path()} missing — "
f"cannot start Discord audio; retrying in {STREAM_RETRY_DELAY}s."
)
self._schedule_restart()
return
# before_options land ahead of -i, so this becomes:
# ffmpeg -f pulse -i drb_sink.monitor …
source = discord.FFmpegPCMAudio(
stream_url,
before_options="-reconnect 1 -reconnect_streamed 1 -reconnect_delay_max 5",
settings.pulse_source,
before_options="-f pulse",
)
self._voice_client.play(
discord.PCMVolumeTransformer(source, volume=1.0),
after=self._on_stream_end,
)
def _schedule_restart(self, delay: float = STREAM_RETRY_DELAY):
"""Re-arm the audio source after a delay — safe to call from any thread."""
if not self._loop:
return
async def _delayed_restart():
await asyncio.sleep(delay)
vc = self._voice_client
if vc and vc.is_connected() and not vc.is_playing():
self._play_stream()
self._loop.call_soon_threadsafe(lambda: asyncio.ensure_future(_delayed_restart()))
def _on_stream_end(self, error):
if error:
logger.error(f"Stream ended with error: {error}")
@@ -128,12 +165,9 @@ class RadioBot:
if not (self._loop and vc and vc.is_connected() and not vc.is_playing()):
return
if error:
# Back off before retrying — prevents tight loop when PulseAudio is unavailable
async def _delayed_restart():
await asyncio.sleep(5)
if self._voice_client and self._voice_client.is_connected() and not self._voice_client.is_playing():
self._play_stream()
self._loop.call_soon_threadsafe(lambda: asyncio.ensure_future(_delayed_restart()))
# Back off before retrying — prevents a tight loop when PulseAudio is
# unavailable (FFmpeg exits immediately in that case).
self._schedule_restart()
else:
self._loop.call_soon_threadsafe(self._play_stream)
@@ -232,6 +266,9 @@ class RadioBot:
else:
self._voice_client = await vc.connect()
self._channel_id = vc.id
# A fresh connect() has no audio source attached yet.
if not self._voice_client.is_playing():
self._play_stream()
await message.reply(f"Joined {vc.name}.")
except Exception as e:
logger.error(f"joinme failed: {e}")
+8
View File
@@ -7,4 +7,12 @@ logging.basicConfig(
handlers=[logging.StreamHandler(sys.stdout)],
)
# The metadata watcher polls the OP25 terminal twice a second and httpx logs
# every one of those requests at INFO ("HTTP Request: POST http://... 200 OK").
# That is ~170k lines/day of pure noise which buries real events and makes field
# log-reading useless. WARNING keeps genuine transport failures visible.
# httpcore is the transport layer underneath httpx and is just as chatty.
for _noisy in ("httpx", "httpcore"):
logging.getLogger(_noisy).setLevel(logging.WARNING)
logger = logging.getLogger("drb-edge-node")
+826 -54
View File
@@ -1,38 +1,274 @@
"""
Call segmentation: AUDIO decides the boundaries, the CONSOLE decides the label.
START first chunk of audio above the silence threshold (voice onset).
STOP settings.call_silence_timeout seconds of continuous silence HEARD in
that audio.
LABEL talkgroup / alias / rid, resolved from OP25 console observations that
fall inside the recording's window, resolved AT CLOSE TIME.
SPLIT a console talkgroup change still forces a cut, even mid-audio.
WHY THE CONTROL CHANNEL NO LONGER DECIDES BOUNDARIES. The previous design
started a segment on an OP25 `call_log` grant and ended it by inferring from the
control channel: the `srcaddr != 0 -> 0` edge started an idle timer and the
segment closed call_idle_timeout seconds later. Both halves were measured wrong
in the field:
* The grant fires 0.84-1.62 s (variable) before anyone speaks, so a
grant-anchored window is always guessing at the offset.
* `srcaddr` can drop to 0 WHILE SOMEONE IS STILL TALKING. Measured across six
recordings, five had healthy trailing silence trimmed (-0.53 s to -2.48 s)
but one reported "-1.61s lead, -0.00s tail" — the trim found nothing to
remove because the capture window had closed on top of live speech. The
recording ends on an unfinished word. Working backwards from its lead trim,
the audio pipeline lag was at most 1.36 s, so the window should have held
~1.6 s more; the only consistent explanation is a false early `srcaddr -> 0`.
Audio is the ground truth for WHEN. It cannot answer WHO, so the console is
still the only source of talkgroup, alias and radio id.
WHY ATTRIBUTION HAPPENS AT CLOSE, NOT AT OPEN. There is no guaranteed ordering
between a grant and the audio it belongs to: the console is polled every 500 ms
and the audio pipeline lag is variable, so the grant can land after voice onset
just as easily as before it. A segment may therefore open unattributed and
acquire its talkgroup part-way through, which is expected and fine. At close we
have seen the whole window and ask the rolling console history "what was active
during this audio, give or take a few seconds" — see _attribute and the
ATTRIBUTION_* constants.
ORPHAN AUDIO. If nothing in the console history overlaps the window, the audio
is unattributed: Liquidsoap fallback, a test tone, stray noise, or a dropped
`call_log`. Policy is DISCARD AND SHOUT — the recording is not uploaded and no
call_start/call_end is published, because a call with no talkgroup silently
poisons incident correlation downstream, and that is worse than losing the
audio. It is logged at ERROR with the window and everything nearby that was
considered, and counted on /api/status so it cannot pass unnoticed.
FALLBACK MODE. When PulseAudio capture is NOT producing audio there is nothing
to segment on, so the old console state machine still runs (grant opens,
srcaddr edge + call_idle_timeout closes). It produces no audio — capture is
down — but it keeps the node reporting real radio activity to C2 while the
audio path is broken. This is the only remaining consumer of
settings.call_idle_timeout.
SEGMENTS: one emitted call (= one recording, one Firestore doc) spans a whole
conversation, not a single transmission. It stays open across repeated grants on
the same talkgroup and closes when the talkgroup changes or the AUDIO goes quiet
for settings.call_silence_timeout seconds.
CLOCKS: `call_log["time"]` is time.time() inside the op25 container. All three
client containers run network_mode: host and share the host kernel clock, so that
value is directly comparable to time.time() here — no offset mapping needed. The
call recorder's chunk timestamps are the same clock, with the caveat that they
are ARRIVAL times and therefore lag the moment of speech by the pipeline
latency. Comparisons between two audio timestamps are exact; comparisons between
audio and console timestamps carry that lag, which is what _tail_pad() covers.
"""
import asyncio
import time
import uuid
from collections import deque
from dataclasses import dataclass, field
from datetime import datetime, timezone
from typing import Optional, Callable, Awaitable
from typing import Optional, Callable, Awaitable, Any, List, Dict
from app.config import settings
from app.internal.call_recorder import AudioActivity
from app.internal.op25_client import op25_client
from app.internal.logger import logger
CallbackFn = Callable[[dict], Awaitable[None]]
ActivityFn = Callable[[], AudioActivity]
HANG_THRESHOLD = 2 # polls before declaring a call ended (0.5s poll → 1s hang time)
POLL_INTERVAL = 0.5 # seconds
# 500 ms. Do NOT lower: audio boundaries come from the recorder's own chunk
# timestamps (~46 ms resolution), not from when this loop happens to notice
# them, and http_server.py's request handler has a ~200 ms blocking floor anyway.
POLL_INTERVAL = 0.5
# Seconds of unreachable OP25 before an open segment is force-closed. Applies in
# both modes: without the console there is no attribution, and unattributed
# audio is discarded anyway.
OP25_OFFLINE_GRACE = 3.0
# Hard ceiling on a single segment; mirrors MAX_RECORDING_SECONDS in call_recorder
# so a talkgroup that never goes quiet cannot produce an unbounded recording. In
# audio mode a new segment is opened immediately afterwards if voice is still
# present, so a genuinely long transmission is split rather than truncated.
MAX_SEGMENT_SECONDS = 600
# How far either side of the AUDIO window console observations are still
# accepted as attribution evidence. "Plus or minus some seconds", made explicit:
#
# LOOKBACK the grant normally PRECEDES the audio — 0.84-1.62 s of
# grant-to-speech delay, plus up to ~1.4 s of audio pipeline lag,
# plus one 0.5 s poll of detection slack. 4.0 s covers the worst
# case measured with margin.
# LOOKAHEAD the grant can also FOLLOW voice onset, because the console is only
# polled every 500 ms and OP25 logs the grant on its own schedule.
# 2.0 s is four poll intervals.
#
# Both are deliberately asymmetric: the "grant first" direction is the common
# one and has the larger physical spread.
ATTRIBUTION_LOOKBACK_SECONDS = 4.0
ATTRIBUTION_LOOKAHEAD_SECONDS = 2.0
# Rolling console history. Bounded twice — by age and by entry count — so a busy
# system cannot grow it without limit. At ~2 observations per poll this is a few
# minutes of history for a few tens of KB.
CONSOLE_HISTORY_SECONDS = 180.0
CONSOLE_HISTORY_MAX = 1200
# How far back to look for a duplicate before appending a grant. OP25's call_log
# deque drains on read so repeats should not happen, but a re-delivered entry
# would otherwise inflate the transmission count and the attribution score.
_GRANT_DEDUPE_DEPTH = 24
# Close reasons where the console explicitly told us the talkgroup changed, so
# the segment's label is already known first-hand and close-time attribution
# would only be able to make it worse (the window extends past the split).
_SPLIT_REASONS = ("tgid_change", "tgid_change_unlogged")
def _tail_pad() -> float:
"""
Audio kept past a CONSOLE-DERIVED segment boundary, to cover the fact that
buffered audio lags control-channel timestamps.
Under audio-driven segmentation this no longer applies to the normal end of
a call — that boundary now comes from the audio itself and needs no pad. It
still applies wherever a boundary is a control-channel timestamp:
tgid_change close at the new grant's timestamp + pad
tgid_change_unlogged close at the observing poll's timestamp + pad
idle_timeout console fallback mode only
Read live from settings (env CALL_TAIL_PAD_SECONDS) rather than frozen into
a module constant, so it is tunable per node.
An earlier version of this docstring claimed the tgid_change paths close at
"an exact, already-known boundary" and so intentionally added no pad — THAT
REASONING WAS WRONG and produced real truncated recordings. The boundary is
exact only in CONTROL-CHANNEL time; the buffered AUDIO lags control-channel
timestamps by ~1.5 s (measured: 0.84-1.62 s of lead trimmed across 7 field
calls), so slicing the outgoing call at the new grant's exact timestamp cut
roughly the last 1.5 s of its real speech. Do not reintroduce a zero-pad
close for tgid_change or tgid_change_unlogged; if the outgoing and incoming
recordings end up overlapping in the underlying audio because of this pad,
that is correct — the audio genuinely contains both.
"""
return settings.call_tail_pad_seconds
def _as_int(value: Any) -> Optional[int]:
"""Coerce an OP25 field to a positive int, or None. Rejects 0/""/"None"."""
if value is None:
return None
try:
number = int(value)
except (TypeError, ValueError):
return None
return number if number > 0 else None
def _as_float(value: Any) -> Optional[float]:
try:
return float(value)
except (TypeError, ValueError):
return None
def _iso(epoch: Optional[float]) -> Optional[str]:
if epoch is None:
return None
return datetime.fromtimestamp(epoch, timezone.utc).isoformat()
@dataclass(frozen=True)
class ConsoleEvent:
"""One thing the OP25 console said, kept so a closing segment can ask about it."""
epoch: float
tgid: int
name: str = ""
freq: Any = None
rid: Optional[int] = None
# True for a `call_log` grant, False for an active `channel_update` row.
is_grant: bool = False
@dataclass
class Attribution:
"""Who a stretch of audio belonged to, and how confident we are."""
tgid: int
name: str = ""
freq: Any = None
rid: Optional[int] = None
grants: int = 0
# Observations that fall strictly inside the audio window (vs only inside
# the tolerance band around it).
overlap: int = 0
nearby: int = 0
competing: List[int] = field(default_factory=list)
class MetadataWatcher:
def __init__(self):
self._running = False
# Open segment state
self._active_call_id: Optional[str] = None
self._current_tgid: Optional[int] = None
self._current_tgid_name: Optional[str] = None
self._hang_counter: int = 0
self._active_call_id: Optional[str] = None
self._call_started_at: Optional[datetime] = None
self._current_freq: Any = None
self._current_srcaddr: Optional[int] = None
self._started_at: Optional[float] = None # audio onset, or grant epoch in fallback mode
self._transmissions: int = 0
# True when the open segment is governed by audio, False for the
# console fallback. Fixed at open so capture flapping cannot switch the
# rules underneath a live segment.
self._audio_driven: bool = False
# Transmission tracking within the open segment (console fallback mode)
self._tx_active: bool = False # last poll saw srcaddr != 0
self._last_activity: float = 0.0 # epoch of last evidence of traffic
self._last_tx_end: Optional[float] = None # epoch of the srcaddr 1→0 edge
self._last_ok_poll: float = 0.0
# Rolling console history for close-time attribution.
self._console: deque[ConsoleEvent] = deque(maxlen=CONSOLE_HISTORY_MAX)
# Onset of the voice run the last audio-driven segment covered, so the
# same run cannot immediately re-open a second segment.
self._consumed_onset: Optional[float] = None
# Field-visible counter of discarded orphan audio.
self._unattributed_segments: int = 0
# Injectable for tests; production is always the host wall clock.
self._clock: Callable[[], float] = time.time
# Set these before calling start()
self.on_call_start: Optional[CallbackFn] = None
self.on_call_end: Optional[CallbackFn] = None
# Supplies the audio-activity snapshot. None (or a snapshot reporting
# capturing=False) puts the watcher in console fallback mode.
self.audio_activity: Optional[ActivityFn] = None
# ------------------------------------------------------------------
# Lifecycle
# ------------------------------------------------------------------
async def start(self):
self._running = True
self._last_ok_poll = self._clock()
asyncio.create_task(self._poll_loop())
logger.info("Metadata watcher started.")
logger.info("Metadata watcher started (audio-driven segmentation, console attribution).")
async def stop(self):
self._running = False
if self._active_call_id:
await self._end_call()
await self._close_segment(self._clock(), reason="shutdown")
async def _poll_loop(self):
while self._running:
@@ -42,79 +278,610 @@ class MetadataWatcher:
logger.warning(f"Metadata poll error: {e}")
await asyncio.sleep(POLL_INTERVAL)
async def _tick(self):
status = await op25_client.get_terminal_status()
# ------------------------------------------------------------------
# One poll
# ------------------------------------------------------------------
if not status:
# OP25 not responding — hang-out any active call
if self._active_call_id:
self._hang_counter += 1
if self._hang_counter >= HANG_THRESHOLD:
await self._end_call()
async def _tick(self):
now = self._clock()
update = await op25_client.poll_terminal()
if update is None:
# OP25 unreachable. Don't kill an open segment on a single blip.
if self._active_call_id and (now - self._last_ok_poll) >= OP25_OFFLINE_GRACE:
await self._close_segment(now, reason="op25_unreachable")
return
# OP25 terminal returns either a list of channels or a single dict
channels = status if isinstance(status, list) else [status]
active_tgid: Optional[int] = None
active_meta: dict = {}
self._last_ok_poll = now
self._record_console(update, now)
for ch in channels:
tgid = ch.get("tgid") or ch.get("tg_id")
if tgid and str(tgid) not in ("0", "", "None"):
active_tgid = int(tgid)
active_meta = ch
break
activity = self._snapshot()
if activity is None or not activity.capturing:
await self._console_tick(update, now)
return
if active_tgid:
self._hang_counter = 0
if self._current_tgid != active_tgid:
# Talkgroup changed — close previous call and open a new one
if self._active_call_id:
await self._end_call()
self._current_tgid = active_tgid
await self._start_call(active_tgid, active_meta)
else:
# No active talkgroup
if self._active_call_id:
self._hang_counter += 1
if self._hang_counter >= HANG_THRESHOLD:
await self._end_call()
await self._audio_tick(update, activity, now)
async def _start_call(self, tgid: int, meta: dict):
def _snapshot(self) -> Optional[AudioActivity]:
if self.audio_activity is None:
return None
try:
return self.audio_activity()
except Exception as e:
logger.warning(f"Audio activity unavailable ({e}) — falling back to console segmentation.")
return None
# ------------------------------------------------------------------
# Console history (feeds close-time attribution)
# ------------------------------------------------------------------
def _record_console(self, update: Any, now: float) -> None:
for entry in update.call_log:
tgid = _as_int(entry.get("tgid"))
if tgid is None:
continue
epoch = _as_float(entry.get("time"))
event = ConsoleEvent(
epoch=now if epoch is None else epoch,
tgid=tgid,
name=entry.get("tgtag") or "",
freq=entry.get("freq"),
rid=_as_int(entry.get("rid")),
is_grant=True,
)
if not self._is_duplicate_grant(event):
self._console.append(event)
for channel in update.channels:
tgid = _as_int(channel.get("tgid"))
srcaddr = _as_int(channel.get("srcaddr"))
if tgid is None or srcaddr is None:
continue # idle channel says nothing about who is talking
self._console.append(ConsoleEvent(
epoch=now,
tgid=tgid,
name=channel.get("tag") or "",
freq=channel.get("freq"),
rid=srcaddr,
is_grant=False,
))
cutoff = now - CONSOLE_HISTORY_SECONDS
while self._console and self._console[0].epoch < cutoff:
self._console.popleft()
def _is_duplicate_grant(self, event: ConsoleEvent) -> bool:
for index in range(len(self._console) - 1, -1, -1):
if len(self._console) - index > _GRANT_DEDUPE_DEPTH:
return False
known = self._console[index]
if known.is_grant and known.tgid == event.tgid and known.epoch == event.epoch:
return True
return False
def _attribute(self, start: float, end: float) -> Optional[Attribution]:
"""
Resolve which talkgroup a stretch of audio belongs to.
Scores every talkgroup seen in [start - LOOKBACK, end + LOOKAHEAD] by
how well its console activity overlaps the audio itself, preferring
real overlap over merely being nearby, and grants over channel rows.
Returns None only when NOTHING was observed in that band at all — the
orphan-audio case.
"""
low = start - ATTRIBUTION_LOOKBACK_SECONDS
high = end + ATTRIBUTION_LOOKAHEAD_SECONDS
candidates: Dict[int, Attribution] = {}
firsts: Dict[int, float] = {}
for event in self._console:
if event.epoch < low or event.epoch > high:
continue
found = candidates.get(event.tgid)
if found is None:
found = Attribution(tgid=event.tgid)
candidates[event.tgid] = found
firsts[event.tgid] = event.epoch
found.nearby += 1
if start <= event.epoch <= end:
found.overlap += 1
if event.is_grant:
found.grants += 1
if event.name and not found.name:
found.name = event.name
if event.freq and found.freq is None:
found.freq = event.freq
if event.rid is not None:
found.rid = event.rid # most recent wins
if not candidates:
return None
best = max(
candidates.values(),
key=lambda a: (a.overlap, a.grants, a.nearby, -firsts[a.tgid]),
)
best.competing = sorted(
tgid for tgid, a in candidates.items() if tgid != best.tgid and a.overlap > 0
)
return best
# ------------------------------------------------------------------
# Audio-driven segmentation
# ------------------------------------------------------------------
async def _audio_tick(self, update: Any, activity: AudioActivity, now: float) -> None:
if self._active_call_id is not None and not self._audio_driven:
# A segment that opened while capture was down finishes under the
# rules it started with rather than switching mid-flight.
await self._console_tick(update, now)
return
# 1. Console first: a talkgroup change must still force a split even
# when the audio never went quiet, and a grant may be the thing that
# finally attributes an already-open segment.
for entry in sorted(update.call_log, key=lambda e: _as_float(e.get("time")) or 0.0):
await self._handle_grant(entry, now)
await self._scan_channels(update.channels, now)
# 2. Then the audio decides the boundaries.
last_voice = activity.last_voice_epoch
voice_active = last_voice is not None and (now - last_voice) < settings.call_silence_timeout
if self._active_call_id is None:
onset = activity.voice_onset_epoch
if voice_active and onset is not None and (
self._consumed_onset is None or onset > self._consumed_onset
):
await self._open_from_audio(onset, now)
return
if not voice_active:
silence = (now - last_voice) if last_voice is not None else settings.call_silence_timeout
# The measured trailing silence, in the AUDIO's own clock. This is
# the number to tune settings.call_silence_timeout from — unlike the
# old control-channel idle it contains no grant-to-speech delay, so
# it means exactly what it says.
logger.info(
f"Audio silence close for tgid {self._current_tgid}: measured trailing silence "
f"{silence:.2f}s (threshold {settings.call_silence_timeout:.2f}s at "
f"{settings.call_silence_threshold_db:.1f}dBFS)."
)
self._consumed_onset = activity.voice_onset_epoch
end = (last_voice + settings.call_silence_timeout) if last_voice is not None else now
await self._close_segment(min(end, now), reason="audio_silence")
return
if self._started_at is not None and (now - self._started_at) >= MAX_SEGMENT_SECONDS:
logger.warning(
f"Segment for tgid {self._current_tgid} hit the {MAX_SEGMENT_SECONDS}s cap while audio "
"was still live — closing and immediately reopening so nothing is dropped. If this "
"repeats, the silence threshold may be low enough that noise reads as voice."
)
await self._close_segment(now, reason="max_length")
await self._open_from_audio(now, now)
async def _open_from_audio(self, onset: float, now: float) -> None:
"""Open a segment at a detected voice onset, attributing it if we can."""
found = self._attribute(onset, now)
await self._open_segment(
started_at=onset,
now=now,
tgid=found.tgid if found else None,
tgid_name=found.name if found else "",
freq=found.freq if found else None,
srcaddr=found.rid if found else None,
audio_driven=True,
transmissions=found.grants if found else 0,
)
async def _handle_grant(self, entry: Dict[str, Any], now: float) -> None:
"""A `call_log` grant, interpreted in audio mode: label or split, never start."""
tgid = _as_int(entry.get("tgid"))
if tgid is None:
return # a grant with no talkgroup is nothing we can label with
started_at = _as_float(entry.get("time"))
if started_at is None:
logger.warning(f"call_log entry for tgid={tgid} has no usable time — using local clock.")
started_at = now
if self._active_call_id is None:
# Audio starts recordings, not grants. The grant is already in the
# console history and will attribute the segment when audio arrives.
return
if self._current_tgid is None:
self._current_tgid = tgid
self._current_tgid_name = entry.get("tgtag") or ""
self._current_freq = entry.get("freq")
self._current_srcaddr = _as_int(entry.get("rid"))
self._transmissions += 1
logger.info(
f"Late attribution: segment {self._active_call_id} adopted tgid {tgid} from a grant "
f"logged {started_at - (self._started_at or started_at):+.2f}s from audio onset."
)
return
if tgid == self._current_tgid:
# CONTINUE: same talkgroup, keep one recording so the back-and-forth
# of a single conversation lands in one file.
self._transmissions += 1
self._refresh_meta_from_log(entry)
return
# FORCED SPLIT. Two talkgroups can be back to back with no silence
# between them; pure audio segmentation would merge them into one file
# under one label, which is exactly the kind of wrong that corrupts
# incident correlation. The console change is authoritative here.
await self._close_segment(started_at + _tail_pad(), reason="tgid_change")
await self._open_segment(
started_at=started_at,
now=now,
tgid=tgid,
tgid_name=entry.get("tgtag") or "",
freq=entry.get("freq"),
srcaddr=_as_int(entry.get("rid")),
audio_driven=True,
transmissions=1,
)
async def _scan_channels(self, channels: List[Dict[str, Any]], now: float) -> None:
"""Channel rows in audio mode: refresh metadata, catch an unlogged split."""
if self._active_call_id is None:
return
active: List[Dict[str, Any]] = []
ours = False
for channel in channels:
tgid = _as_int(channel.get("tgid"))
srcaddr = _as_int(channel.get("srcaddr"))
if tgid is None or srcaddr is None:
continue
active.append(channel)
if tgid == self._current_tgid:
ours = True
self._current_srcaddr = srcaddr
self._last_activity = now
self._refresh_meta_from_channel(channel)
if self._current_tgid is None:
# Late attribution from a channel row — this is the path that saves
# us when the grant itself was dropped from OP25's capped deque.
if len(active) == 1:
tgid = _as_int(active[0].get("tgid"))
self._current_tgid = tgid
self._current_tgid_name = active[0].get("tag") or ""
self._current_freq = active[0].get("freq")
self._current_srcaddr = _as_int(active[0].get("srcaddr"))
logger.info(f"Late attribution: segment {self._active_call_id} adopted tgid {tgid} from channel state.")
return
if ours or not active or len(channels) != 1:
# Restricted to single-receiver setups on purpose: with several
# receivers, another channel being busy says nothing about ours.
return
foreign = _as_int(active[0].get("tgid"))
if foreign is None or foreign == self._current_tgid:
return
logger.warning(
f"tgid {foreign} active without a call_log entry — splitting segment for tgid "
f"{self._current_tgid} (call_log event likely dropped)."
)
await self._close_segment(now + _tail_pad(), reason="tgid_change_unlogged")
await self._open_segment(
started_at=now,
now=now,
tgid=foreign,
tgid_name=active[0].get("tag") or "",
freq=active[0].get("freq"),
srcaddr=_as_int(active[0].get("srcaddr")),
audio_driven=True,
transmissions=1,
)
# ------------------------------------------------------------------
# Console fallback segmentation (capture down)
# ------------------------------------------------------------------
async def _console_tick(self, update: Any, now: float) -> None:
if self._active_call_id is not None and self._audio_driven:
logger.warning(
f"PulseAudio capture stopped while recording {self._active_call_id} — closing the "
"segment at the last captured audio; segmentation falls back to the control channel."
)
await self._close_segment(now, reason="capture_lost")
return
# 1. call_log first — these are the authoritative starts, and processing
# them before the channel scan means a same-poll grant+state pair is
# already attributed to the new segment by the time we scan channels.
# Sorted defensively: multi-receiver setups append per receiver.
for entry in sorted(update.call_log, key=lambda e: _as_float(e.get("time")) or 0.0):
await self._handle_call_log(entry, now)
# 2. channel_update — the only external end signal available here.
await self._handle_channels(update.channels, now)
async def _handle_call_log(self, entry: Dict[str, Any], now: float) -> None:
tgid = _as_int(entry.get("tgid"))
if tgid is None:
return
started_at = _as_float(entry.get("time"))
if started_at is None:
logger.warning(f"call_log entry for tgid={tgid} has no usable time — using local clock.")
started_at = now
if self._active_call_id is None:
await self._open_from_console(entry, tgid, started_at, now)
return
if tgid == self._current_tgid:
self._transmissions += 1
self._tx_active = True
self._last_tx_end = None
self._last_activity = now
self._refresh_meta_from_log(entry)
return
await self._close_segment(started_at + _tail_pad(), reason="tgid_change")
await self._open_from_console(entry, tgid, started_at, now)
async def _handle_channels(self, channels: List[Dict[str, Any]], now: float) -> None:
if self._active_call_id is None:
return
tx_active = False
foreign_active_tgid: Optional[int] = None
for channel in channels:
srcaddr = _as_int(channel.get("srcaddr"))
chan_tgid = _as_int(channel.get("tgid"))
if srcaddr is None:
continue
if chan_tgid == self._current_tgid:
tx_active = True
self._current_srcaddr = srcaddr
self._refresh_meta_from_channel(channel)
elif chan_tgid is not None:
foreign_active_tgid = chan_tgid
if tx_active:
self._tx_active = True
self._last_tx_end = None
self._last_activity = now
elif self._tx_active:
# The srcaddr != 0 → 0 edge. Note this is NOT trusted as an end of
# speech any more (it fires mid-word in the field) — in fallback
# mode there is simply nothing better available.
self._tx_active = False
self._last_tx_end = now
self._last_activity = now
if not tx_active and foreign_active_tgid is not None and len(channels) == 1:
logger.warning(
f"tgid {foreign_active_tgid} active without a call_log entry — "
f"closing segment for tgid {self._current_tgid} (call_log event likely dropped)."
)
await self._close_segment(now + _tail_pad(), reason="tgid_change_unlogged")
return
if (now - self._last_activity) >= settings.call_idle_timeout:
if self._last_tx_end is not None:
measured_idle = now - self._last_tx_end
end = self._last_tx_end + _tail_pad()
logger.info(
f"Idle timeout for tgid {self._current_tgid}: measured control-channel idle "
f"{measured_idle:.2f}s (threshold {settings.call_idle_timeout:.2f}s, "
f"tail pad {_tail_pad():.2f}s)."
)
else:
end = now
logger.info(
f"Idle timeout for tgid {self._current_tgid}: no srcaddr end edge observed, "
f"idle {now - self._last_activity:.2f}s measured from last activity."
)
await self._close_segment(min(end, now), reason="idle_timeout")
return
if self._started_at is not None and (now - self._started_at) >= MAX_SEGMENT_SECONDS:
logger.warning(f"Segment for tgid {self._current_tgid} hit the {MAX_SEGMENT_SECONDS}s cap — closing.")
await self._close_segment(now, reason="max_length")
async def _open_from_console(self, entry: Dict[str, Any], tgid: int, started_at: float, now: float) -> None:
await self._open_segment(
started_at=started_at,
now=now,
tgid=tgid,
tgid_name=entry.get("tgtag") or "",
freq=entry.get("freq"),
srcaddr=_as_int(entry.get("rid")),
audio_driven=False,
transmissions=1,
)
# ------------------------------------------------------------------
# Segment open / close
# ------------------------------------------------------------------
def _refresh_meta_from_log(self, entry: Dict[str, Any]) -> None:
if not self._current_tgid_name:
self._current_tgid_name = entry.get("tgtag") or ""
if entry.get("freq"):
self._current_freq = entry.get("freq")
rid = _as_int(entry.get("rid"))
if rid is not None:
self._current_srcaddr = rid
def _refresh_meta_from_channel(self, channel: Dict[str, Any]) -> None:
if not self._current_tgid_name:
self._current_tgid_name = channel.get("tag") or ""
if not self._current_freq and channel.get("freq"):
self._current_freq = channel.get("freq")
async def _open_segment(
self,
started_at: float,
now: float,
tgid: Optional[int],
tgid_name: str,
freq: Any,
srcaddr: Optional[int],
audio_driven: bool,
transmissions: int = 1,
) -> None:
self._active_call_id = str(uuid.uuid4())
self._call_started_at = datetime.now(timezone.utc)
self._current_tgid_name = meta.get("tag") or meta.get("tgid_tag") or ""
self._current_tgid = tgid
self._current_tgid_name = tgid_name
self._current_freq = freq
self._current_srcaddr = srcaddr
self._started_at = started_at
self._transmissions = max(1, transmissions)
self._audio_driven = audio_driven
# Console fallback assumes the transmission is still up; it learns
# otherwise from the next channel scan.
self._tx_active = not audio_driven
self._last_tx_end = None
self._last_activity = now
payload = {
"call_id": self._active_call_id,
"tgid": tgid,
"tgid_name": self._current_tgid_name,
"freq": meta.get("freq"),
"srcaddr": meta.get("srcaddr"),
"started_at": self._call_started_at.isoformat(),
"tgid_name": tgid_name,
"freq": freq,
"srcaddr": srcaddr,
"started_at": _iso(started_at),
# Raw epoch for the recorder's ring-buffer slice — same clock domain.
"started_at_epoch": started_at,
"attributed": tgid is not None,
"driver": "audio" if audio_driven else "console",
}
logger.info(f"Call start: tgid={tgid} id={self._active_call_id}")
source = "audio onset" if audio_driven else "op25 grant"
logger.info(
f"Call start: tgid={tgid} id={self._active_call_id} "
f"({source} t={started_at:.3f}, detected {now - started_at:+.2f}s later)"
)
if self.on_call_start:
await self.on_call_start(payload)
async def _end_call(self):
async def _close_segment(self, end_epoch: float, reason: str) -> None:
if not self._active_call_id:
return
started_at = self._started_at
if started_at is not None:
end_epoch = max(end_epoch, started_at)
if self._audio_driven and reason not in _SPLIT_REASONS:
self._resolve_attribution(started_at if started_at is not None else end_epoch, end_epoch)
attributed = self._current_tgid is not None
payload = {
"call_id": self._active_call_id,
"tgid": self._current_tgid,
"tgid_name": self._current_tgid_name or "",
"started_at": self._call_started_at.isoformat() if self._call_started_at else None,
"ended_at": datetime.now(timezone.utc).isoformat(),
"freq": self._current_freq,
"srcaddr": self._current_srcaddr,
"started_at": _iso(started_at),
"started_at_epoch": started_at,
"ended_at": _iso(end_epoch),
"ended_at_epoch": end_epoch,
"transmissions": self._transmissions,
"end_reason": reason,
"attributed": attributed,
"driver": "audio" if self._audio_driven else "console",
}
logger.info(f"Call end: id={self._active_call_id}")
duration = (end_epoch - started_at) if started_at is not None else 0.0
if not attributed:
self._unattributed_segments += 1
window_start = started_at if started_at is not None else end_epoch
logger.error(
f"ORPHAN AUDIO: {duration:.2f}s of audio ({self._active_call_id}, reason={reason}, "
f"window {window_start:.3f}-{end_epoch:.3f}) had NO OP25 talkgroup anywhere within "
f"{ATTRIBUTION_LOOKBACK_SECONDS:.0f}s before or {ATTRIBUTION_LOOKAHEAD_SECONDS:.0f}s "
f"after it. It will be DISCARDED, not uploaded — an untagged call would poison "
f"incident correlation. Causes: Liquidsoap fallback/test audio on drb_sink, OP25 not "
f"decoding the control channel, or a dropped call_log. Console history holds "
f"{len(self._console)} recent observations; total orphans this run: "
f"{self._unattributed_segments}."
)
else:
logger.info(
f"Call end: id={self._active_call_id} tgid={self._current_tgid} "
f"reason={reason} transmissions={self._transmissions} duration={duration:.2f}s"
)
# Clear state before awaiting so a re-entrant tick can't see a half-closed
# segment (and so an immediately-following _open_segment is clean).
self._active_call_id = None
self._current_tgid = None
self._current_tgid_name = None
self._hang_counter = 0
self._call_started_at = None
self._current_freq = None
self._current_srcaddr = None
self._started_at = None
self._transmissions = 0
self._tx_active = False
self._last_tx_end = None
self._audio_driven = False
if self.on_call_end:
await self.on_call_end(payload)
def _resolve_attribution(self, start: float, end: float) -> None:
"""
Last chance to label an audio-driven segment, run at close.
Only ADOPTS a talkgroup when the segment still has none. A tgid we
already hold came from a grant or a channel row — the console stating
outright who was transmitting — and an inference over a window is not
allowed to overrule a direct statement. This matters because the window
deliberately extends past the audio (ATTRIBUTION_LOOKAHEAD_SECONDS, and
the tail pad on a split), so a neighbouring call's console activity can
legitimately fall inside it.
A disagreement is still worth knowing about, so it is logged: it means
two talkgroups' console activity overlaps one recording, i.e. the split
logic should have fired and did not.
"""
found = self._attribute(start, end)
if found is None:
return
if self._current_tgid is None:
logger.info(
f"Attributed {self._active_call_id} at close to tgid {found.tgid} "
f"(overlap {found.overlap}, grants {found.grants}, nearby {found.nearby})."
)
self._current_tgid = found.tgid
if found.name:
self._current_tgid_name = found.name
if found.freq is not None and not self._current_freq:
self._current_freq = found.freq
if found.rid is not None and self._current_srcaddr is None:
self._current_srcaddr = found.rid
self._transmissions = max(self._transmissions, found.grants)
return
others = sorted(set(found.competing) | ({found.tgid} if found.tgid != self._current_tgid else set()))
others = [tgid for tgid in others if tgid != self._current_tgid]
if others:
logger.warning(
f"Segment {self._active_call_id} (tgid {self._current_tgid}) overlaps console "
f"activity for {others} as well — the split logic should have fired and did not. "
"Keeping the talkgroup the console stated directly."
)
if not self._current_tgid_name and found.tgid == self._current_tgid and found.name:
self._current_tgid_name = found.name
# ------------------------------------------------------------------
# Public state (consumed by routers/api.py, main.py and the dashboards)
# ------------------------------------------------------------------
@property
def active_call_id(self) -> Optional[str]:
return self._active_call_id
@@ -131,5 +898,10 @@ class MetadataWatcher:
def is_active(self) -> bool:
return self._active_call_id is not None
@property
def unattributed_segments(self) -> int:
"""Orphan-audio segments discarded since start. Surfaced on /api/status."""
return self._unattributed_segments
metadata_watcher = MetadataWatcher()
+101 -3
View File
@@ -1,6 +1,7 @@
import asyncio
import json
from datetime import datetime, timezone
from collections import deque
from typing import Optional, Callable, Awaitable, Dict, Any
import paho.mqtt.client as mqtt
from app.config import settings
@@ -23,12 +24,24 @@ class MQTTManager:
self.on_config_push: Optional[ConfigCallback] = None
self.on_api_key: Optional[ApiKeyCallback] = None
self._offline_buffer = deque(maxlen=settings.offline_call_buffer_size)
nid = settings.node_id
self._t_checkin = f"nodes/{nid}/checkin"
self._t_status = f"nodes/{nid}/status"
self._t_metadata = f"nodes/{nid}/metadata"
self._t_commands = f"nodes/{nid}/commands"
self._t_config = f"nodes/{nid}/config"
# TODO(mqtt-cutover): dead once enrollment lands client-side. This
# was the pre-dynsec key-delivery path (server retain-publishes the
# api_key here after admin approval; node asks for redelivery via
# _t_key_request if none shows up). Under dynsec a node with no
# api_key can't authenticate to the broker at all — see
# _build_client() — so this subscribe is only ever reachable while
# still using the legacy mqtt_user/mqtt_pass fallback against a
# pre-cutover broker. Left in as the rollback path per
# MQTT-PUBLIC-AUTH-PLAN.md; remove together with the server's
# matching TODO(mqtt-cutover) markers once enrollment replaces it.
self._t_api_key = f"nodes/{nid}/api_key"
self._t_key_request = f"nodes/{nid}/key_request"
self._t_discovery = "nodes/discovery/request"
@@ -38,8 +51,47 @@ class MQTTManager:
callback_api_version=mqtt.CallbackAPIVersion.VERSION2,
client_id=settings.node_id,
)
if settings.mqtt_user:
api_key = credentials.get_api_key()
if api_key:
# Post-cutover auth: broker's dynsec plugin authenticates this
# exact (username, password) pair as this node's own client — see
# Server/drb-c2-core/app/internal/dynsec.py upsert_node_client()
# and MQTT-PUBLIC-AUTH-PLAN.md. node_id doubles as the dynsec
# username AND the %u substitution in the "node" role's
# nodes/%u/# ACL pattern, so it must match exactly what C2 has on
# file for this node (it always does — node_id is not operator
# editable post-provisioning).
client.username_pw_set(settings.node_id, api_key)
elif settings.mqtt_user:
# Legacy fallback — only valid against a pre-cutover broker still
# using mosquitto's old password_file auth. See config.py's
# mqtt_user/mqtt_pass docstring. Not accepted by a dynsec broker.
client.username_pw_set(settings.mqtt_user, settings.mqtt_pass)
else:
# No api_key on disk and no legacy shared login configured. A
# dynsec broker (allow_anonymous false) refuses this outright —
# expected, not a bug to route around here: this node hasn't been
# enrolled/approved yet, and the enrollment flow that would fix
# that client-side is a later, separate pass (out of scope here;
# see MQTT-PUBLIC-AUTH-PLAN.md). paho's reconnect_delay_set()
# below bounds the retry rate (2..60s exponential backoff), so
# this degrades to a slow, clearly-logged refusal loop via
# _on_connect's "MQTT connect refused" line — not a hot spin.
logger.warning(
"No API key on disk and no legacy MQTT_USER configured — "
"connecting without credentials; the broker is expected to "
"refuse this until the node is enrolled/approved."
)
if settings.mqtt_tls:
# No arguments = system CA store + ssl.CERT_REQUIRED (verified
# against paho's tls_set() source/docstring — unverified by
# running anything, per instruction). The broker presents a real
# Let's Encrypt cert for mqtt.<domain>:8883, so default
# verification is exactly correct: do not pass ca_certs, do not
# call tls_insecure_set(True).
client.tls_set()
lwt = json.dumps({
"node_id": settings.node_id,
@@ -59,11 +111,13 @@ class MQTTManager:
self._connected = True
client.subscribe(self._t_commands, qos=1)
client.subscribe(self._t_config, qos=1)
client.subscribe(self._t_api_key, qos=2)
client.subscribe(self._t_api_key, qos=2) # TODO(mqtt-cutover): see _t_api_key comment above
client.subscribe(self._t_discovery, qos=0)
logger.info("MQTT connected.")
asyncio.run_coroutine_threadsafe(self._publish_checkin(), self._loop)
# TODO(mqtt-cutover): see _t_api_key comment above
asyncio.run_coroutine_threadsafe(self._maybe_request_key(), self._loop)
asyncio.run_coroutine_threadsafe(self._flush_offline_buffer(), self._loop)
else:
logger.error(f"MQTT connect refused: {reason_code}")
@@ -130,7 +184,23 @@ class MQTTManager:
"timestamp": datetime.now(timezone.utc).isoformat(),
**data,
}
self._publish(self._t_metadata, payload, qos=1)
if not self._connected:
if event_type == "call_end":
self._offline_buffer.append((self._t_metadata, payload))
logger.warning(f"MQTT offline. Buffered call_end event for {data.get('call_id')}")
else:
logger.debug(f"MQTT offline. Dropping metadata event: {event_type}")
else:
self._publish(self._t_metadata, payload, qos=1)
async def _flush_offline_buffer(self):
if not self._offline_buffer:
return
count = len(self._offline_buffer)
logger.info(f"Relaying {count} buffered call_end events from offline queue.")
while self._offline_buffer:
topic, payload = self._offline_buffer.popleft()
self._publish(topic, payload, qos=1)
async def _maybe_request_key(self):
"""After connecting, wait for any retained api_key message to arrive.
@@ -140,10 +210,16 @@ class MQTTManager:
logger.info("No API key on disk — requesting re-delivery from C2 server.")
self._publish(self._t_key_request, {}, qos=1)
async def publish_checkin(self):
await self._publish_checkin()
async def _publish_checkin(self):
from app.internal.discord_radio import radio_bot
from app.internal.config_manager import load_node_config
from app.internal.op25_client import op25_client
from app.internal.secondary_sdr_client import secondary_sdr_client
config = load_node_config()
devices = await op25_client.devices()
payload = {
"node_id": settings.node_id,
"name": settings.node_name,
@@ -155,7 +231,29 @@ class MQTTManager:
"is_overridden": config.override_system_id is not None and config.node_type != "portable",
"override_system_id": config.override_system_id,
"enforce_override_timeout": config.enforce_override_timeout,
"secondary_sdr_mode": config.secondary_sdr_mode,
"secondary_sdr_priority": config.secondary_sdr_priority,
}
payload["sdr_pins"] = config.sdr_pins
secondary = await secondary_sdr_client.status()
if secondary is not None:
payload["secondary_sdr_running"] = [r["mode"] for r in secondary.get("running", [])]
if secondary.get("devices") is not None:
payload["sdr_devices"] = [
{k: d.get(k) for k in ("index", "serial", "name", "duplicate_serial")}
for d in secondary["devices"]
]
from app.internal.sdr_settings import op25_serial
payload["op25_sdr_serial"] = op25_serial(config, secondary["devices"])
# Best-effort hardware report — omit rather than guess. Prefer the
# secondary-sdr container's count: op25's :stable image has no lsusb and
# its /op25/devices answers 0 rather than "unknown" when that fails. A
# node running op25 has at least one SDR, so 0 is never a real reading.
count = secondary.get("sdr_count") if secondary is not None else None
if not count and devices:
count = devices.get("count") or None
if count is not None:
payload["sdr_count"] = count
self._publish(self._t_checkin, payload, qos=1)
def _publish(self, topic: str, payload: dict, qos: int = 0, retain: bool = False):
+98 -15
View File
@@ -1,8 +1,34 @@
import httpx
from typing import Optional, Dict, Any
from dataclasses import dataclass, field
from typing import Optional, Dict, Any, List
from app.config import settings
from app.internal.logger import logger
# The OP25 HTTP terminal answers a single "update" command with a LIST of
# messages, each tagged with a `json_type`. We care about two of them:
#
# channel_update — current receiver state. `channels` holds the channel ids and
# each id is also a top-level key holding that channel's dict
# (freq/tgid/tag/srcaddr/svcopts/hold_tgid/…).
#
# call_log — an EVENT QUEUE, not a snapshot. `log` holds entries appended
# by tk_p25.log_call() at channel-grant time, each stamped with
# OP25's own time.time(). get_call_log() DRAINS the deque, so
# every entry is delivered exactly once and a missed poll loses
# it forever. The deque is capped at CALL_LOG_MAX_LEN = 10, so
# the consumer must keep up.
#
# Everything else (trunk_update, rx_update, terminal_config, …) is ignored.
TERMINAL_UPDATE_COMMAND = [{"command": "update", "arg1": 0, "arg2": 0}]
@dataclass
class TerminalUpdate:
"""One decoded poll of the OP25 HTTP terminal."""
channels: List[Dict[str, Any]] = field(default_factory=list)
call_log: List[Dict[str, Any]] = field(default_factory=list)
class OP25Client:
def __init__(self):
@@ -39,34 +65,91 @@ class OP25Client:
logger.error(f"OP25 status failed: {e}")
return None
async def devices(self) -> Optional[Dict[str, Any]]:
try:
async with httpx.AsyncClient(timeout=5) as client:
r = await client.get(f"{self.api_url}/op25/devices")
r.raise_for_status()
return r.json()
except Exception as e:
logger.error(f"OP25 device enumeration failed: {e}")
return None
async def generate_config(self, config: Dict[str, Any]) -> bool:
try:
async with httpx.AsyncClient(timeout=10) as client:
r = await client.post(f"{self.api_url}/op25/generate-config", json=config)
r.raise_for_status()
return True
except Exception as e:
logger.error(f"OP25 generate-config failed: {e}")
return False
# Every generated config opens its dongle by serial, never "first
# found" (node-26#11) — here so no generation path can skip it.
from app.internal.sdr_settings import pin_op25_device
await pin_op25_device()
return True
async def get_terminal_status(self) -> Optional[Any]:
"""Poll the OP25 HTTP terminal for current call metadata."""
async def poll_terminal(self) -> Optional[TerminalUpdate]:
"""
Poll the OP25 HTTP terminal once and decode every message we understand.
Returns None only when OP25 is unreachable / returned garbage — callers
use that to distinguish "no traffic" from "no OP25".
"""
try:
async with httpx.AsyncClient(timeout=3) as client:
r = await client.post(
self.terminal_url,
json=[{"command": "update", "arg1": 0, "arg2": 0}],
)
r = await client.post(self.terminal_url, json=TERMINAL_UPDATE_COMMAND)
r.raise_for_status()
messages = r.json()
for msg in messages:
if msg.get("json_type") == "channel_update":
channels = msg.get("channels", [])
if channels:
return msg.get(str(channels[0]), {})
return None
return parse_terminal_messages(r.json())
except Exception:
return None
async def get_terminal_status(self) -> Optional[Dict[str, Any]]:
"""
Compatibility shim: the first channel's state dict, as this used to return.
Prefer poll_terminal() — this discards the call_log, which is the only
source of exact call-start timestamps.
"""
update = await self.poll_terminal()
if not update or not update.channels:
return None
return update.channels[0]
def parse_terminal_messages(messages: Any) -> TerminalUpdate:
"""
Decode an OP25 terminal response into channel state + call-log events.
Deliberately permissive: the response may be a bare dict instead of a list,
may contain json_type values we have never seen, and individual entries may
be malformed. Anything unrecognised is skipped rather than raising, because
dropping a whole poll would drop call_log events that are never re-sent.
"""
update = TerminalUpdate()
if isinstance(messages, dict):
messages = [messages]
if not isinstance(messages, list):
return update
for msg in messages:
if not isinstance(msg, dict):
continue
json_type = msg.get("json_type")
if json_type == "channel_update":
for chan_id in msg.get("channels") or []:
channel = msg.get(str(chan_id))
if isinstance(channel, dict):
update.channels.append(channel)
elif json_type == "call_log":
for entry in msg.get("log") or []:
if isinstance(entry, dict):
update.call_log.append(entry)
return update
op25_client = OP25Client()
+131
View File
@@ -0,0 +1,131 @@
"""
Raw PCM primitives: the one place that knows the capture format.
The capture pipeline buffers RAW PCM (signed 16-bit little-endian, mono,
22050 Hz) instead of MP3. Three things fall out of that, and they are the whole
reason for the change:
1. Silence detection is integer arithmetic over the bytes as they arrive —
no decode, no FFmpeg, no second process. That is what makes an
AUDIO-DRIVEN call boundary possible at all.
2. Trimming becomes a byte-offset slice instead of a second encode pass.
3. MP3 encoding happens exactly ONCE, at save time, so uploads stop being
double-encoded.
WHY SILENCE IS UNAMBIGUOUS HERE: between transmissions the captured stream is
the monitor of a PulseAudio *null sink*, which emits digital silence, not an
analog noise floor. Measured on a live node, the gap between transmissions sits
at about -91 dBFS — that is 20*log10(1/32768), i.e. one least-significant bit,
the quietest thing a 16-bit sample can be without being exactly zero. Speech on
the same node averages about -18 dBFS. There is therefore ~70 dB of daylight
between "silence" and "voice", and the threshold does NOT need field
calibration against radio noise the way an analog squelch tail would.
MEASUREMENT IS RMS, NOT PEAK. Peak would be cheaper but a single decoder click
would read as voice for a whole window; RMS over a window is the honest
"is there signal here" answer. The cost is a Python loop over the window's
samples, which is affordable because of how little audio is ever scanned:
one ~46 ms chunk per chunk arrival at capture time, and only the head/tail of a
finished recording at trim time (see audio_trim.MAX_SCAN_SECONDS). A cheap
all-zero fast path in C skips the loop entirely for exactly-silent windows.
BYTE ORDER: FFmpeg is asked for s16le. `array("h")` is native-endian, so on a
big-endian host the samples are byte-swapped before use. Every DRB target is
little-endian today; this is three lines of insurance, not a real scenario.
"""
import math
import sys
from array import array
from typing import Union
# Capture format. MP3_SAMPLE_RATE in call_recorder must stay equal to
# SAMPLE_RATE — the encode at save time is a straight pass with no resample.
SAMPLE_RATE = 22050
SAMPLE_WIDTH = 2
CHANNELS = 1
FRAME_BYTES = SAMPLE_WIDTH * CHANNELS
BYTES_PER_SECOND = SAMPLE_RATE * FRAME_BYTES # 44100 B/s
# 16-bit full scale. A sample of 32768 (or -32768) is 0 dBFS.
FULL_SCALE = 32768.0
# Reported for a window with no signal at all. Any real threshold is far above
# this, so it always compares as "silent" without special-casing log10(0).
SILENT_DBFS = -120.0
_NEEDS_BYTESWAP = sys.byteorder != "little"
Buffer = Union[bytes, bytearray]
def align(nbytes: int) -> int:
"""Round a byte count DOWN to a whole number of samples."""
if nbytes <= 0:
return 0
return nbytes - (nbytes % FRAME_BYTES)
def seconds(nbytes: int) -> float:
"""Duration of `nbytes` of PCM."""
return nbytes / BYTES_PER_SECOND
def byte_offset(sec: float) -> int:
"""Sample-aligned byte offset of `sec` seconds into a PCM buffer."""
return align(int(sec * BYTES_PER_SECOND))
def samples(buf: Buffer) -> array:
"""View a PCM buffer as signed 16-bit samples, dropping any partial frame."""
usable = align(len(buf))
data = array("h")
if usable:
data.frombytes(bytes(buf[:usable]))
if _NEEDS_BYTESWAP:
data.byteswap()
return data
def is_all_zero(buf: Buffer) -> bool:
"""
True when every byte is zero — exact digital silence.
`bytes.count` runs in C, so this is the cheap path that lets a long scan
over silence stay fast without touching the per-sample loop below.
"""
return len(buf) > 0 and buf.count(0) == len(buf)
def rms(buf: Buffer) -> float:
"""Root-mean-square amplitude in raw sample units (0 .. 32768)."""
data = samples(buf)
if not data:
return 0.0
total = 0
for sample in data:
total += sample * sample
return math.sqrt(total / len(data))
def rms_dbfs(buf: Buffer) -> float:
"""RMS level of a PCM window in dBFS. SILENT_DBFS for an empty/zero window."""
if not buf or is_all_zero(buf):
return SILENT_DBFS
value = rms(buf)
if value <= 0.0:
return SILENT_DBFS
return 20.0 * math.log10(min(value, FULL_SCALE) / FULL_SCALE)
def is_silent(buf: Buffer, threshold_db: float) -> bool:
"""
True when a PCM window carries no signal above `threshold_db` (dBFS RMS).
An empty buffer counts as silence: "no audio arrived" must never read as
"someone is talking", or a stalled capture would hold a segment open.
"""
if not buf:
return True
if is_all_zero(buf):
return True
return rms_dbfs(buf) < threshold_db
+141
View File
@@ -0,0 +1,141 @@
"""
PulseAudio readiness helpers.
The PulseAudio daemon lives in the `op25` container and exposes its native
socket on the shared `pulse_socket` docker volume (mounted at /run/pulse in
both containers, with PULSE_SERVER=unix:/run/pulse/native).
`op25-container/docker-entrypoint.sh` waits (bounded) for that daemon to
actually answer before starting its own app, and the edge-node needs the same
guarantee before launching FFmpeg: FFmpeg with `-f pulse` fails instantly if
nothing is listening, and used to stay dead for the lifetime of the process.
This module is the wait.
HISTORY / WHY THIS CHECKS LIVENESS, NOT FILE EXISTENCE: the `pulse_socket`
named volume survives container recreation, but the daemon process that
created the socket does not. Observed on live hardware: a stale
`/run/pulse/native` socket file and `/run/pulse/pid` from a killed daemon
were still in the volume after `docker compose up -d --build` recreated the
containers. PulseAudio refused to start ("Daemon already running") because of
the stale pid file, so nothing was actually listening on the socket — but the
socket *file* still existed. An earlier version of this module (and of the
op25 entrypoint) only checked `stat.S_ISSOCK` on the path, so it reported
"ready" against a dead daemon, FFmpeg launched anyway, and immediately failed
with "No such process" in a tight restart loop. Readiness here means "a
PulseAudio connection actually succeeds," never "a file exists at this path."
NOTE on the source name: the op25 entrypoint starts pulseaudio with `-n`, which
skips /etc/pulse/system.pa entirely and loads modules from the command line
instead. That means the `set-default-source drb_sink.monitor` line in system.pa
is NOT applied at runtime, so `-i default` is unreliable. Always address the
monitor explicitly via settings.pulse_source (default "drb_sink.monitor").
"""
import asyncio
import os
import shutil
import subprocess
from typing import Optional
from app.config import settings
from app.internal.logger import logger
DEFAULT_SOCKET_PATH = "/run/pulse/native"
POLL_INTERVAL = 0.5
# Bounded timeout for a single `pactl info` liveness probe. Kept short: this
# runs synchronously on the calling thread (see is_ready()), and callers of
# is_ready() include a sync code path inside the Discord voice bot, so a slow
# probe would stall its event loop. wait_until_ready() runs probes off-thread
# via asyncio.to_thread and can afford this bound comfortably within its own
# much larger PULSE_WAIT_TIMEOUT.
PROBE_TIMEOUT_SECONDS = 1.5
def socket_path() -> str:
"""Resolve the PulseAudio socket path from PULSE_SERVER (`unix:/path` form)."""
server = os.environ.get("PULSE_SERVER", "")
if server.startswith("unix:"):
candidate = server[len("unix:"):].strip()
if candidate:
return candidate
return DEFAULT_SOCKET_PATH
def _probe_env(path: str) -> dict:
env = dict(os.environ)
env["PULSE_SERVER"] = f"unix:{path}"
return env
def _daemon_responds() -> bool:
"""
True only when a PulseAudio daemon actually answers on the configured
socket. Shells out to `pactl info` (from `pulseaudio-utils`, installed
alongside `libpulse0` in the edge-node image) rather than re-implementing
the native protocol handshake in Python — this container has no other use
for talking to PulseAudio directly, so a subprocess call is the smallest
correct implementation.
Deliberately does NOT check `os.path.exists`/`stat.S_ISSOCK` first: a
stale socket file from a killed daemon passes that check and always did,
which is the exact defect this function replaces.
"""
path = socket_path()
pactl = shutil.which("pactl")
if pactl is None:
logger.error("pactl not found in PATH — cannot verify PulseAudio liveness.")
return False
try:
result = subprocess.run(
[pactl, "info"],
env=_probe_env(path),
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
timeout=PROBE_TIMEOUT_SECONDS,
)
return result.returncode == 0
except (subprocess.TimeoutExpired, OSError):
return False
def is_ready() -> bool:
"""
True when a PulseAudio daemon is alive and answering right now.
Synchronous and bounded by PROBE_TIMEOUT_SECONDS — used from a sync call
site (discord_radio._play_stream). Prefer `wait_until_ready()` from async
code so the probe doesn't block the event loop.
"""
return _daemon_responds()
async def wait_until_ready(timeout: Optional[float] = None) -> bool:
"""
Block until a PulseAudio daemon actually answers, or `timeout` seconds
elapse.
Bounded on purpose — never hang the caller forever. Returns True once a
live connection succeeds, False on timeout (caller decides whether to
retry). Each probe runs via asyncio.to_thread so the subprocess call never
blocks the event loop.
"""
limit = settings.pulse_wait_timeout if timeout is None else timeout
path = socket_path()
if await asyncio.to_thread(_daemon_responds):
return True
logger.info(f"Waiting up to {limit:.0f}s for a live PulseAudio daemon at {path}…")
waited = 0.0
while waited < limit:
await asyncio.sleep(POLL_INTERVAL)
waited += POLL_INTERVAL
if await asyncio.to_thread(_daemon_responds):
logger.info(f"PulseAudio daemon live after {waited:.1f}s.")
return True
logger.error(
f"PulseAudio daemon at {path} not responding after {limit:.0f}s — "
"is the op25 container's daemon actually up? Audio capture will retry."
)
return False
+102
View File
@@ -0,0 +1,102 @@
"""
Which SDR does what on this node (node-26#9, node-26#11).
- OP25 always has exactly one dongle. `sdr_pins["op25"]` names it by serial;
unpinned, OP25 keeps its historical "first dongle" — but that dongle is now
named by serial too, so the decoders can never take it out from under OP25.
- Every other dongle runs the next enabled service in `secondary_sdr_priority`.
`sdr_pins[mode]` optionally binds a service to the dongle carrying its
antenna; unpinned services take any spare.
One apply path for the local dashboard, C2 commands and config pushes.
"""
import asyncio
import json
from pathlib import Path
from typing import Any, Dict, Iterable, List, Optional
from app.config import settings
from app.internal.config_manager import load_node_config, save_node_config
from app.internal.logger import logger
from app.internal.secondary_sdr_client import secondary_sdr_client
from app.models import NodeConfig, normalize_sdr_pins, normalize_secondary_priority
_OP25_CONFIG = Path(settings.config_path) / "active.cfg.json"
def op25_serial(cfg: NodeConfig, devs: Optional[List[Dict[str, Any]]]) -> Optional[str]:
"""The serial of OP25's dongle: the pin, else the first detected dongle
(what OP25's plain "rtl" device string has always opened)."""
if cfg.sdr_pins.get("op25"):
return cfg.sdr_pins["op25"]
return devs[0]["serial"] if devs else None
async def pin_op25_device() -> Optional[str]:
"""Rewrite OP25's generated config to open its dongle by serial.
Runs after every generate-config (op25_client.generate_config). A plain
"rtl" means "device 0", and which dongle is device 0 depends on who opened
what first — that is how OP25 lost its SDR to readsb on 2026-09-27.
Left as "rtl" only when the serial is unknown or shared by two dongles.
"""
devs = await secondary_sdr_client.devices()
serial = op25_serial(load_node_config(), devs)
if not serial or (devs and sum(d["serial"] == serial for d in devs) > 1):
return None
try:
cfg = json.loads(_OP25_CONFIG.read_text())
for dev in cfg.get("devices", []):
dev["args"] = f"rtl={serial}"
_OP25_CONFIG.write_text(json.dumps(cfg, indent=2))
except Exception as e:
logger.error(f"Could not pin OP25 to SDR {serial}: {e}")
return None
return serial
async def apply_secondaries() -> Optional[List[str]]:
"""Start/stop decoders to match the saved priority and pins, never on OP25's dongle."""
cfg = load_node_config()
if not cfg.secondary_sdr_priority:
return await secondary_sdr_client.apply([], {}, [])
devs = await secondary_sdr_client.devices()
reserved = [s for s in [op25_serial(cfg, devs)] if s]
pins = {m: s for m, s in cfg.sdr_pins.items() if m != "op25"}
return await secondary_sdr_client.apply(cfg.secondary_sdr_priority, pins, reserved)
async def set_sdr_settings(priority: Optional[Iterable[str]] = None,
pins: Optional[Dict[str, Optional[str]]] = None) -> Optional[List[str]]:
"""Persist and apply priority and/or pins. OP25 restarts only if its own
dongle changed; reordering the spare dongles never interrupts P25.
Returns the decoders now running (None if the container is unreachable).
"""
cfg = load_node_config()
old_op25 = cfg.sdr_pins.get("op25")
if priority is not None:
ordered = normalize_secondary_priority(priority)
cfg.secondary_sdr_priority = ordered
cfg.secondary_sdr_mode = ordered[0] if ordered else "none"
if pins is not None:
cfg.sdr_pins = normalize_sdr_pins(pins)
save_node_config(cfg)
if cfg.sdr_pins.get("op25") != old_op25:
from app.internal.op25_client import op25_client
# Free the new dongle first if a decoder holds it, then move OP25.
await secondary_sdr_client.apply([], {}, [])
serial = await pin_op25_device()
await op25_client.stop()
await asyncio.sleep(2)
await op25_client.start()
logger.info(f"OP25 moved to SDR {serial or 'rtl (first dongle)'}")
running = await apply_secondaries()
logger.info(f"SDR settings: priority={cfg.secondary_sdr_priority!r} pins={cfg.sdr_pins!r}; running {running!r}")
# Report straight away so C2's view doesn't wait for the next heartbeat.
from app.internal.mqtt_manager import mqtt_manager
asyncio.create_task(mqtt_manager.publish_checkin())
return running
@@ -0,0 +1,84 @@
import httpx
from typing import Any, Dict, List, Optional
from app.config import settings
from app.internal.logger import logger
class SecondarySdrClient:
"""Talks to the secondary-sdr-container (node-26#9) over its control API.
Mirrors op25_client.py's shape on purpose — same failure handling (log
and return None/False rather than raise), since this container is
optional and its absence must never break the primary op25 radio path.
"""
def __init__(self):
self.api_url = settings.secondary_sdr_api_url
async def devices(self) -> Optional[List[Dict[str, Any]]]:
"""Every RTL-SDR on the node with its serial, or None if unreachable."""
try:
async with httpx.AsyncClient(timeout=5) as client:
r = await client.get(f"{self.api_url}/secondary/devices")
r.raise_for_status()
return r.json().get("devices", [])
except Exception as e:
logger.error(f"Secondary SDR device list failed: {e}")
return None
async def apply(self, priority: List[str], pins: Dict[str, str], reserved: List[str]) -> Optional[List[str]]:
"""Run decoders down the priority list until SDRs run out, honouring
pins and never touching `reserved` (op25's) dongles. Returns the modes
actually running, or None if the container is unreachable."""
body = {"priority": priority, "pins": pins, "reserved": reserved}
try:
async with httpx.AsyncClient(timeout=30) as client:
r = await client.post(f"{self.api_url}/secondary/apply", json=body)
r.raise_for_status()
return r.json().get("running", [])
except Exception as e:
logger.error(f"Secondary SDR apply (priority={priority!r}) failed: {e}")
return None
async def start(self, mode: str) -> bool:
try:
async with httpx.AsyncClient(timeout=10) as client:
r = await client.post(f"{self.api_url}/secondary/start", json={"mode": mode})
r.raise_for_status()
return True
except Exception as e:
logger.error(f"Secondary SDR start (mode={mode!r}) failed: {e}")
return False
async def stop(self) -> bool:
try:
async with httpx.AsyncClient(timeout=10) as client:
r = await client.post(f"{self.api_url}/secondary/stop")
r.raise_for_status()
return True
except Exception as e:
logger.error(f"Secondary SDR stop failed: {e}")
return False
async def status(self) -> Optional[Dict[str, Any]]:
try:
async with httpx.AsyncClient(timeout=5) as client:
r = await client.get(f"{self.api_url}/secondary/status")
r.raise_for_status()
return r.json()
except Exception as e:
logger.error(f"Secondary SDR status failed: {e}")
return None
async def data(self) -> Optional[Dict[str, Any]]:
try:
async with httpx.AsyncClient(timeout=5) as client:
r = await client.get(f"{self.api_url}/secondary/data")
r.raise_for_status()
return r.json()
except Exception as e:
logger.error(f"Secondary SDR data fetch failed: {e}")
return None
secondary_sdr_client = SecondarySdrClient()
+4 -1
View File
@@ -8,6 +8,7 @@ from app.internal import credentials
_CACHE_FILE = Path(settings.config_path) / "systems_cache.json"
async def fetch_and_cache_systems() -> bool:
"""Fetch all systems from the C2 server and cache them locally."""
if not settings.c2_url:
@@ -23,7 +24,7 @@ async def fetch_and_cache_systems() -> bool:
r = await client.get(url, headers=headers)
r.raise_for_status()
systems = r.json()
_CACHE_FILE.parent.mkdir(parents=True, exist_ok=True)
_CACHE_FILE.write_text(json.dumps(systems, indent=2))
logger.info(f"Cached {len(systems)} systems from C2.")
@@ -32,6 +33,7 @@ async def fetch_and_cache_systems() -> bool:
logger.warning(f"Failed to fetch systems from C2: {e}. Offline cache will be used.")
return False
def load_cached_systems() -> List[Dict[str, Any]]:
"""Load cached systems from disk."""
if _CACHE_FILE.exists():
@@ -41,6 +43,7 @@ def load_cached_systems() -> List[Dict[str, Any]]:
logger.error(f"Failed to read systems cache: {e}")
return []
def get_cached_system(system_id: str) -> Optional[Dict[str, Any]]:
"""Retrieve a single system config from the cache."""
systems = load_cached_systems()
@@ -0,0 +1,45 @@
import asyncio
import httpx
from app.config import settings
from app.internal import credentials
from app.internal.config_manager import load_node_config
from app.internal.logger import logger
from app.internal.secondary_sdr_client import secondary_sdr_client
# How often the second-SDR decoder's current snapshot is forwarded to C2
# (node-26#9). This is a live-map overlay, not a flight/vessel history, so
# there is no backlog/retry on a missed tick — the next one supersedes it.
UPLINK_INTERVAL_SECONDS = 10
async def _post_snapshot(path: str, body: dict) -> None:
if not settings.c2_url:
return
api_key = credentials.get_api_key()
if not api_key:
return
headers = {"Authorization": f"Bearer {api_key}"}
try:
async with httpx.AsyncClient(timeout=10) as client:
r = await client.post(f"{settings.c2_url}{path}", json=body, headers=headers)
r.raise_for_status()
except Exception as e:
logger.debug(f"Telemetry uplink to {path} failed: {e}")
async def telemetry_uplink_loop():
while True:
await asyncio.sleep(UPLINK_INTERVAL_SECONDS)
if not load_node_config().secondary_sdr_priority:
continue
snapshot = await secondary_sdr_client.data()
if not snapshot:
continue
# Several decoders can run at once (one per spare SDR), so forward
# whatever each produced rather than keying off a single mode.
if snapshot.get("aircraft"):
await _post_snapshot("/telemetry/adsb", {"aircraft": snapshot["aircraft"]})
if snapshot.get("vessels"):
await _post_snapshot("/telemetry/ais", {"vessels": snapshot["vessels"]})
+143 -10
View File
@@ -1,8 +1,11 @@
import asyncio
from contextlib import asynccontextmanager
from datetime import datetime, timezone
from typing import Optional
from fastapi import FastAPI
from app.config import settings
from app.models import SystemConfig
from app.models import SystemConfig, normalize_sdr_pins, normalize_secondary_priority
from app.internal.logger import logger
from app.internal.mqtt_manager import mqtt_manager
from app.internal import credentials
@@ -20,34 +23,119 @@ from app.routers import api, ui
# Event handlers wired up at startup
# ---------------------------------------------------------------------------
def _iso(epoch: Optional[float]) -> Optional[str]:
"""Epoch → UTC ISO-8601, matching metadata_watcher's timestamp format."""
if epoch is None:
return None
return datetime.fromtimestamp(epoch, timezone.utc).isoformat()
# call_ids whose `call_start` has already gone out over MQTT. A segment can open
# before its talkgroup is known (audio onset can precede the OP25 grant), and
# C2's _on_call_start writes talkgroup_id straight into a new Firestore `calls`
# doc — publishing early with tgid=None would create a permanently untagged call.
# So the start is held back until attribution succeeds, and replayed just before
# the end event if it resolved late.
_published_starts: set = set()
async def on_call_start(data: dict):
radio_bot.start_stream()
await mqtt_manager.publish_status("recording")
await mqtt_manager.publish_metadata("call_start", data)
await call_recorder.start_recording(data["call_id"])
# started_at_epoch is the detected voice onset (or, in console fallback mode,
# OP25's call_log timestamp). The recorder slices the ring buffer back to it
# minus the pre-roll, so however late the poll loop noticed, the audio still
# starts in the right place.
await call_recorder.start_recording(
data["call_id"],
start_epoch=data.get("started_at_epoch"),
)
if data.get("attributed", True):
_published_starts.add(data["call_id"])
await mqtt_manager.publish_metadata("call_start", data)
else:
logger.info(
f"Call {data['call_id']} started on audio onset with no talkgroup yet — holding the "
"call_start event until the console attributes it."
)
async def on_call_end(data: dict):
radio_bot.stop_stream()
file_path = await call_recorder.stop_recording()
if file_path:
call_id = data["call_id"]
published_start = call_id in _published_starts
_published_starts.discard(call_id)
if not data.get("attributed", True):
# ORPHAN AUDIO. metadata_watcher has already logged the details at ERROR.
# The audio is dropped rather than uploaded: a call with no talkgroup is
# worse than no call at all, because it silently poisons correlation.
await call_recorder.discard_recording()
if published_start:
# Should not happen (attribution only ever improves), but if a start
# did go out, the doc must not be left hanging in "active".
data["audio_skipped"] = "unattributed"
await mqtt_manager.publish_metadata("call_end", data)
await mqtt_manager.publish_status("online")
return
recording = await call_recorder.stop_recording(end_epoch=data.get("ended_at_epoch"))
if recording is not None and recording.path is not None:
# Silence trimming shortens the audio, so the audio's own bounds no
# longer equal the call's. `started_at`/`ended_at` keep meaning the CALL
# (what OP25 observed on the control channel) — these extra fields carry
# the AUDIO's wall-clock bounds so playback, correlation and incident
# timelines can still map an audio offset back to real time:
# wall_clock_of(audio_offset_t) == audio_start_epoch + t
data["audio_start_at"] = _iso(recording.audio_start_epoch)
data["audio_end_at"] = _iso(recording.audio_end_epoch)
data["audio_start_epoch"] = recording.audio_start_epoch
data["audio_end_epoch"] = recording.audio_end_epoch
data["audio_lead_trimmed"] = round(recording.lead_trimmed, 3)
data["audio_tail_trimmed"] = round(recording.tail_trimmed, 3)
if recording.clamped_seconds:
data["audio_clamped_seconds"] = round(recording.clamped_seconds, 3)
if recording is not None and recording.path is not None:
node_cfg = load_node_config()
audio_url = await call_recorder.upload_recording(
file_path,
recording.path,
data["call_id"],
talkgroup_id=data.get("tgid"),
talkgroup_name=data.get("tgid_name"),
system_id=node_cfg.assigned_system_id,
audio_start_epoch=recording.audio_start_epoch,
audio_end_epoch=recording.audio_end_epoch,
)
if audio_url:
data["audio_url"] = audio_url
else:
logger.error(f"Audio upload failed for call {data['call_id']}. Verify C2_URL and Node API Key.")
elif recording is not None and recording.all_silence:
# Explicit policy: an all-silence recording is not uploaded. It has no
# transcript value and silence is what makes Whisper invent text.
data["audio_skipped"] = "all_silence"
logger.warning(f"Call {data['call_id']} was pure silence — no upload. Investigate the audio path.")
else:
logger.warning(
f"No recording file generated for call {data['call_id']} "
"— call may have been too short or Icecast unreachable."
"— PulseAudio capture may be down (check the op25 container and "
f"the {settings.pulse_source} source)."
)
if not published_start:
# Attribution arrived after the segment opened. Replay the start so C2
# creates the `calls` doc with the right talkgroup before the end event
# updates it.
start_payload = {
key: data[key]
for key in ("call_id", "tgid", "tgid_name", "freq", "srcaddr",
"started_at", "started_at_epoch", "attributed", "driver")
if key in data
}
await mqtt_manager.publish_metadata("call_start", start_payload)
await mqtt_manager.publish_metadata("call_end", data)
await mqtt_manager.publish_status("online")
@@ -70,6 +158,12 @@ async def on_command(payload: dict):
)
elif action == "discord_leave":
await radio_bot.leave()
elif action in ("set_sdr_config", "set_secondary_priority"):
from app.internal.sdr_settings import set_sdr_settings
try:
await set_sdr_settings(payload.get("priority"), payload.get("pins"))
except ValueError as e:
logger.error(f"Rejected SDR settings from C2: {e}")
elif action == "op25_restart":
from app.internal.op25_client import op25_client
await op25_client.stop()
@@ -136,6 +230,9 @@ async def on_config_push(payload: dict):
hardware_preset = payload.pop("hardware_preset", None)
ppm_override = payload.pop("ppm_override", None)
node_type = payload.pop("node_type", None)
secondary_sdr_priority = payload.pop("secondary_sdr_priority", None)
sdr_pins = payload.pop("sdr_pins", None)
secondary_sdr_mode = payload.pop("secondary_sdr_mode", None) # legacy single-mode C2
enforce_override_timeout = payload.pop("enforce_override_timeout", None)
try:
config = SystemConfig(**payload)
@@ -157,6 +254,16 @@ async def on_config_push(payload: dict):
node_cfg.node_type = node_type
if enforce_override_timeout is not None:
node_cfg.enforce_override_timeout = bool(enforce_override_timeout)
if secondary_sdr_priority is None and secondary_sdr_mode is not None:
secondary_sdr_priority = [secondary_sdr_mode]
if secondary_sdr_priority is not None:
node_cfg.secondary_sdr_priority = normalize_secondary_priority(secondary_sdr_priority)
node_cfg.secondary_sdr_mode = (node_cfg.secondary_sdr_priority or ["none"])[0]
if sdr_pins is not None:
try:
node_cfg.sdr_pins = normalize_sdr_pins(sdr_pins)
except ValueError as e:
logger.error(f"Ignoring invalid sdr_pins in config push: {e}")
save_node_config(node_cfg)
from app.internal.op25_client import op25_client
@@ -164,10 +271,16 @@ async def on_config_push(payload: dict):
logger.error(f"Failed to generate OP25 config for {config.name}")
return
# OP25's (pinned) dongle may be one a decoder holds right now: free the
# spares, restart OP25 on its own dongle, then hand the rest back out.
from app.internal.secondary_sdr_client import secondary_sdr_client
from app.internal.sdr_settings import apply_secondaries
await secondary_sdr_client.apply([], {}, [])
await op25_client.stop()
await asyncio.sleep(2)
await op25_client.start()
logger.info(f"Config push applied: {config.name}")
await apply_secondaries()
# ---------------------------------------------------------------------------
@@ -178,12 +291,19 @@ async def on_config_push(payload: dict):
async def lifespan(app: FastAPI):
logger.info(f"Edge node starting — ID: {settings.node_id}")
# Load persisted credentials (API key provisioned by C2 after approval)
# Load persisted credentials (API key provisioned by C2 after approval;
# also generates/loads the local dashboard's auth salt + session secret)
credentials.load()
from app.internal import auth
auth.warn_if_default_password()
# Wire callbacks
metadata_watcher.on_call_start = on_call_start
metadata_watcher.on_call_end = on_call_end
# Segment boundaries come from the audio itself; this is how the watcher
# sees it. Without this the watcher falls back to control-channel
# segmentation, which is measurably wrong in both directions.
metadata_watcher.audio_activity = call_recorder.audio_activity
mqtt_manager.on_command = on_command
mqtt_manager.on_config_push = on_config_push
mqtt_manager.on_api_key = on_api_key
@@ -191,7 +311,7 @@ async def lifespan(app: FastAPI):
# Start services (radio_bot starts on-demand when a discord_join command arrives)
await mqtt_manager.connect()
await metadata_watcher.start()
await call_recorder.start() # persistent Icecast stream buffer
await call_recorder.start() # persistent PulseAudio ring buffer
# Start system caching in background
from app.internal.system_cacher import fetch_and_cache_systems
@@ -202,7 +322,11 @@ async def lifespan(app: FastAPI):
initial_status = "online" if node_cfg.configured else "unconfigured"
await mqtt_manager.publish_status(initial_status)
active_config = node_cfg.override_config if (node_cfg.override_system_id and node_cfg.override_config) else node_cfg.system_config
active_config = (
node_cfg.override_config
if (node_cfg.override_system_id and node_cfg.override_config)
else node_cfg.system_config
)
if node_cfg.configured and active_config:
from app.internal.op25_client import op25_client
logger.info("Node is configured — waiting for OP25 API then generating config.")
@@ -213,12 +337,21 @@ async def lifespan(app: FastAPI):
logger.warning(f"OP25 not ready yet (attempt {attempt + 1}/10), retrying in 3s…")
await asyncio.sleep(3)
# After op25 has claimed its dongle, so the decoders only get the spares.
if node_cfg.secondary_sdr_priority:
from app.internal.sdr_settings import apply_secondaries
logger.info(f"Resuming secondary SDRs (priority={node_cfg.secondary_sdr_priority!r}) after restart.")
await apply_secondaries()
heartbeat_task = asyncio.create_task(mqtt_manager.heartbeat_loop())
from app.internal.telemetry_uplink import telemetry_uplink_loop
telemetry_task = asyncio.create_task(telemetry_uplink_loop())
yield # --- app running ---
logger.info("Edge node shutting down.")
heartbeat_task.cancel()
telemetry_task.cancel()
await metadata_watcher.stop()
await call_recorder.stop()
await radio_bot.stop()
+43 -2
View File
@@ -1,5 +1,5 @@
from pydantic import BaseModel
from typing import Optional, Dict, Any
from pydantic import BaseModel, model_validator
from typing import Optional, Dict, Any, Iterable, List
from enum import Enum
from datetime import datetime
@@ -23,6 +23,34 @@ class SystemConfig(BaseModel):
config: Dict[str, Any] # OP25-compatible config blob passed through to op25-container
# Decoders a node can run on the SDRs beyond op25's (node-26#9). op25 always
# keeps its own dongle; each further dongle runs the next entry of the node's
# secondary_sdr_priority, so a 3-SDR node with ["adsb", "ais"] runs both.
SECONDARY_SDR_MODES = ("adsb", "ais")
def normalize_secondary_priority(items: Iterable[str]) -> List[str]:
"""Known modes only, first occurrence wins, order preserved."""
out: List[str] = []
for m in items or []:
if m in SECONDARY_SDR_MODES and m not in out:
out.append(m)
return out
# Services that can be bound to a specific dongle by its USB serial.
SDR_PIN_KEYS = ("op25",) + SECONDARY_SDR_MODES
def normalize_sdr_pins(pins: Dict[str, Optional[str]]) -> Dict[str, str]:
"""Known services only, blank = automatic (dropped). Two services can't
share a dongle, so a serial claimed twice raises."""
out = {k: str(v).strip() for k, v in (pins or {}).items() if k in SDR_PIN_KEYS and v and str(v).strip()}
if len(set(out.values())) != len(out):
raise ValueError("Two services can't be pinned to the same SDR.")
return out
class NodeConfig(BaseModel):
node_id: str
node_name: str
@@ -34,9 +62,22 @@ class NodeConfig(BaseModel):
hardware_preset: str = "rtl-sdr-v3"
ppm_override: Optional[float] = None
node_type: str = "fixed" # fixed or portable
secondary_sdr_priority: List[str] = [] # ordered; SDRs beyond op25's run these top-down
sdr_pins: Dict[str, str] = {} # service (op25/adsb/ais) -> dongle serial; absent = automatic
# Legacy single-mode field (pre-priority). Still written as priority[0] so
# an older C2 reading checkins sees something sensible; read only to
# migrate a node_config.json saved before priority existed.
secondary_sdr_mode: str = "none"
enforce_override_timeout: bool = True
override_system_id: Optional[str] = None
override_config: Optional[SystemConfig] = None
offline_call_buffer_size: int = 35 # max call_end events to buffer while MQTT is offline
@model_validator(mode="after")
def _migrate_secondary_mode(self) -> "NodeConfig":
if not self.secondary_sdr_priority and self.secondary_sdr_mode in SECONDARY_SDR_MODES:
self.secondary_sdr_priority = [self.secondary_sdr_mode]
return self
class CallEvent(BaseModel):
+78 -18
View File
@@ -1,9 +1,10 @@
from fastapi import APIRouter, HTTPException, Body
from typing import Optional
from fastapi import APIRouter, Depends, HTTPException, Body
from pydantic import BaseModel
from typing import Dict, List, Optional
import asyncio
import httpx
from app.config import settings
from app.models import SystemConfig
from app.models import SystemConfig, SECONDARY_SDR_MODES
from app.internal.op25_client import op25_client
from app.internal.config_manager import load_node_config, save_node_config, apply_system_config
from app.internal.call_recorder import call_recorder
@@ -11,20 +12,30 @@ from app.internal.discord_radio import radio_bot
from app.internal.metadata_watcher import metadata_watcher
from app.internal import credentials
from app.internal.mqtt_manager import mqtt_manager
from app.internal import auth
router = APIRouter(prefix="/api", tags=["api"])
# Every route in this router requires auth — a valid dashboard session cookie
# or HTTP Basic (see app/internal/auth.py). No exemption exists for any route
# here: there is no health/liveness endpoint in this file or anywhere else in
# the edge node (confirmed against source — no docker healthcheck references
# one either), so nothing needs to stay open for a container healthcheck.
router = APIRouter(prefix="/api", tags=["api"], dependencies=[Depends(auth.require_auth)])
@router.get("/status")
async def get_status():
node_cfg = load_node_config()
op25_status = await op25_client.status()
active_tgid = metadata_watcher.current_tgid
active_tgid_name = metadata_watcher.current_tgid_name
system_name = None
active_config = node_cfg.override_config if (node_cfg.override_system_id and node_cfg.override_config) else node_cfg.system_config
active_config = (
node_cfg.override_config
if (node_cfg.override_system_id and node_cfg.override_config)
else node_cfg.system_config
)
if active_config:
system_name = active_config.name
if active_tgid:
@@ -48,6 +59,17 @@ async def get_status():
"assigned_system_id": node_cfg.assigned_system_id,
"system_name": system_name,
"is_recording": call_recorder.is_recording,
# Health of the PulseAudio capture that feeds every recording — the single
# most useful signal when recordings come back empty. Segment boundaries
# come from this stream, so audio_silence_seconds is also how far the
# node currently is from closing whatever it is recording.
"audio_capture": call_recorder.is_capturing,
"buffered_seconds": round(call_recorder.buffered_seconds, 1),
"audio_silence_seconds": round(call_recorder.audio_activity().silence_seconds, 1),
# Audio that was recorded but had no OP25 talkgroup anywhere near it, so
# it was discarded rather than uploaded. Non-zero means either the
# console is not decoding or something else is feeding drb_sink.
"unattributed_segments": metadata_watcher.unattributed_segments,
"active_tgid": active_tgid,
"active_tgid_name": active_tgid_name,
"active_call_id": metadata_watcher.active_call_id,
@@ -109,7 +131,7 @@ async def set_override(
):
node_cfg = load_node_config()
config = None
if system_id:
from app.internal.system_cacher import get_cached_system
cached = get_cached_system(system_id)
@@ -122,19 +144,19 @@ async def set_override(
config = SystemConfig(**system_config)
else:
raise HTTPException(400, "Must specify system_id or system_config.")
node_cfg.override_system_id = config.system_id
node_cfg.override_config = config
save_node_config(node_cfg)
from app.main import _generate_op25_config
if not await _generate_op25_config(config):
raise HTTPException(500, f"Failed to generate OP25 config for override: {config.name}")
await op25_client.stop()
await asyncio.sleep(2)
await op25_client.start()
await mqtt_manager._publish_checkin()
return {"ok": True}
@@ -144,20 +166,20 @@ async def revert_config():
node_cfg = load_node_config()
if not node_cfg.override_system_id:
return {"ok": True, "message": "No override active."}
node_cfg.override_system_id = None
node_cfg.override_config = None
save_node_config(node_cfg)
if node_cfg.system_config:
from app.main import _generate_op25_config
if not await _generate_op25_config(node_cfg.system_config):
raise HTTPException(500, "Failed to regenerate original OP25 config.")
await op25_client.stop()
await asyncio.sleep(2)
await op25_client.start()
await mqtt_manager._publish_checkin()
return {"ok": True}
@@ -166,10 +188,10 @@ async def revert_config():
async def ack_override(timeout_minutes: int = Body(1440)):
if not settings.c2_url:
raise HTTPException(400, "C2_URL not configured.")
api_key = credentials.get_api_key()
headers = {"Authorization": f"Bearer {api_key}"} if api_key else {}
try:
async with httpx.AsyncClient(timeout=10) as client:
r = await client.post(
@@ -183,6 +205,44 @@ async def ack_override(timeout_minutes: int = Body(1440)):
raise HTTPException(500, f"Failed to contact C2: {e}")
@router.get("/sdr")
async def get_sdr():
"""Which SDR does what: detected dongles, pins, priority, what's running."""
from app.internal.secondary_sdr_client import secondary_sdr_client
from app.internal.sdr_settings import op25_serial
cfg = load_node_config()
devs = await secondary_sdr_client.devices()
status = await secondary_sdr_client.status()
return {
"devices": devs, # None = secondary-sdr service unreachable
"op25_serial": op25_serial(cfg, devs),
"pins": cfg.sdr_pins,
"priority": cfg.secondary_sdr_priority,
"modes": list(SECONDARY_SDR_MODES),
"running": [r["mode"] for r in status.get("running", [])] if status else None,
}
class SdrSettingsBody(BaseModel):
priority: Optional[List[str]] = None
pins: Optional[Dict[str, Optional[str]]] = None # service -> serial; null/"" = automatic
@router.post("/sdr")
async def set_sdr(body: SdrSettingsBody):
"""Set priority and/or pins. Restarts OP25 only if OP25's own SDR changes."""
from app.internal.sdr_settings import set_sdr_settings
if body.priority is not None:
unknown = [m for m in body.priority if m not in SECONDARY_SDR_MODES]
if unknown:
raise HTTPException(400, f"Unknown secondary SDR mode(s): {unknown}")
try:
running = await set_sdr_settings(body.priority, body.pins)
except ValueError as e:
raise HTTPException(400, str(e))
return {"ok": True, "running": running}
@router.post("/discord/join")
async def discord_join(guild_id: int, channel_id: int):
ok = await radio_bot.join(guild_id, channel_id)
+50 -4
View File
@@ -1,18 +1,64 @@
from pathlib import Path
from fastapi import APIRouter
from fastapi.responses import HTMLResponse
from typing import Optional
from fastapi import APIRouter, Depends, Form
from fastapi.responses import HTMLResponse, RedirectResponse
from app.internal import auth
router = APIRouter(tags=["ui"])
_TEMPLATE = Path(__file__).parent.parent / "templates" / "index.html"
_SCANNER_TEMPLATE = Path(__file__).parent.parent / "templates" / "scanner.html"
_LOGIN_TEMPLATE = Path(__file__).parent.parent / "templates" / "login.html"
@router.get("/login", response_class=HTMLResponse)
async def login_page(error: Optional[str] = None):
html = _LOGIN_TEMPLATE.read_text()
banner = (
'<div class="error">Invalid username or password.</div>' if error else ""
)
return html.replace("<!--ERROR_BANNER-->", banner)
@router.post("/login")
async def login_submit(username: str = Form(...), password: str = Form(...)):
if not auth.verify_credentials(username, password):
return RedirectResponse("/login?error=1", status_code=303)
token = auth.create_session_token(username)
resp = RedirectResponse("/", status_code=303)
resp.set_cookie(
auth.SESSION_COOKIE_NAME,
token,
max_age=auth.SESSION_TTL_SECONDS,
httponly=True,
samesite="lax",
# No TLS termination on this port (LAN dashboard on :80) — `secure`
# would make the cookie never get sent at all.
secure=False,
)
return resp
@router.post("/logout")
@router.get("/logout")
async def logout():
resp = RedirectResponse("/login", status_code=303)
resp.delete_cookie(auth.SESSION_COOKIE_NAME)
return resp
@router.get("/", response_class=HTMLResponse)
async def index():
async def index(authed: bool = Depends(auth.require_session)):
if not authed:
return RedirectResponse("/login")
return _TEMPLATE.read_text()
@router.get("/scanner", response_class=HTMLResponse)
async def scanner():
async def scanner(authed: bool = Depends(auth.require_session)):
if not authed:
return RedirectResponse("/login")
return _SCANNER_TEMPLATE.read_text()
+187
View File
@@ -105,6 +105,40 @@
background: rgba(255, 255, 255, 0.1);
}
.sdr-row {
display: flex;
align-items: center;
gap: 0.6rem;
padding: 0.55rem 0;
border-bottom: 1px solid var(--glass-border);
}
.sdr-row:last-child { border-bottom: none; }
.sdr-rank { width: 1.2rem; color: var(--text-muted); font-size: 0.8rem; text-align: right; }
.sdr-name { flex: 1; }
.sdr-name small { display: block; color: var(--text-muted); font-size: 0.75rem; }
.sdr-move {
background: rgba(255, 255, 255, 0.05);
border: 1px solid var(--glass-border);
color: var(--text-main);
border-radius: 6px;
width: 1.9rem;
height: 1.9rem;
cursor: pointer;
}
.sdr-move:disabled { opacity: 0.3; cursor: default; }
.sdr-state { font-size: 0.75rem; min-width: 6.5rem; text-align: right; color: var(--text-muted); }
.sdr-state.on { color: var(--success); }
.sdr-select {
background: rgba(255, 255, 255, 0.05);
border: 1px solid var(--glass-border);
color: var(--text-main);
border-radius: 6px;
padding: 0.3rem 0.4rem;
font-size: 0.8rem;
max-width: 13rem;
}
.sdr-select option { background: #111827; }
.grid {
display: grid;
grid-template-columns: repeat(auto-fit, minmax(320px, 1fr));
@@ -268,6 +302,10 @@
<svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" style="margin-right:8px"><rect x="2" y="3" width="20" height="14" rx="2" ry="2"></rect><line x1="8" y1="21" x2="16" y2="21"></line><line x1="12" y1="17" x2="12" y2="21"></line></svg>
Scanner Mode
</a>
<a href="/logout" class="btn btn-secondary">
<svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" style="margin-right:8px"><path d="M9 21H5a2 2 0 0 1-2-2V5a2 2 0 0 1 2-2h4"></path><polyline points="16 17 21 12 16 7"></polyline><line x1="21" y1="12" x2="9" y2="12"></line></svg>
Logout
</a>
</div>
</header>
@@ -356,6 +394,28 @@
</div>
</div>
<!-- SDRs Card (node-26#9, #11) -->
<div class="glass-card">
<div class="card-header">
<div class="card-title">SDRs</div>
</div>
<div class="data-row" style="align-items:center;">
<span class="data-label">OP25 SDR</span>
<select id="sdr-op25" class="sdr-select" aria-label="OP25 SDR"></select>
</div>
<p id="sdr-op25-note" style="color: var(--text-muted); font-size: 0.75rem; margin: 0.25rem 0 0.75rem;"></p>
<p style="color: var(--text-muted); font-size: 0.8rem; margin: 0 0 0.5rem;">
Every other SDR runs the next enabled service, top first. Pin a service to the SDR
that has its antenna, or leave it on "Any spare SDR".
</p>
<div id="sdr-list"></div>
<p id="sdr-warning" style="color: var(--danger); font-size: 0.8rem; margin: 0.5rem 0 0; display:none;"></p>
<div style="display:flex; gap:0.75rem; align-items:center; margin-top:0.75rem;">
<button id="sdr-save" class="btn btn-primary" style="padding:0.5rem 1rem;" onclick="saveSdr()" disabled>Save</button>
<span id="sdr-msg" style="color: var(--text-muted); font-size: 0.8rem;"></span>
</div>
</div>
<!-- Local Listening Card -->
<div class="glass-card player-card">
<div class="card-header" style="margin-bottom:0.5rem; border:none;">
@@ -488,6 +548,133 @@
}
}
// ── SDRs: OP25's dongle, secondary priority and per-service pins ────────
const SDR_LABELS = {
adsb: ['ADS-B', 'Aircraft · 1090 MHz'],
ais: ['AIS', 'Vessels · 162 MHz'],
};
let sdr = null; // last GET /api/sdr
let sdrRows = []; // [{mode, enabled, pin}] in display order
let sdrOp25Pin = ''; // '' = automatic
let sdrDirty = false;
function esc(v) {
return String(v ?? '').replace(/[&<>"']/g, c => ({'&': '&amp;', '<': '&lt;', '>': '&gt;', '"': '&quot;', "'": '&#39;'}[c]));
}
function deviceOptions(selected, autoLabel) {
const devs = sdr.devices || [];
const known = devs.some(d => d.serial === selected);
return `<option value="">${esc(autoLabel)}</option>` +
devs.map(d => `<option value="${esc(d.serial)}" ${d.serial === selected ? 'selected' : ''}>` +
`SDR ${d.index + 1} · serial ${esc(d.serial)}${d.duplicate_serial ? ' (shared serial!)' : ''}</option>`).join('') +
(selected && !known ? `<option value="${esc(selected)}" selected>serial ${esc(selected)} (not plugged in)</option>` : '');
}
function renderSdr() {
const op25 = document.getElementById('sdr-op25');
op25.innerHTML = deviceOptions(sdrOp25Pin, 'Automatic (first SDR)');
op25.onchange = () => { sdrOp25Pin = op25.value; markSdrDirty(); };
document.getElementById('sdr-op25-note').textContent = sdrOp25Pin !== (sdr.pins.op25 || '')
? 'Saving will restart OP25 on the selected SDR.'
: (sdr.op25_serial ? `OP25 is using serial ${sdr.op25_serial}.` : '');
const list = document.getElementById('sdr-list');
list.innerHTML = '';
let rank = 0;
const running = sdr.running || [];
sdrRows.forEach((row, i) => {
const [name, hint] = SDR_LABELS[row.mode] || [row.mode, ''];
const pinMissing = row.pin && !(sdr.devices || []).some(d => d.serial === row.pin);
const state = !row.enabled ? 'Off'
: sdrDirty ? 'Unsaved'
: sdr.running === null ? 'Unknown'
: running.includes(row.mode) ? 'Running'
: pinMissing ? 'Pinned SDR missing' : 'Waiting for SDR';
const el = document.createElement('div');
el.className = 'sdr-row';
el.innerHTML = `
<span class="sdr-rank">${row.enabled ? ++rank : ''}</span>
<input type="checkbox" ${row.enabled ? 'checked' : ''} aria-label="Enable ${name}">
<span class="sdr-name">${name}<small>${hint}</small></span>
<select class="sdr-select" aria-label="${name} SDR">${deviceOptions(row.pin, 'Any spare SDR')}</select>
<button class="sdr-move" aria-label="Move ${name} up" ${i === 0 ? 'disabled' : ''}>▲</button>
<button class="sdr-move" aria-label="Move ${name} down" ${i === sdrRows.length - 1 ? 'disabled' : ''}>▼</button>
<span class="sdr-state ${state === 'Running' ? 'on' : ''}">${state}</span>`;
const [box] = el.getElementsByTagName('input');
const [pick] = el.getElementsByTagName('select');
const [up, down] = el.getElementsByTagName('button');
box.onchange = () => { row.enabled = box.checked; markSdrDirty(); };
pick.onchange = () => { row.pin = pick.value; markSdrDirty(); };
up.onclick = () => { [sdrRows[i - 1], sdrRows[i]] = [sdrRows[i], sdrRows[i - 1]]; markSdrDirty(); };
down.onclick = () => { [sdrRows[i + 1], sdrRows[i]] = [sdrRows[i], sdrRows[i + 1]]; markSdrDirty(); };
list.appendChild(el);
});
const warn = document.getElementById('sdr-warning');
const pinned = [sdrOp25Pin, ...sdrRows.map(r => r.pin)].filter(Boolean);
const problems = [];
if (sdr.devices === null) problems.push('The secondary SDR service is not responding, so SDRs can\'t be listed.');
if ((sdr.devices || []).some(d => d.duplicate_serial)) {
problems.push('Two SDRs share a serial number, so they can\'t be told apart. Give each a unique serial (rtl_eeprom -s) before pinning.');
}
if (new Set(pinned).size !== pinned.length) problems.push('Two services are pinned to the same SDR.');
warn.textContent = problems.join(' ');
warn.style.display = problems.length ? 'block' : 'none';
document.getElementById('sdr-save').disabled = !sdrDirty || new Set(pinned).size !== pinned.length;
}
function markSdrDirty() {
sdrDirty = true;
document.getElementById('sdr-msg').textContent = '';
renderSdr();
}
async function loadSdr() {
if (sdrDirty) return; // never clobber unsaved edits with a poll
try {
const r = await fetch('/api/sdr');
if (!r.ok) return;
sdr = await r.json();
sdrOp25Pin = sdr.pins.op25 || '';
sdrRows = [
...sdr.priority.map(mode => ({ mode, enabled: true, pin: sdr.pins[mode] || '' })),
...sdr.modes.filter(m => !sdr.priority.includes(m)).map(mode => ({ mode, enabled: false, pin: sdr.pins[mode] || '' })),
];
renderSdr();
} catch (e) {
console.error('SDR load failed:', e);
}
}
async function saveSdr() {
const pins = { op25: sdrOp25Pin || null };
sdrRows.forEach(r => { pins[r.mode] = r.pin || null; });
const btn = document.getElementById('sdr-save');
const msg = document.getElementById('sdr-msg');
btn.disabled = true;
msg.textContent = 'Applying…';
try {
const r = await fetch('/api/sdr', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ priority: sdrRows.filter(r => r.enabled).map(r => r.mode), pins }),
});
if (!r.ok) throw new Error((await r.json().catch(() => ({}))).detail || r.statusText);
const d = await r.json();
sdrDirty = false;
msg.textContent = d.running === null ? 'Saved, but the secondary SDR service did not respond.' : 'Saved.';
await loadSdr();
} catch (e) {
console.error('SDR save failed:', e);
msg.textContent = `Save failed: ${e.message}`;
btn.disabled = false;
}
}
loadSdr();
setInterval(loadSdr, 10000);
refresh();
setInterval(refresh, 2000); // Polling every 2 seconds
</script>
+133
View File
@@ -0,0 +1,133 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>DRB Edge Node — Login</title>
<link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700;800&family=JetBrains+Mono:wght@400;700&display=swap" rel="stylesheet">
<style>
:root {
--bg: #0b0f19;
--glass-bg: rgba(20, 25, 40, 0.6);
--glass-border: rgba(255, 255, 255, 0.08);
--accent: #3b82f6;
--accent-hover: #2563eb;
--danger: #ef4444;
--text-main: #f8fafc;
--text-muted: #94a3b8;
}
*, *::before, *::after {
box-sizing: border-box;
margin: 0;
padding: 0;
}
body {
font-family: 'Inter', sans-serif;
background: var(--bg);
background-image:
radial-gradient(circle at 15% 50%, rgba(59, 130, 246, 0.15), transparent 25%),
radial-gradient(circle at 85% 30%, rgba(139, 92, 246, 0.15), transparent 25%);
color: var(--text-main);
min-height: 100vh;
display: flex;
align-items: center;
justify-content: center;
padding: 1.5rem;
}
.login-card {
width: 100%;
max-width: 360px;
background: var(--glass-bg);
backdrop-filter: blur(12px);
-webkit-backdrop-filter: blur(12px);
border: 1px solid var(--glass-border);
border-radius: 16px;
padding: 2rem;
box-shadow: 0 4px 6px -1px rgba(0, 0, 0, 0.1), 0 2px 4px -1px rgba(0, 0, 0, 0.06);
}
h1 {
font-size: 1.5rem;
font-weight: 800;
margin-bottom: 0.25rem;
}
.subtitle {
color: var(--text-muted);
font-family: 'JetBrains Mono', monospace;
font-size: 0.8rem;
margin-bottom: 1.5rem;
}
label {
display: block;
font-size: 0.8rem;
color: var(--text-muted);
margin-bottom: 0.35rem;
margin-top: 1rem;
}
input[type="text"], input[type="password"] {
width: 100%;
padding: 0.65rem 0.75rem;
border-radius: 8px;
border: 1px solid var(--glass-border);
background: rgba(255, 255, 255, 0.05);
color: var(--text-main);
font-family: inherit;
font-size: 0.95rem;
}
input[type="text"]:focus, input[type="password"]:focus {
outline: none;
border-color: var(--accent);
}
button {
width: 100%;
margin-top: 1.5rem;
padding: 0.75rem 1.5rem;
border-radius: 8px;
font-weight: 600;
font-size: 0.9rem;
border: none;
cursor: pointer;
background: var(--accent);
color: white;
box-shadow: 0 4px 14px 0 rgba(59, 130, 246, 0.39);
transition: all 0.2s ease;
}
button:hover {
background: var(--accent-hover);
}
.error {
margin-top: 1rem;
padding: 0.6rem 0.8rem;
border-radius: 8px;
background: rgba(239, 68, 68, 0.15);
border: 1px solid rgba(239, 68, 68, 0.2);
color: var(--danger);
font-size: 0.85rem;
}
</style>
</head>
<body>
<div class="login-card">
<h1>DRB Edge Node</h1>
<p class="subtitle">Sign in to the local dashboard</p>
<form method="post" action="/login">
<label for="username">Username</label>
<input type="text" id="username" name="username" autocomplete="username" required autofocus>
<label for="password">Password</label>
<input type="password" id="password" name="password" autocomplete="current-password" required>
<button type="submit">Sign in</button>
</form>
<!--ERROR_BANNER-->
</div>
</body>
</html>
+1
View File
@@ -5,5 +5,6 @@ paho-mqtt>=2.0.0
httpx
discord.py[voice]
PyNaCl
python-multipart
pytest
pytest-asyncio
+218
View File
@@ -0,0 +1,218 @@
"""
Unit tests for silence trimming, now a byte-offset slice of raw PCM.
`keep_window` is pure on purpose so the "what do we keep" decision — the part
that can destroy a transmission if it is wrong — stays testable without any
audio at all. The rest of the file drives the real detector over synthesised
buffers shaped like the six real recordings measured off a live P25 node:
1.71-2.45 s of leading silence and 0.00-1.11 s trailing.
The old implementation shelled out to FFmpeg twice (silencedetect, then a
re-encode) and these tests parsed its stderr. Both passes are gone; the recorder
buffers PCM, so detection is arithmetic and the cut is a slice.
"""
from array import array
import pytest
from app.config import settings
from app.internal import audio_trim, pcm
from app.internal.audio_trim import (
TrimResult,
first_signal_offset,
keep_window,
last_signal_offset,
trim_pcm,
)
GUARD = 0.25
SPEECH_LEVEL = 4096 # -18 dBFS, the measured field average
FLOOR_LEVEL = 1 # -90.3 dBFS, the measured digital-silence floor
def speech(seconds: float) -> bytes:
count = int(pcm.SAMPLE_RATE * seconds)
return array("h", [SPEECH_LEVEL, -SPEECH_LEVEL] * (count // 2)).tobytes()
def silence(seconds: float, level: int = FLOOR_LEVEL) -> bytes:
count = int(pcm.SAMPLE_RATE * seconds)
return array("h", [level, -level] * (count // 2)).tobytes()
# ---------------------------------------------------------------------------
# keep_window — the pure decision
# ---------------------------------------------------------------------------
def test_guard_margin_is_kept_around_detected_speech():
guard = pcm.byte_offset(GUARD)
start, end = keep_window(
first_signal=pcm.byte_offset(1.85),
last_signal=pcm.byte_offset(4.20),
total_bytes=pcm.byte_offset(4.54),
guard_bytes=guard,
)
assert pcm.seconds(start) == pytest.approx(1.85 - GUARD, abs=0.001)
assert pcm.seconds(end) == pytest.approx(4.20 + GUARD, abs=0.001)
def test_guard_margin_never_runs_past_the_buffer_bounds():
total = pcm.byte_offset(4.0)
start, end = keep_window(
first_signal=pcm.byte_offset(0.10),
last_signal=pcm.byte_offset(3.95),
total_bytes=total,
guard_bytes=pcm.byte_offset(1.0),
)
assert (start, end) == (0, total)
def test_keep_window_offsets_are_sample_aligned():
start, end = keep_window(3, 9, 21, 1)
assert start % pcm.FRAME_BYTES == 0
assert end % pcm.FRAME_BYTES == 0
def test_no_signal_found_keeps_everything():
total = pcm.byte_offset(6.0)
assert keep_window(None, None, total, pcm.byte_offset(GUARD)) == (0, total)
def test_an_inverted_window_degrades_to_keeping_everything():
"""Never return an empty slice, whatever the inputs say."""
total = pcm.byte_offset(4.0)
assert keep_window(pcm.byte_offset(3.0), pcm.byte_offset(0.5), total, 0) == (0, total)
# ---------------------------------------------------------------------------
# Scanning
# ---------------------------------------------------------------------------
def test_first_and_last_signal_are_found_in_a_realistic_recording():
audio = silence(1.85) + speech(2.35) + silence(0.34)
threshold = -40.0
first = first_signal_offset(audio, threshold)
last = last_signal_offset(audio, threshold)
assert pcm.seconds(first) == pytest.approx(1.85, abs=audio_trim.ANALYSIS_WINDOW_SECONDS)
assert pcm.seconds(last) == pytest.approx(4.20, abs=audio_trim.ANALYSIS_WINDOW_SECONDS)
def test_internal_pauses_are_not_treated_as_the_tail():
"""Trimming the middle out of a conversation would be unrecoverable."""
audio = silence(1.9) + speech(3.1) + silence(2.5) + speech(4.5)
last = last_signal_offset(audio, -40.0)
assert pcm.seconds(last) == pytest.approx(12.0, abs=audio_trim.ANALYSIS_WINDOW_SECONDS)
def test_all_silence_returns_no_signal_offset():
assert first_signal_offset(silence(4.0), -40.0) is None
assert last_signal_offset(silence(4.0), -40.0) is None
def test_the_scan_is_bounded_so_a_long_buffer_cannot_stall_the_upload():
"""The per-sample loop is the only unbounded cost; it must have a ceiling."""
audio = silence(2.0)
assert first_signal_offset(audio, -40.0, limit_seconds=0.5) is None
assert first_signal_offset(speech(0.1) + silence(1.9), -40.0, limit_seconds=0.5) == 0
# ---------------------------------------------------------------------------
# trim_pcm — end to end over synthesised audio
# ---------------------------------------------------------------------------
def test_leading_and_trailing_silence_are_trimmed_to_the_guard_margin():
audio = silence(1.85) + speech(2.35) + silence(0.34)
kept, result = trim_pcm(audio, threshold_db=-40.0, guard=GUARD)
assert result.applied and not result.all_silence
assert result.lead == pytest.approx(1.85 - GUARD, abs=0.05)
assert result.tail == pytest.approx(0.34 - GUARD, abs=0.05)
assert pcm.seconds(len(kept)) == pytest.approx(result.duration_after, abs=0.001)
assert result.duration_after < result.duration_before
def test_the_guard_margin_never_eats_into_speech():
audio = silence(2.0) + speech(1.0) + silence(2.0)
kept, result = trim_pcm(audio, threshold_db=-40.0, guard=GUARD)
# Everything removed from the head must be silence, and the first sample of
# real speech must survive.
assert result.lead < 2.0
assert pcm.seconds(len(kept)) > 1.0
def test_measured_trailing_silence_is_reported_for_field_tuning():
"""
The recorder deliberately over-captures the tail (it closes only after the
silence timeout has actually elapsed in the audio), so `tail` is how the
real silence run reaches the logs.
"""
audio = silence(0.5) + speech(2.0) + silence(3.0)
_, result = trim_pcm(audio, threshold_db=-40.0, guard=GUARD)
assert result.tail == pytest.approx(3.0 - GUARD, abs=0.05)
assert result.trimmed_seconds == pytest.approx(result.lead + result.tail)
def test_an_all_silence_buffer_is_reported_not_truncated_to_nothing():
audio = silence(4.0)
kept, result = trim_pcm(audio, threshold_db=-40.0, guard=GUARD)
assert result.all_silence
assert not result.applied
assert kept == audio, "an all-silence recording must not become zero-length"
def test_digital_silence_at_the_measured_field_floor_is_detected():
"""
The -91 dBFS floor is the whole reason this needs no field calibration.
Detection must not depend on the threshold being tuned to a noise floor.
"""
audio = silence(1.0, level=1) + speech(1.0) + silence(1.0, level=1)
for threshold in (-70.0, -60.0, -50.0, -40.0):
_, result = trim_pcm(audio, threshold_db=threshold, guard=GUARD)
assert result.applied, f"threshold {threshold} should still find the speech"
assert result.lead == pytest.approx(0.75, abs=0.05)
def test_audio_with_no_silence_at_either_end_is_left_alone():
audio = speech(3.0)
kept, result = trim_pcm(audio, threshold_db=-40.0, guard=GUARD)
assert not result.applied
assert kept == audio
assert result.duration_before == pytest.approx(result.duration_after)
def test_an_empty_buffer_is_handled():
kept, result = trim_pcm(b"", threshold_db=-40.0, guard=GUARD)
assert kept == b"" and not result.applied and not result.all_silence
def test_thresholds_default_to_settings():
audio = silence(1.0) + speech(1.0) + silence(1.0)
_, result = trim_pcm(audio)
assert result.applied
assert settings.trim_silence_threshold_db == -40.0
assert settings.trim_silence_guard_seconds == 0.25
assert result.lead == pytest.approx(1.0 - settings.trim_silence_guard_seconds, abs=0.05)
def test_a_scan_that_gives_up_leaves_the_audio_untouched_and_says_so():
"""
Refusing to guess is the point: an untrimmed upload is always better than a
wrongly-truncated one, and better than dropping a call as "all silence"
without having actually looked at all of it.
"""
long_silence = silence(audio_trim.MAX_SCAN_SECONDS + 5.0)
kept, result = trim_pcm(long_silence, threshold_db=-40.0, guard=GUARD)
assert result.scan_truncated
assert not result.all_silence
assert not result.applied
assert kept == long_silence
def test_trim_result_reports_total_trimmed():
assert TrimResult(lead=1.9, tail=0.35).trimmed_seconds == pytest.approx(2.25)
+248
View File
@@ -0,0 +1,248 @@
"""
Unit tests for local dashboard/API auth (app.internal.auth), plus the
credentials.py additions that persist its signing material (auth_salt,
session_secret) alongside the existing node API key.
This file is pure logic: password hashing/constant-time comparison, session
token signing/expiry, and HTTP Basic header parsing. See test_auth_endpoints.py
for the HTTP-level login/redirect/protection round trip through the routers.
"""
import base64
import secrets
import time
from unittest.mock import patch
import pytest
from fastapi import HTTPException
from app.config import settings
from app.internal import auth, credentials
@pytest.fixture(autouse=True)
def isolated_credentials(tmp_path, monkeypatch):
"""Every test gets a fresh, on-disk-isolated credentials store so the
generated auth salt / session secret never leak between tests, and a
known username/password instead of the shipped default."""
creds_file = tmp_path / "credentials.json"
monkeypatch.setattr(credentials, "_CREDS_FILE", creds_file)
monkeypatch.setattr(credentials, "_api_key", None)
monkeypatch.setattr(credentials, "_auth_salt", None)
monkeypatch.setattr(credentials, "_session_secret", None)
monkeypatch.setattr(settings, "dashboard_username", "tester")
monkeypatch.setattr(settings, "dashboard_password", "s3cret-pass")
yield
def _basic_header(username: str, password: str) -> str:
encoded = base64.b64encode(f"{username}:{password}".encode()).decode()
return f"Basic {encoded}"
# ---------------------------------------------------------------------------
# credentials.py: auth salt / session secret generation + persistence
# ---------------------------------------------------------------------------
def test_auth_material_is_generated_on_first_access():
salt = credentials.get_auth_salt()
secret = credentials.get_session_secret()
assert isinstance(salt, bytes) and len(salt) == 16
assert isinstance(secret, bytes) and len(secret) == 32
def test_auth_material_is_stable_across_repeated_calls():
assert credentials.get_auth_salt() == credentials.get_auth_salt()
assert credentials.get_session_secret() == credentials.get_session_secret()
def test_auth_material_persists_to_disk_and_survives_reload():
salt = credentials.get_auth_salt()
secret = credentials.get_session_secret()
# Simulate a container restart: drop in-memory state, reload from disk.
credentials._api_key = None
credentials._auth_salt = None
credentials._session_secret = None
credentials.load()
assert credentials.get_auth_salt() == salt
assert credentials.get_session_secret() == secret
def test_saving_api_key_does_not_clobber_auth_material():
"""save_api_key() used to json.dumps({"api_key": key}) directly, which
would have wiped auth_salt/session_secret out of credentials.json the
moment C2 provisioned an API key after this feature was added."""
salt = credentials.get_auth_salt()
secret = credentials.get_session_secret()
credentials.save_api_key("some-node-api-key")
assert credentials.get_api_key() == "some-node-api-key"
assert credentials.get_auth_salt() == salt
assert credentials.get_session_secret() == secret
# ---------------------------------------------------------------------------
# verify_credentials() — password hashing + constant-time compare
# ---------------------------------------------------------------------------
def test_verify_credentials_accepts_correct_username_and_password():
assert auth.verify_credentials("tester", "s3cret-pass") is True
def test_verify_credentials_rejects_wrong_password():
assert auth.verify_credentials("tester", "wrong") is False
def test_verify_credentials_rejects_wrong_username():
assert auth.verify_credentials("someone-else", "s3cret-pass") is False
def test_verify_credentials_rejects_empty_password():
assert auth.verify_credentials("tester", "") is False
def test_password_is_hashed_not_compared_in_plaintext():
with patch.object(auth, "_hash_password", wraps=auth._hash_password) as spy:
auth.verify_credentials("tester", "s3cret-pass")
# Once for the configured password, once for the submitted one — neither
# side is ever compared as a raw string.
assert spy.call_count == 2
def test_is_using_default_password_detects_the_shipped_default(monkeypatch):
monkeypatch.setattr(settings, "dashboard_password", auth.DEFAULT_PASSWORD)
assert auth.is_using_default_password() is True
def test_is_using_default_password_false_once_changed():
assert auth.is_using_default_password() is False # fixture already changed it
# ---------------------------------------------------------------------------
# session tokens
# ---------------------------------------------------------------------------
def test_session_token_round_trips():
token = auth.create_session_token("tester")
assert auth._verify_session_token(token) == "tester"
def test_session_token_rejects_tampered_payload():
token = auth.create_session_token("tester")
tampered = ("X" if token[0] != "X" else "Y") + token[1:]
assert auth._verify_session_token(tampered) is None
def test_session_token_rejects_expired_token(monkeypatch):
token = auth.create_session_token("tester")
future = time.time() + auth.SESSION_TTL_SECONDS + 1
monkeypatch.setattr(time, "time", lambda: future)
assert auth._verify_session_token(token) is None
def test_session_token_rejects_username_mismatch(monkeypatch):
token = auth.create_session_token("tester")
monkeypatch.setattr(settings, "dashboard_username", "someone-else")
assert auth._verify_session_token(token) is None
def test_session_token_garbage_input_does_not_raise():
assert auth._verify_session_token("not-a-real-token") is None
assert auth._verify_session_token("") is None
def test_session_token_signed_with_a_different_secret_is_rejected():
token = auth.create_session_token("tester")
# As if the node restarted without a persisted credentials.json.
credentials._session_secret = secrets.token_bytes(32)
assert auth._verify_session_token(token) is None
# ---------------------------------------------------------------------------
# HTTP Basic parsing
# ---------------------------------------------------------------------------
def test_basic_auth_accepts_valid_header():
assert auth._verify_basic_auth(_basic_header("tester", "s3cret-pass")) is True
def test_basic_auth_rejects_wrong_credentials():
assert auth._verify_basic_auth(_basic_header("tester", "wrong")) is False
def test_basic_auth_rejects_non_basic_scheme():
assert auth._verify_basic_auth("Bearer sometoken") is False
def test_basic_auth_tolerates_garbage_without_raising():
assert auth._verify_basic_auth("Basic not-valid-base64!!") is False
assert auth._verify_basic_auth("") is False
# ---------------------------------------------------------------------------
# is_authenticated() — the combined check require_auth is built on
# ---------------------------------------------------------------------------
def test_is_authenticated_true_with_valid_session_cookie():
token = auth.create_session_token("tester")
assert auth.is_authenticated(token, None) is True
def test_is_authenticated_true_with_valid_basic_header():
assert auth.is_authenticated(None, _basic_header("tester", "s3cret-pass")) is True
def test_is_authenticated_false_with_neither():
assert auth.is_authenticated(None, None) is False
def test_is_authenticated_false_with_invalid_session_and_no_header():
assert auth.is_authenticated("garbage", None) is False
# ---------------------------------------------------------------------------
# FastAPI dependencies: require_session / require_auth
# ---------------------------------------------------------------------------
async def test_require_session_false_with_no_cookie():
assert await auth.require_session(None) is False
async def test_require_session_true_with_valid_cookie():
token = auth.create_session_token("tester")
assert await auth.require_session(token) is True
async def test_require_auth_raises_401_with_no_credentials():
with pytest.raises(HTTPException) as exc_info:
await auth.require_auth(None, None)
assert exc_info.value.status_code == 401
assert exc_info.value.headers["WWW-Authenticate"] == "Basic"
async def test_require_auth_passes_with_valid_session_cookie():
token = auth.create_session_token("tester")
await auth.require_auth(token, None) # must not raise
async def test_require_auth_passes_with_valid_basic_header():
await auth.require_auth(None, _basic_header("tester", "s3cret-pass")) # must not raise
# ---------------------------------------------------------------------------
# startup warning
# ---------------------------------------------------------------------------
def test_warn_if_default_password_logs_when_default(monkeypatch):
monkeypatch.setattr(settings, "dashboard_password", auth.DEFAULT_PASSWORD)
with patch("app.internal.auth.logger") as mock_logger:
auth.warn_if_default_password()
mock_logger.warning.assert_called_once()
def test_warn_if_default_password_silent_once_changed():
with patch("app.internal.auth.logger") as mock_logger:
auth.warn_if_default_password()
mock_logger.warning.assert_not_called()
+157
View File
@@ -0,0 +1,157 @@
"""
HTTP-level tests for the auth-protected dashboard/API surface: login/logout
flow, session-cookie protection of the HTML pages, and Basic-auth protection
of the JSON API.
Built as a standalone FastAPI app assembling the real api/ui routers — NOT
app.main:app, which wires a lifespan that connects to MQTT, starts the
PulseAudio capture loop, and pings OP25/C2. None of that belongs in a unit
test, and none of it is needed to exercise the auth layer: the auth
dependency runs (and short-circuits with a redirect/401) before any route
body that would touch those services.
"""
import base64
import pytest
from fastapi import FastAPI
from fastapi.testclient import TestClient
from app.config import settings
from app.internal import auth, credentials
from app.routers import api, ui
@pytest.fixture(autouse=True)
def isolated_credentials(tmp_path, monkeypatch):
creds_file = tmp_path / "credentials.json"
monkeypatch.setattr(credentials, "_CREDS_FILE", creds_file)
monkeypatch.setattr(credentials, "_api_key", None)
monkeypatch.setattr(credentials, "_auth_salt", None)
monkeypatch.setattr(credentials, "_session_secret", None)
monkeypatch.setattr(settings, "dashboard_username", "tester")
monkeypatch.setattr(settings, "dashboard_password", "s3cret-pass")
yield
@pytest.fixture
def client():
test_app = FastAPI()
test_app.include_router(api.router)
test_app.include_router(ui.router)
with TestClient(test_app) as c:
yield c
def _basic_header(username: str, password: str) -> dict:
encoded = base64.b64encode(f"{username}:{password}".encode()).decode()
return {"Authorization": f"Basic {encoded}"}
# ---------------------------------------------------------------------------
# /api/* — machine-facing JSON API
# ---------------------------------------------------------------------------
def test_api_route_rejects_unauthenticated_requests(client):
r = client.get("/api/status", follow_redirects=False)
assert r.status_code == 401
assert r.headers["www-authenticate"] == "Basic"
def test_api_route_accepts_valid_basic_auth(client):
r = client.get("/api/config", headers=_basic_header("tester", "s3cret-pass"))
assert r.status_code == 200
def test_api_route_rejects_wrong_basic_auth_password(client):
r = client.get("/api/config", headers=_basic_header("tester", "wrong"))
assert r.status_code == 401
def test_api_route_accepts_dashboard_session_cookie(client):
login = client.post(
"/login", data={"username": "tester", "password": "s3cret-pass"}, follow_redirects=False
)
assert login.status_code == 303
assert auth.SESSION_COOKIE_NAME in login.cookies
r = client.get("/api/config") # cookie jar carries the session cookie
assert r.status_code == 200
def test_every_api_route_is_registered_behind_the_auth_dependency():
"""Structural guard: catches a future route added to api.py that forgets
the router is meant to protect everything in it."""
assert any(
getattr(dep, "dependency", None) is auth.require_auth
for dep in api.router.dependencies
)
# ---------------------------------------------------------------------------
# / and /scanner — the HTML dashboard
# ---------------------------------------------------------------------------
def test_index_redirects_to_login_when_unauthenticated(client):
r = client.get("/", follow_redirects=False)
assert r.status_code in (302, 307)
assert r.headers["location"] == "/login"
def test_scanner_redirects_to_login_when_unauthenticated(client):
r = client.get("/scanner", follow_redirects=False)
assert r.status_code in (302, 307)
assert r.headers["location"] == "/login"
def test_index_served_with_a_valid_session_cookie(client):
client.post("/login", data={"username": "tester", "password": "s3cret-pass"})
r = client.get("/")
assert r.status_code == 200
assert "text/html" in r.headers["content-type"]
# ---------------------------------------------------------------------------
# /login, /logout
# ---------------------------------------------------------------------------
def test_login_page_loads_without_auth(client):
r = client.get("/login")
assert r.status_code == 200
def test_login_with_correct_credentials_sets_cookie_and_redirects_home(client):
r = client.post(
"/login", data={"username": "tester", "password": "s3cret-pass"}, follow_redirects=False
)
assert r.status_code == 303
assert r.headers["location"] == "/"
cookie = r.cookies.get(auth.SESSION_COOKIE_NAME)
assert cookie
assert auth._verify_session_token(cookie) == "tester"
def test_login_with_wrong_password_redirects_back_with_error_and_no_cookie(client):
r = client.post(
"/login", data={"username": "tester", "password": "wrong"}, follow_redirects=False
)
assert r.status_code == 303
assert r.headers["location"] == "/login?error=1"
assert auth.SESSION_COOKIE_NAME not in r.cookies
def test_login_error_banner_renders_on_the_login_page(client):
r = client.get("/login?error=1")
assert r.status_code == 200
assert "Invalid username or password" in r.text
def test_logout_clears_the_session_cookie_and_redirects_to_login(client):
client.post("/login", data={"username": "tester", "password": "s3cret-pass"})
assert client.get("/").status_code == 200 # confirm we were logged in
r = client.get("/logout", follow_redirects=False)
assert r.status_code == 303
assert r.headers["location"] == "/login"
r2 = client.get("/", follow_redirects=False)
assert r2.status_code in (302, 307) # session cookie was cleared
+644
View File
@@ -0,0 +1,644 @@
"""
Unit tests for the CallRecorder: PCM ring buffer, per-call accumulator, the
continuous voice-activity signal the segmenter reads, and the single encode.
No FFmpeg and no PulseAudio: chunks are pushed through _ingest() with a patched
clock, which is exactly what the capture loop does at runtime, and the MP3
encoder is replaced with a stub that writes the raw PCM it was handed. That stub
is also how "exactly one encode per call" is asserted — the old design captured
MP3 and then re-encoded it to trim, so every upload was double-encoded.
"""
import asyncio
import itertools
import time
from array import array
from typing import List
from unittest.mock import patch
import pytest
from app.config import settings
from app.internal import call_recorder as recorder_mod
from app.internal import pcm
from app.internal.call_recorder import (
CallRecorder,
MAX_RECORDING_BYTES,
MAX_RECORDING_SECONDS,
PRE_ROLL_SECONDS,
RING_BUFFER_SECONDS,
)
T0 = 1_700_000_000.0
# One synthetic chunk carries exactly CHUNK_INTERVAL seconds of audio AND
# arrives CHUNK_INTERVAL apart, so arrival-timestamp arithmetic (slicing) and
# byte-offset arithmetic (trimming) agree with each other.
CHUNK_INTERVAL = 0.1
CHUNK_SAMPLES = int(pcm.SAMPLE_RATE * CHUNK_INTERVAL)
CHUNK_BYTES = CHUNK_SAMPLES * pcm.FRAME_BYTES
SPEECH_LEVEL = 4096 # -18 dBFS, the measured field average
FLOOR_LEVEL = 1 # -90.3 dBFS, the measured digital-silence floor
def block(level: int, samples: int = CHUNK_SAMPLES) -> bytes:
return array("h", [level, -level] * (samples // 2)).tobytes()
VOICE = block(SPEECH_LEVEL)
QUIET = block(FLOOR_LEVEL)
@pytest.fixture
def encodes(monkeypatch):
"""Replace the one encode with a stub that writes the PCM it was given."""
calls: List[tuple] = []
async def _encode(audio: bytes, path):
calls.append((audio, path))
path.write_bytes(audio)
return True
monkeypatch.setattr(recorder_mod, "encode_recording", _encode)
return calls
@pytest.fixture
def recorder(tmp_path, monkeypatch, encodes):
monkeypatch.setattr(settings, "trim_silence", False)
r = CallRecorder()
r._recordings_dir = tmp_path
r._capturing = True
return r
def ingest(recorder, start: float, end: float, chunk: bytes = VOICE) -> None:
"""Feed one chunk every CHUNK_INTERVAL seconds over [start, end)."""
stamps: List[float] = []
ts = start
while ts < end:
stamps.append(ts)
ts = round(ts + CHUNK_INTERVAL, 6)
if not stamps:
return
with patch("app.internal.call_recorder.time.time", side_effect=stamps):
for _ in stamps:
recorder._ingest(chunk)
def duration_of(path) -> float:
return pcm.seconds(len(path.read_bytes()))
def timestamps(recorder):
return [ts for ts, _ in recorder._buffer]
# ---------------------------------------------------------------------------
# Ring buffer trimming (pre-roll duty only)
# ---------------------------------------------------------------------------
def test_idle_buffer_keeps_only_the_rolling_window(recorder):
ingest(recorder, T0, T0 + RING_BUFFER_SECONDS + 20)
assert recorder.buffered_seconds <= RING_BUFFER_SECONDS + CHUNK_INTERVAL
newest = max(timestamps(recorder))
assert min(timestamps(recorder)) >= newest - RING_BUFFER_SECONDS
@pytest.mark.asyncio
async def test_ring_buffer_is_trimmed_even_while_recording(recorder):
"""
The ring buffer serves PRE-ROLL only. An open recording must not pin it —
that was the mechanism that made call length depend on buffer size.
"""
ingest(recorder, T0, T0 + 2.0)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
ingest(recorder, T0 + 2.0, T0 + 2.0 + RING_BUFFER_SECONDS + 10)
assert recorder.buffered_seconds <= RING_BUFFER_SECONDS + CHUNK_INTERVAL
# ...and the audio the ring buffer dropped is safe in the accumulator.
assert recorder._active is not None
assert recorder._active.chunks[0][0] == pytest.approx(T0 + 1.0 - PRE_ROLL_SECONDS, abs=CHUNK_INTERVAL)
# ---------------------------------------------------------------------------
# Call length must not be bounded by the ring buffer
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_call_longer_than_the_ring_buffer_is_captured_whole(recorder):
call_length = RING_BUFFER_SECONDS * 2 + 5 # 65s against a 30s ring buffer
grant = T0 + 1.0
end = grant + call_length
ingest(recorder, T0, grant)
await recorder.start_recording("call-long", start_epoch=grant)
ingest(recorder, grant, end + 1.0)
rec = await recorder.stop_recording(end_epoch=end)
assert rec is not None and rec.path is not None
captured = duration_of(rec.path)
assert captured > RING_BUFFER_SECONDS, "call length must not be clamped by the ring buffer"
assert captured == pytest.approx(call_length + PRE_ROLL_SECONDS, abs=2 * CHUNK_INTERVAL)
@pytest.mark.asyncio
async def test_accumulator_stops_growing_at_the_memory_ceiling(recorder, caplog):
"""A runaway call must not be able to exhaust RAM on a Pi."""
await recorder.start_recording("call-runaway", start_epoch=T0)
# Silent blocks on purpose: this test is about bytes, not content, and the
# all-zero fast path keeps it from spending seconds in the RMS loop.
big = b"\x00" * 64_000
needed = (MAX_RECORDING_BYTES // len(big)) + 5
# An unbounded clock: patching time.time patches it for everything running
# inside the block, not only for our calls.
ticks = itertools.count()
with caplog.at_level("WARNING", logger="drb-edge-node"):
with patch("app.internal.call_recorder.time.time",
side_effect=lambda: T0 + next(ticks) * 0.1):
for _ in range(needed):
recorder._ingest(big)
assert recorder._active is not None
assert recorder._active.total_bytes <= MAX_RECORDING_BYTES
assert recorder._active.truncated_by_cap
assert any("memory ceiling" in r.message for r in caplog.records)
def test_the_byte_ceiling_can_never_truncate_a_legal_call():
"""
PCM costs 44.1 KB/s where MP3 cost 2 KB/s, so this had to be re-derived.
The TIME cap must always bite before the BYTE cap, or a long pursuit would
be silently cut short by a memory limit.
"""
assert MAX_RECORDING_BYTES > MAX_RECORDING_SECONDS * pcm.BYTES_PER_SECOND
# ...and it still has to be a deliberate, bounded number on a Pi.
assert MAX_RECORDING_BYTES <= 48 * 1024 * 1024
def test_the_ring_buffer_memory_cost_is_bounded():
assert RING_BUFFER_SECONDS * pcm.BYTES_PER_SECOND < 2 * 1024 * 1024
# ---------------------------------------------------------------------------
# Voice activity — the signal the segmenter starts and stops on
# ---------------------------------------------------------------------------
def test_digital_silence_produces_no_voice_marks(recorder):
ingest(recorder, T0, T0 + 5.0, chunk=QUIET)
activity = recorder.audio_activity()
assert activity.last_voice_epoch is None
assert activity.voice_onset_epoch is None
def test_voice_onset_is_the_arrival_of_the_first_non_silent_chunk(recorder):
ingest(recorder, T0, T0 + 2.0, chunk=QUIET)
ingest(recorder, T0 + 2.0, T0 + 3.0, chunk=VOICE)
activity = recorder.audio_activity()
assert activity.voice_onset_epoch == pytest.approx(T0 + 2.0)
assert activity.last_voice_epoch == pytest.approx(T0 + 3.0 - CHUNK_INTERVAL)
def test_a_gap_shorter_than_the_silence_timeout_does_not_start_a_new_run(recorder, monkeypatch):
"""Back-and-forth inside the window is ONE run, hence one recording."""
monkeypatch.setattr(settings, "call_silence_timeout", 3.0)
ingest(recorder, T0, T0 + 1.0, chunk=VOICE)
ingest(recorder, T0 + 1.0, T0 + 2.5, chunk=QUIET)
ingest(recorder, T0 + 2.5, T0 + 3.5, chunk=VOICE)
assert recorder.audio_activity().voice_onset_epoch == pytest.approx(T0)
def test_a_gap_longer_than_the_silence_timeout_starts_a_new_run(recorder, monkeypatch):
monkeypatch.setattr(settings, "call_silence_timeout", 3.0)
ingest(recorder, T0, T0 + 1.0, chunk=VOICE)
ingest(recorder, T0 + 1.0, T0 + 6.0, chunk=QUIET)
ingest(recorder, T0 + 6.0, T0 + 7.0, chunk=VOICE)
assert recorder.audio_activity().voice_onset_epoch == pytest.approx(T0 + 6.0)
def test_activity_snapshot_reports_capture_and_recording_state(recorder):
activity = recorder.audio_activity()
assert activity.capturing is True
assert activity.recording is False
recorder._capturing = False
assert recorder.audio_activity().capturing is False
# ---------------------------------------------------------------------------
# Pre-roll and slicing
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_slice_starts_pre_roll_before_the_detected_onset(recorder):
ingest(recorder, T0, T0 + 10.0)
onset = T0 + 5.0
await recorder.start_recording("call-1", start_epoch=onset)
assert recorder._active.slice_start == pytest.approx(onset - PRE_ROLL_SECONDS)
rec = await recorder.stop_recording(end_epoch=onset + 2.0)
assert rec is not None and rec.path is not None and rec.path.exists()
first_ts = recorder._active.chunks[0][0] if recorder._active else None
assert first_ts is None # recording closed
# The audio actually covered must begin at or before the requested slice
# start — erring early is safe, erring late loses speech.
assert rec.audio_start_epoch <= onset - PRE_ROLL_SECONDS + 1e-6
@pytest.mark.asyncio
async def test_tail_chunk_straddling_the_end_is_included(recorder):
ingest(recorder, T0, T0 + 10.0)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
rec = await recorder.stop_recording(end_epoch=T0 + 3.05)
# The chunk covering the end instant must be kept, so the captured audio
# reaches past the requested end rather than stopping short of it.
assert rec.audio_end_epoch >= T0 + 3.05
@pytest.mark.asyncio
async def test_max_recording_seconds_caps_the_slice(recorder):
ingest(recorder, T0, T0 + 1.0)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
ingest(recorder, T0 + 1.0, T0 + MAX_RECORDING_SECONDS + 60)
rec = await recorder.stop_recording(end_epoch=T0 + MAX_RECORDING_SECONDS + 50)
assert duration_of(rec.path) <= MAX_RECORDING_SECONDS + PRE_ROLL_SECONDS + CHUNK_INTERVAL
# ---------------------------------------------------------------------------
# Tail wait — still needed for control-channel-derived ends (tgid splits)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_stop_waits_for_captured_audio_to_reach_the_call_end(recorder, caplog):
"""
A tgid_change close pads past a control-channel timestamp that is ~now, so
the audio it asks for has not been captured yet. Slicing immediately would
cut the last word off.
"""
ingest(recorder, T0, T0 + 4.0)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
async def late_tail():
await asyncio.sleep(0.15)
with patch("app.internal.call_recorder.time.time", return_value=T0 + 4.6):
recorder._ingest(block(SPEECH_LEVEL))
task = asyncio.create_task(late_tail())
with caplog.at_level("INFO", logger="drb-edge-node"):
rec = await recorder.stop_recording(end_epoch=T0 + 4.5)
await task
assert rec is not None and rec.path is not None
# The slice covers the chunks stamped T0+0.8 .. T0+3.9 (32 of them) plus the
# one that arrived late — and that last one is where the final word of the
# transmission lives. Without the wait it would have been cut.
chunk_seconds = pcm.seconds(len(VOICE))
assert duration_of(rec.path) == pytest.approx(33 * chunk_seconds, abs=0.01)
assert duration_of(rec.path) > 32 * chunk_seconds
assert any("Waited" in r.message and "tail" in r.message for r in caplog.records), \
"a tail wait must be observable in the field logs"
@pytest.mark.asyncio
async def test_tail_wait_is_bounded_and_warns_when_audio_never_arrives(recorder, caplog, monkeypatch):
monkeypatch.setattr(recorder_mod, "TAIL_WAIT_TIMEOUT_SECONDS", 0.2)
ingest(recorder, T0, T0 + 4.0)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
started = time.monotonic()
with caplog.at_level("WARNING", logger="drb-edge-node"):
rec = await recorder.stop_recording(end_epoch=T0 + 10.0)
elapsed = time.monotonic() - started
assert elapsed < 2.0, "the wait must be bounded, never open-ended"
assert rec is not None and rec.path is not None, "a short tail still beats no recording"
messages = [r.message for r in caplog.records]
assert any("Tail wait" in m and "gave up" in m for m in messages)
assert any("BUFFER CLAMP" in m for m in messages), "silent truncation must be loud"
@pytest.mark.asyncio
async def test_no_wait_when_the_buffer_already_covers_the_end(recorder, caplog):
"""
An audio-driven close derives its end epoch from audio that is already
buffered, so the common path must never pay the tail wait at all.
"""
ingest(recorder, T0, T0 + 10.0)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
started = time.monotonic()
with caplog.at_level("INFO", logger="drb-edge-node"):
await recorder.stop_recording(end_epoch=T0 + 3.0)
assert (time.monotonic() - started) < 0.1
assert not any("Waited" in r.message for r in caplog.records)
# ---------------------------------------------------------------------------
# Clamping must be loud
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pre_roll_earlier_than_buffer_start_is_clamped_and_warned(recorder, caplog):
"""An onset older than anything buffered must still produce a file, loudly."""
ingest(recorder, T0 + 5.0, T0 + 10.0) # buffer only covers T0+5 onwards
with caplog.at_level("WARNING", logger="drb-edge-node"):
await recorder.start_recording("call-1", start_epoch=T0) # 5 s before the head
rec = await recorder.stop_recording(end_epoch=T0 + 8.0)
assert rec is not None and rec.path is not None and rec.path.stat().st_size > 0
assert rec.clamped_seconds == pytest.approx(5.0 + PRE_ROLL_SECONDS, abs=CHUNK_INTERVAL)
assert any("BUFFER CLAMP" in r.message for r in caplog.records)
@pytest.mark.asyncio
async def test_no_buffered_audio_returns_none(recorder, encodes):
await recorder.start_recording("call-1", start_epoch=T0)
assert await recorder.stop_recording(end_epoch=T0 + 2.0) is None
assert encodes == [], "nothing to encode means no encoder subprocess"
@pytest.mark.asyncio
async def test_start_epoch_omitted_falls_back_to_now(recorder):
now = time.time()
ingest(recorder, now - 5.0, now)
await recorder.start_recording("call-1")
assert recorder._active.slice_start == pytest.approx(now - PRE_ROLL_SECONDS, abs=1.0)
# ---------------------------------------------------------------------------
# Encode — exactly once, at save time
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_audio_is_encoded_exactly_once_per_call(recorder, encodes, monkeypatch):
"""
The old pipeline captured MP3 and then re-encoded it to trim, so every
upload was double-encoded. Capture is PCM now and MP3 happens once, after
trimming, at save time.
"""
monkeypatch.setattr(settings, "trim_silence", True)
ingest(recorder, T0, T0 + 1.0, chunk=QUIET)
ingest(recorder, T0 + 1.0, T0 + 3.0, chunk=VOICE)
ingest(recorder, T0 + 3.0, T0 + 6.0, chunk=QUIET)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
rec = await recorder.stop_recording(end_epoch=T0 + 6.0)
assert rec is not None and rec.path is not None
assert len(encodes) == 1, "exactly one encode per recording"
encoded_audio, encoded_path = encodes[0]
assert encoded_path == rec.path
# What was encoded is the TRIMMED audio, not the raw slice.
assert pcm.seconds(len(encoded_audio)) < 5.0
@pytest.mark.asyncio
async def test_a_failed_encode_leaves_no_file_and_no_recording(recorder, monkeypatch):
async def _fail(audio, path):
return False
monkeypatch.setattr(recorder_mod, "encode_recording", _fail)
ingest(recorder, T0, T0 + 5.0)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
assert await recorder.stop_recording(end_epoch=T0 + 3.0) is None
assert list(recorder._recordings_dir.glob("*.flac")) == []
def test_encoder_command_contract_matches_what_c2_expects():
"""
The saved file is what Whisper transcribes, so the encode must stay
LOSSLESS and must not resample. It was 16 kbps MP3 — a bitrate copied from
Icecast's live stream, i.e. the listening path's budget applied to the
accuracy path — which put a second lossy stage on top of the P25 vocoder.
The sample rate must equal pcm.SAMPLE_RATE or the encode stops being a
straight pass and the byte-offset trim arithmetic no longer lines up.
"""
assert recorder_mod.AUDIO_SAMPLE_RATE == str(pcm.SAMPLE_RATE) == "22050"
assert recorder_mod.AUDIO_FORMAT == "flac"
assert recorder_mod.AUDIO_SUFFIX == ".flac"
assert recorder_mod.AUDIO_MIME == "audio/flac"
# No bitrate constant should exist: a bitrate on a lossless codec would mean
# someone reintroduced lossy encoding.
assert not hasattr(recorder_mod, "MP3_BITRATE")
# ---------------------------------------------------------------------------
# Silence trimming and timing metadata
# ---------------------------------------------------------------------------
async def _recorded(recorder, lead_silence=1.0, voice=2.0, tail_silence=1.0):
start = T0
ingest(recorder, start, start + lead_silence, chunk=QUIET)
ingest(recorder, start + lead_silence, start + lead_silence + voice, chunk=VOICE)
ingest(recorder, start + lead_silence + voice,
start + lead_silence + voice + tail_silence, chunk=QUIET)
await recorder.start_recording("call-1", start_epoch=start + 0.5)
return await recorder.stop_recording(end_epoch=start + lead_silence + voice + tail_silence)
@pytest.mark.asyncio
async def test_trimming_is_off_when_the_setting_is_off(recorder, monkeypatch):
monkeypatch.setattr(settings, "trim_silence", False)
called = False
def _never(*args, **kwargs):
nonlocal called
called = True
return b"", None
monkeypatch.setattr(recorder_mod.audio_trim, "trim_pcm", _never)
rec = await _recorded(recorder)
assert rec is not None and not called
@pytest.mark.asyncio
async def test_trim_shifts_the_audio_bounds_but_not_the_call_bounds(recorder, monkeypatch):
"""
Trimming changes audio duration, so the AUDIO's wall-clock bounds move.
The call's own started_at/ended_at (owned by metadata_watcher) must not be
redefined — the recorder only reports where the audio now sits.
"""
monkeypatch.setattr(settings, "trim_silence", True)
rec = await _recorded(recorder, lead_silence=1.0, voice=2.0, tail_silence=1.5)
assert rec is not None and rec.path is not None
assert rec.lead_trimmed > 0.0 and rec.tail_trimmed > 0.0
guard = settings.trim_silence_guard_seconds
# Slice began at T0+0.25; speech begins at T0+1.0, so the audio now starts
# one guard margin before the speech.
assert rec.audio_start_epoch == pytest.approx(T0 + 1.0 - guard, abs=3 * CHUNK_INTERVAL)
assert rec.audio_end_epoch == pytest.approx(T0 + 3.0 + guard, abs=3 * CHUNK_INTERVAL)
assert rec.audio_end_epoch > rec.audio_start_epoch
@pytest.mark.asyncio
async def test_all_silence_recording_is_dropped_and_logged(recorder, monkeypatch, caplog, encodes):
monkeypatch.setattr(settings, "trim_silence", True)
ingest(recorder, T0, T0 + 5.0, chunk=QUIET)
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
with caplog.at_level("WARNING", logger="drb-edge-node"):
rec = await recorder.stop_recording(end_epoch=T0 + 4.0)
assert rec is not None
assert rec.all_silence is True
assert rec.path is None, "an all-silence recording must not be uploaded"
assert encodes == [], "and must not be encoded either"
assert any("no speech" in r.message for r in caplog.records)
# ---------------------------------------------------------------------------
# Recording lifecycle
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_second_start_is_rejected_while_recording(recorder):
ingest(recorder, T0, T0 + 5.0)
assert await recorder.start_recording("call-1", start_epoch=T0 + 1.0) is True
assert await recorder.start_recording("call-2", start_epoch=T0 + 2.0) is False
assert recorder.is_recording
@pytest.mark.asyncio
async def test_stop_without_start_is_a_noop(recorder):
assert await recorder.stop_recording() is None
assert not recorder.is_recording
@pytest.mark.asyncio
async def test_discard_drops_the_audio_without_writing_anything(recorder, encodes):
"""The orphan-audio path: unattributed audio must never reach a file."""
ingest(recorder, T0, T0 + 5.0)
await recorder.start_recording("call-orphan", start_epoch=T0 + 1.0)
await recorder.discard_recording()
assert not recorder.is_recording
assert encodes == []
assert list(recorder._recordings_dir.glob("*.flac")) == []
# ...and the recorder is immediately reusable.
assert await recorder.start_recording("call-next", start_epoch=T0 + 2.0) is True
@pytest.mark.asyncio
async def test_split_then_immediate_restart_keeps_both_slices(recorder):
"""
A tgid change closes one recording and opens the next at the same instant —
the second must still find its pre-roll in the buffer.
"""
ingest(recorder, T0, T0 + 10.0)
split = T0 + 5.0
await recorder.start_recording("call-1", start_epoch=T0 + 1.0)
first = await recorder.stop_recording(end_epoch=split)
await recorder.start_recording("call-2", start_epoch=split)
second = await recorder.stop_recording(end_epoch=T0 + 8.0)
assert first is not None and first.path.stat().st_size > 0
assert second is not None and second.path.stat().st_size > 0
assert first.path != second.path
# ---------------------------------------------------------------------------
# FFmpeg invocation
# ---------------------------------------------------------------------------
def test_capture_command_asks_for_raw_pcm_not_mp3(recorder):
cmd = recorder._ffmpeg_command()
joined = " ".join(cmd)
assert "-f pulse" in joined
assert "drb_sink.monitor" in joined, "must address the monitor explicitly, not 'default'"
assert cmd[-1] == "-" and cmd[-2] == "s16le", "capture must emit raw PCM on stdout"
assert "mp3" not in joined, "MP3 now happens once at save time, not in the capture"
assert "-ar" in cmd and str(pcm.SAMPLE_RATE) in cmd
assert "-ac" in cmd and str(pcm.CHANNELS) in cmd
def test_read_chunk_is_finer_than_the_pre_roll(recorder):
"""
Chunk size is both the ring buffer's timestamp resolution and the window
silence detection runs over, so it has to stay well under the pre-roll.
"""
chunk_seconds = pcm.seconds(recorder_mod.READ_CHUNK_BYTES)
assert chunk_seconds < PRE_ROLL_SECONDS / 4
assert recorder_mod.READ_CHUNK_BYTES % pcm.FRAME_BYTES == 0
# ---------------------------------------------------------------------------
# Capture-exit classification — the two failure modes must be told apart
# instead of both logging the same generic "restarting" line. This is what
# let a wrong PULSE_SOURCE hide behind normal-looking startup retries before.
# ---------------------------------------------------------------------------
def _log_levels(caplog, logger_name="drb-edge-node"):
return [r.levelname for r in caplog.records if r.name == logger_name]
def test_capture_exit_logs_error_when_source_missing(recorder, caplog):
"""FFmpeg's pulse input prints 'No such process' when the daemon is up
but the configured source name does not exist — a real misconfiguration,
not a startup race, so this must stand out as an error naming the source."""
recorder._last_stderr_lines.append(
"[pulse @ 0x...] pa_stream_connect_record failed: No such process"
)
with caplog.at_level("INFO", logger="drb-edge-node"):
recorder._log_capture_exit()
assert "ERROR" in _log_levels(caplog)
error_messages = [r.message for r in caplog.records if r.levelname == "ERROR"]
assert any(settings.pulse_source in m for m in error_messages)
def test_capture_exit_logs_info_when_no_daemon(recorder, caplog):
"""Connection refused means nothing is listening yet — expected during
startup, so it must NOT be logged at the same severity as a real
misconfiguration."""
recorder._last_stderr_lines.append(
"[pulse @ 0x...] pa_context_connect() failed: Connection refused"
)
with caplog.at_level("INFO", logger="drb-edge-node"):
recorder._log_capture_exit()
levels = _log_levels(caplog)
assert "ERROR" not in levels
assert "INFO" in levels
def test_capture_exit_falls_back_to_generic_warning(recorder, caplog):
"""An FFmpeg failure that matches neither known marker keeps the original
generic behavior rather than guessing."""
recorder._last_stderr_lines.append("[pulse @ 0x...] some other unexpected failure")
with caplog.at_level("INFO", logger="drb-edge-node"):
recorder._log_capture_exit()
assert _log_levels(caplog) == ["WARNING"]
def test_capture_exit_with_no_stderr_captured_is_generic_warning(recorder, caplog):
"""No stderr at all (e.g. FFmpeg killed before printing anything) must not
crash the classifier and must fall back to the generic message."""
assert list(recorder._last_stderr_lines) == []
with caplog.at_level("INFO", logger="drb-edge-node"):
recorder._log_capture_exit()
assert _log_levels(caplog) == ["WARNING"]
File diff suppressed because it is too large Load Diff
+142
View File
@@ -0,0 +1,142 @@
"""
Unit tests for mqtt_manager's per-node auth + TLS wiring
(MQTT-PUBLIC-AUTH-PLAN.md dynsec cutover).
Pure client-construction tests — _build_client() only builds a paho Client
object, it never calls .connect(), so no real broker is involved. What's
verified here is the credential/TLS *selection logic*, matching what the
server's dynsec plugin now expects (username=node_id, password=api_key,
default-verified TLS on the public listener) — see
Server/drb-c2-core/app/internal/dynsec.py and mosquitto.conf (read-only
reference, not touched by this change).
"""
import ssl
from unittest.mock import patch
import pytest
from app.config import settings
from app.internal import credentials
from app.internal.mqtt_manager import mqtt_manager
@pytest.fixture(autouse=True)
def isolated_mqtt_settings(monkeypatch):
"""Every test gets known, isolated mqtt_* settings and a clean
credentials._api_key so tests can't see real .env values or leak state
between tests (mirrors the isolated_credentials fixture in test_auth.py)."""
monkeypatch.setattr(settings, "mqtt_user", None)
monkeypatch.setattr(settings, "mqtt_pass", None)
monkeypatch.setattr(settings, "mqtt_tls", False)
monkeypatch.setattr(credentials, "_api_key", None)
yield
# ---------------------------------------------------------------------------
# Credential selection: api_key > legacy mqtt_user > anonymous
# ---------------------------------------------------------------------------
def test_build_client_uses_node_id_and_api_key_when_present(monkeypatch):
monkeypatch.setattr(credentials, "_api_key", "the-api-key")
client = mqtt_manager._build_client()
assert client._username == settings.node_id.encode()
assert client._password == b"the-api-key"
def test_build_client_falls_back_to_legacy_mqtt_user_without_api_key(monkeypatch):
monkeypatch.setattr(settings, "mqtt_user", "drb-node")
monkeypatch.setattr(settings, "mqtt_pass", "legacy-pass")
client = mqtt_manager._build_client()
assert client._username == b"drb-node"
assert client._password == b"legacy-pass"
def test_build_client_api_key_takes_priority_over_legacy_mqtt_user(monkeypatch):
"""Once a node has a real api_key, it must never fall back to the shared
legacy login even if MQTT_USER/MQTT_PASS are still set in .env."""
monkeypatch.setattr(credentials, "_api_key", "the-api-key")
monkeypatch.setattr(settings, "mqtt_user", "drb-node")
monkeypatch.setattr(settings, "mqtt_pass", "legacy-pass")
client = mqtt_manager._build_client()
assert client._username == settings.node_id.encode()
assert client._password == b"the-api-key"
def test_build_client_with_no_credentials_connects_anonymously(monkeypatch):
"""No api_key on disk, no legacy login configured: _build_client() must
still return a usable client (paho, not this code, decides what happens
on the wire — the dynsec broker refuses it, see the warning test below).
This must never raise."""
client = mqtt_manager._build_client()
assert client._username is None
assert client._password is None
def test_build_client_warns_when_no_credentials_available(caplog):
with caplog.at_level("WARNING", logger="drb-edge-node"):
mqtt_manager._build_client()
messages = [r.message for r in caplog.records]
assert any("No API key" in m for m in messages), \
"an unenrolled node must log a clear, greppable warning, not fail silently"
def test_build_client_does_not_warn_when_api_key_present(monkeypatch, caplog):
monkeypatch.setattr(credentials, "_api_key", "the-api-key")
with caplog.at_level("WARNING", logger="drb-edge-node"):
mqtt_manager._build_client()
assert not any("No API key" in r.message for r in caplog.records)
# ---------------------------------------------------------------------------
# TLS
# ---------------------------------------------------------------------------
def test_build_client_no_tls_by_default(monkeypatch):
monkeypatch.setattr(credentials, "_api_key", "the-api-key")
monkeypatch.setattr(settings, "mqtt_tls", False)
client = mqtt_manager._build_client()
assert client._ssl_context is None
def test_build_client_enables_tls_with_default_verification(monkeypatch):
monkeypatch.setattr(credentials, "_api_key", "the-api-key")
monkeypatch.setattr(settings, "mqtt_tls", True)
client = mqtt_manager._build_client()
assert isinstance(client._ssl_context, ssl.SSLContext)
# The whole point: default CA verification against the broker's real
# Let's Encrypt cert must stay ON. tls_insecure_set(True) must never be
# called — that would defeat verification entirely.
assert client._ssl_context.verify_mode == ssl.CERT_REQUIRED
assert client._tls_insecure is False
# ---------------------------------------------------------------------------
# Offline call buffer must be untouched by the auth/TLS change
# ---------------------------------------------------------------------------
def test_build_client_does_not_touch_offline_buffer(monkeypatch):
"""_build_client() is called fresh on every connect(); it must never
reset or otherwise touch the offline call-buffer deque — that survives
reconnects/auth changes by design (the whole point of the buffer)."""
monkeypatch.setattr(credentials, "_api_key", "the-api-key")
mqtt_manager._offline_buffer.append(("nodes/test/metadata", {"call_id": "sentinel"}))
with patch.object(mqtt_manager, "_offline_buffer", mqtt_manager._offline_buffer):
mqtt_manager._build_client()
assert list(mqtt_manager._offline_buffer) == [("nodes/test/metadata", {"call_id": "sentinel"})]
mqtt_manager._offline_buffer.clear()
+134
View File
@@ -0,0 +1,134 @@
"""
Unit tests for the raw-PCM primitives that silence detection rests on.
The whole audio-driven design depends on one field observation: between
transmissions the captured stream is DIGITAL silence (a PulseAudio null-sink
monitor), measured at about -91 dBFS — one least-significant bit — not an analog
noise floor. These tests pin that assumption down in code: a 1-LSB "silent"
buffer must read as silence at every sane threshold, and speech-level audio must
never read as silence.
"""
from array import array
import pytest
from app.internal import pcm
def tone(level: int, samples: int = 1024) -> bytes:
"""A square wave at +/-level, so RMS == level exactly."""
return array("h", [level, -level] * (samples // 2)).tobytes()
def zeros(samples: int = 1024) -> bytes:
return b"\x00\x00" * samples
# ---------------------------------------------------------------------------
# Format arithmetic
# ---------------------------------------------------------------------------
def test_capture_format_is_22050_mono_16bit():
"""MP3_SAMPLE_RATE in call_recorder must match, or the encode resamples."""
assert (pcm.SAMPLE_RATE, pcm.CHANNELS, pcm.SAMPLE_WIDTH) == (22050, 1, 2)
assert pcm.BYTES_PER_SECOND == 44100
def test_seconds_and_byte_offset_round_trip():
assert pcm.seconds(pcm.BYTES_PER_SECOND) == pytest.approx(1.0)
assert pcm.byte_offset(1.0) == pcm.BYTES_PER_SECOND
assert pcm.byte_offset(0.5) == 22050
def test_byte_offset_is_always_sample_aligned():
"""A byte offset that splits a sample would shift every later sample."""
for seconds in (0.001, 0.0137, 0.25, 1.7):
assert pcm.byte_offset(seconds) % pcm.FRAME_BYTES == 0
def test_align_drops_a_trailing_half_sample():
assert pcm.align(9) == 8
assert pcm.align(0) == 0
assert pcm.align(-4) == 0
# ---------------------------------------------------------------------------
# Silence detection
# ---------------------------------------------------------------------------
def test_exact_digital_zero_is_silence():
assert pcm.is_all_zero(zeros())
assert pcm.rms_dbfs(zeros()) == pcm.SILENT_DBFS
assert pcm.is_silent(zeros(), -50.0)
assert pcm.is_silent(zeros(), -90.0)
def test_one_lsb_of_dither_is_the_measured_field_floor():
"""
The gap between transmissions measures ~-91 dBFS on a live node, which is
exactly 20*log10(1/32768) — a single LSB. It must read as silence at any
threshold we would ever configure.
"""
floor = tone(1)
assert pcm.rms_dbfs(floor) == pytest.approx(-90.3, abs=0.2)
assert pcm.is_silent(floor, -50.0)
assert pcm.is_silent(floor, -70.0)
assert not pcm.is_silent(floor, -95.0), "an absurd threshold should still be honoured"
def test_speech_level_audio_is_never_silence():
"""Speech on the live node averages about -18 dBFS."""
speech = tone(4096) # -18.06 dBFS
assert pcm.rms_dbfs(speech) == pytest.approx(-18.06, abs=0.1)
assert not pcm.is_silent(speech, -50.0)
assert not pcm.is_silent(speech, -40.0)
def test_threshold_is_honoured_exactly_at_the_boundary():
# RMS 104 -> -49.96 dBFS, just above a -50 threshold.
assert not pcm.is_silent(tone(104), -50.0)
# RMS 100 -> -50.30 dBFS, just below it.
assert pcm.is_silent(tone(100), -50.0)
def test_empty_buffer_counts_as_silence():
"""
"No audio arrived" must never read as "someone is talking" — otherwise a
stalled capture would hold a segment open forever.
"""
assert pcm.is_silent(b"", -50.0)
assert pcm.rms_dbfs(b"") == pcm.SILENT_DBFS
def test_full_scale_is_zero_dbfs():
assert pcm.rms_dbfs(tone(32767)) == pytest.approx(0.0, abs=0.001)
def test_a_trailing_odd_byte_does_not_break_detection():
"""Short reads at EOF can leave half a sample; it must be dropped, not skew."""
assert not pcm.is_silent(tone(8000) + b"\x00", -50.0)
assert pcm.samples(zeros(4) + b"\x01").itemsize == 2
assert len(pcm.samples(zeros(4) + b"\x01")) == 4
def test_rms_attenuates_an_isolated_click_the_way_peak_would_not():
"""
Why RMS and not peak. A single stray sample in an otherwise silent window is
a decoder click, not speech. Peak would score it at its full amplitude and
hold a recording open; RMS spreads it over the window and divides it down by
sqrt(N) — 30 dB for a 1024-sample window.
A full-scale click still reads as signal even after that attenuation, which
is deliberate: at worst it extends a recording by the silence timeout, and
the trim strips the result before upload. Under-detecting speech is the
failure that loses words permanently.
"""
moderate = bytearray(zeros(1024))
moderate[0:2] = array("h", [1000]).tobytes() # -30 dBFS peak
assert pcm.rms_dbfs(bytes(moderate)) == pytest.approx(-60.4, abs=0.2)
assert pcm.is_silent(bytes(moderate), -50.0)
full_scale = bytearray(zeros(1024))
full_scale[0:2] = array("h", [32767]).tobytes()
assert pcm.rms_dbfs(bytes(full_scale)) == pytest.approx(-30.1, abs=0.2)
assert not pcm.is_silent(bytes(full_scale), -50.0)
+141
View File
@@ -0,0 +1,141 @@
"""
Unit tests for PulseAudio readiness helpers (app.internal.pulse).
The whole point of this module is that readiness means "a PulseAudio
connection actually succeeds," never "a socket file exists at this path" —
that was the exact bug reproduced on live hardware: a killed daemon left its
pid file and native socket behind in the shared `pulse_socket` volume, the
old file-existence check reported "ready", and FFmpeg launched against a
dead daemon.
No real PulseAudio daemon or `pactl` binary is required for these tests:
`_daemon_responds` (the one function that actually shells out) is monkeypatched
everywhere except the dedicated subprocess-layer tests, which fake out
`shutil.which`/`subprocess.run` directly so the "stale file, dead daemon" case
is proven at the layer that matters.
"""
import subprocess
from unittest.mock import Mock
import pytest
from app.internal import pulse
# ---------------------------------------------------------------------------
# socket_path()
# ---------------------------------------------------------------------------
def test_socket_path_defaults_when_pulse_server_unset(monkeypatch):
monkeypatch.delenv("PULSE_SERVER", raising=False)
assert pulse.socket_path() == pulse.DEFAULT_SOCKET_PATH
def test_socket_path_parses_unix_prefixed_pulse_server(monkeypatch):
monkeypatch.setenv("PULSE_SERVER", "unix:/tmp/somewhere/native")
assert pulse.socket_path() == "/tmp/somewhere/native"
def test_socket_path_falls_back_on_malformed_pulse_server(monkeypatch):
monkeypatch.setenv("PULSE_SERVER", "not-a-unix-uri")
assert pulse.socket_path() == pulse.DEFAULT_SOCKET_PATH
# ---------------------------------------------------------------------------
# is_ready() / wait_until_ready() against a monkeypatched probe
# ---------------------------------------------------------------------------
def test_is_ready_true_when_daemon_responds(monkeypatch):
monkeypatch.setattr(pulse, "_daemon_responds", lambda: True)
assert pulse.is_ready() is True
def test_is_ready_false_when_daemon_does_not_respond(monkeypatch):
monkeypatch.setattr(pulse, "_daemon_responds", lambda: False)
assert pulse.is_ready() is False
@pytest.mark.asyncio
async def test_wait_until_ready_short_circuits_when_already_live(monkeypatch):
calls = Mock(return_value=True)
monkeypatch.setattr(pulse, "_daemon_responds", calls)
assert await pulse.wait_until_ready(timeout=5) is True
assert calls.call_count == 1
@pytest.mark.asyncio
async def test_wait_until_ready_polls_until_daemon_comes_up(monkeypatch):
monkeypatch.setattr(pulse, "POLL_INTERVAL", 0.01)
responses = iter([False, False, True])
monkeypatch.setattr(pulse, "_daemon_responds", lambda: next(responses))
assert await pulse.wait_until_ready(timeout=5) is True
@pytest.mark.asyncio
async def test_wait_until_ready_times_out_when_daemon_never_responds(monkeypatch):
monkeypatch.setattr(pulse, "POLL_INTERVAL", 0.01)
monkeypatch.setattr(pulse, "_daemon_responds", lambda: False)
assert await pulse.wait_until_ready(timeout=0.05) is False
@pytest.mark.asyncio
async def test_wait_until_ready_uses_settings_default_when_timeout_omitted(monkeypatch):
monkeypatch.setattr(pulse.settings, "pulse_wait_timeout", 0.05)
monkeypatch.setattr(pulse, "POLL_INTERVAL", 0.01)
monkeypatch.setattr(pulse, "_daemon_responds", lambda: False)
assert await pulse.wait_until_ready() is False
# ---------------------------------------------------------------------------
# _daemon_responds() at the subprocess layer — proves a stale FILE is not
# enough, which is the actual regression this module fixes.
# ---------------------------------------------------------------------------
def test_daemon_responds_false_when_pactl_missing(monkeypatch):
monkeypatch.setattr(pulse.shutil, "which", lambda name: None)
assert pulse._daemon_responds() is False
def test_daemon_responds_false_on_probe_timeout(monkeypatch, tmp_path):
monkeypatch.setattr(pulse.shutil, "which", lambda name: "/usr/bin/pactl")
def fake_run(*args, **kwargs):
raise subprocess.TimeoutExpired(cmd="pactl", timeout=pulse.PROBE_TIMEOUT_SECONDS)
monkeypatch.setattr(pulse.subprocess, "run", fake_run)
assert pulse._daemon_responds() is False
def test_daemon_responds_false_when_stale_socket_file_exists_but_daemon_dead(monkeypatch, tmp_path):
"""
The regression, reproduced at the layer that matters: a plain FILE sits
at the socket path (exactly what a killed daemon leaves behind), but
`pactl info` against it fails (nonzero exit — connection refused). This
must NOT be treated as ready.
"""
stale_socket = tmp_path / "native"
stale_socket.write_bytes(b"") # a stale file, not a live socket
monkeypatch.setenv("PULSE_SERVER", f"unix:{stale_socket}")
monkeypatch.setattr(pulse.shutil, "which", lambda name: "/usr/bin/pactl")
monkeypatch.setattr(
pulse.subprocess, "run",
lambda *a, **k: subprocess.CompletedProcess(args=a, returncode=1),
)
assert stale_socket.exists() # sanity: the old file-existence check would pass
assert pulse._daemon_responds() is False
def test_daemon_responds_true_when_pactl_succeeds(monkeypatch, tmp_path):
live_socket = tmp_path / "native"
live_socket.write_bytes(b"")
monkeypatch.setenv("PULSE_SERVER", f"unix:{live_socket}")
monkeypatch.setattr(pulse.shutil, "which", lambda name: "/usr/bin/pactl")
monkeypatch.setattr(
pulse.subprocess, "run",
lambda *a, **k: subprocess.CompletedProcess(args=a, returncode=0),
)
assert pulse._daemon_responds() is True
+91
View File
@@ -0,0 +1,91 @@
"""
node-26#9 / #11 — which SDR does what: secondary priority, per-service pins,
and OP25 always opening its dongle by serial.
"""
import asyncio
import json
from unittest.mock import AsyncMock, patch
import pytest
from app.models import NodeConfig, normalize_sdr_pins, normalize_secondary_priority
DEVS = [
{"index": 0, "serial": "69420", "name": "RTL", "duplicate_serial": False},
{"index": 1, "serial": "00000001", "name": "RTL", "duplicate_serial": False},
]
def _cfg(**kw) -> NodeConfig:
return NodeConfig(node_id="n1", node_name="N1", lat=0.0, lon=0.0, **kw)
def test_normalize_priority_keeps_order_drops_unknown_and_duplicates():
assert normalize_secondary_priority(["ais", "bogus", "adsb", "ais"]) == ["ais", "adsb"]
def test_legacy_single_mode_migrates_to_priority():
assert _cfg(secondary_sdr_mode="adsb").secondary_sdr_priority == ["adsb"]
assert _cfg(secondary_sdr_mode="adsb", secondary_sdr_priority=["ais"]).secondary_sdr_priority == ["ais"]
def test_pins_drop_blanks_and_unknown_services():
assert normalize_sdr_pins({"op25": "00000001", "adsb": "", "ais": None, "x": "1"}) == {"op25": "00000001"}
def test_two_services_cannot_share_a_dongle():
with pytest.raises(ValueError):
normalize_sdr_pins({"op25": "69420", "adsb": "69420"})
@pytest.fixture
def node(tmp_path):
"""Isolated node_config.json + OP25 active.cfg.json, mocked decoders/op25/mqtt."""
import app.internal.config_manager as cm
from app.internal import sdr_settings as ss
op25_cfg = tmp_path / "active.cfg.json"
op25_cfg.write_text(json.dumps({"devices": [{"args": "rtl", "name": "sdr"}]}))
with patch.object(cm, "_CONFIG_FILE", tmp_path / "node_config.json"), \
patch.object(ss, "_OP25_CONFIG", op25_cfg), \
patch.object(ss.secondary_sdr_client, "devices", AsyncMock(return_value=DEVS)), \
patch.object(ss.secondary_sdr_client, "apply", AsyncMock(return_value=["adsb"])) as apply, \
patch("app.internal.op25_client.op25_client.stop", AsyncMock()) as op25_stop, \
patch("app.internal.op25_client.op25_client.start", AsyncMock()), \
patch("app.internal.mqtt_manager.mqtt_manager.publish_checkin", AsyncMock()), \
patch("asyncio.sleep", AsyncMock()):
cm.save_node_config(_cfg())
yield {"cm": cm, "ss": ss, "apply": apply, "op25_stop": op25_stop,
"op25_args": lambda: json.loads(op25_cfg.read_text())["devices"][0]["args"]}
def test_unpinned_op25_is_still_opened_by_serial_of_the_first_dongle(node):
assert asyncio.run(node["ss"].pin_op25_device()) == "69420"
assert node["op25_args"]() == "rtl=69420"
def test_pinned_op25_opens_its_pinned_dongle(node):
node["cm"].save_node_config(_cfg(sdr_pins={"op25": "00000001"}))
asyncio.run(node["ss"].pin_op25_device())
assert node["op25_args"]() == "rtl=00000001"
def test_shared_serial_leaves_op25_on_plain_rtl(node):
dup = [dict(d, serial="00000001") for d in DEVS]
with patch.object(node["ss"].secondary_sdr_client, "devices", AsyncMock(return_value=dup)):
assert asyncio.run(node["ss"].pin_op25_device()) is None
assert node["op25_args"]() == "rtl"
def test_priority_change_never_restarts_op25_and_reserves_its_dongle(node):
node["cm"].save_node_config(_cfg(sdr_pins={"op25": "00000001"}))
asyncio.run(node["ss"].set_sdr_settings(priority=["adsb", "ais"]))
node["op25_stop"].assert_not_awaited()
node["apply"].assert_awaited_with(["adsb", "ais"], {}, ["00000001"])
def test_moving_op25_restarts_it_on_the_new_dongle(node):
asyncio.run(node["ss"].set_sdr_settings(priority=["adsb"], pins={"op25": "00000001", "adsb": "69420"}))
node["op25_stop"].assert_awaited_once()
assert node["op25_args"]() == "rtl=00000001"
node["apply"].assert_awaited_with(["adsb"], {"adsb": "69420"}, ["00000001"])
assert node["cm"].load_node_config().sdr_pins == {"op25": "00000001", "adsb": "69420"}
+13 -2
View File
@@ -1,8 +1,19 @@
#!/bin/sh
set -e
ICECAST_SOURCE_PASSWORD="${ICECAST_SOURCE_PASSWORD:-hackme}"
ICECAST_ADMIN_PASSWORD="${ICECAST_ADMIN_PASSWORD:-admin}"
# No defaults here on purpose. This container binds all interfaces, so a
# fallback password is a published credential on every node that ever accepted
# it -- and the source password is what lets a caller PUSH audio into the
# stream the frontend plays as live radio. Refuse to start instead.
for var in ICECAST_SOURCE_PASSWORD ICECAST_ADMIN_PASSWORD; do
eval "value=\${$var}"
if [ -z "$value" ]; then
echo "icecast: $var is not set." >&2
echo "icecast: set it in the node's .env -- 'bash setup.sh' generates a random one," >&2
echo "icecast: or run: openssl rand -base64 24" >&2
exit 1
fi
done
export ICECAST_SOURCE_PASSWORD ICECAST_ADMIN_PASSWORD
+545
View File
@@ -0,0 +1,545 @@
#!/usr/bin/env bash
#
# DRB edge node — one-shot bootstrap for a clean Raspberry Pi OS (arm64).
#
# curl -fsSL https://git.vpn.cusano.net/logan/node-26/raw/tag/v1/install.sh \
# | sudo bash -s -- --token DRB-xxxx --node-id node-003 \
# --c2-url https://api.<domain> --mqtt-broker mqtt.<domain>
#
# (fetch install.sh from the same tag it installs; use raw/branch/main with
# --track-main for the owner's own always-latest nodes)
#
# Hosting, the pinned ref, and setup.sh's retirement are settled (owner
# decisions, 2026-09-06). One standing hazard remains — D3 below.
#
# What it does, in order:
# 1. Preflight (root, arch, apt, SDR present)
# 2. Install Docker + compose plugin + git + curl + jq
# 3. Clone logan/node-26 at a PINNED ref into $INSTALL_DIR
# 4. Write .env (non-interactive from flags/env, interactive fallback)
# 5. Enroll with C2 (POST /nodes/enroll) and poll for the api_key
# 6. docker compose pull && up -d (prebuilt; --build opts into the ~1h build)
# 7. Print the approval step
#
# It does NOT set up WireGuard. WireGuard-per-node was evaluated and rejected
# (Server/MQTT-PUBLIC-AUTH-PLAN.md, Server/infra/main.tf:67-73). A field node
# reaches production over the public internet only:
# https://api.<domain> enrollment + /upload (Caddy, real TLS)
# mqtt.<domain>:8883 MQTT over TLS, username=NODE_ID password=api_key
# The two stale "nodes reach it via WireGuard" comments in
# Server/docker-compose.prod.yml:6 and Server/infra/main.tf:69 are leftovers.
#
# ---------------------------------------------------------------------------
# SETTLED (owner decisions, 2026-09-06 — node-26#4)
#
# D1 HOSTING. git.vpn.cusano.net is PUBLIC — it is a CNAME to
# cusano-net.duckdns.org (71.117.95.129) with a real Let's Encrypt cert,
# resolvable from any public resolver. install.sh, `git clone` and
# `docker compose pull` all work from a customer's Pi with no VPN.
# Caveat, not a blocker: that IP is a dynamic-DNS record on the owner's
# home uplink, so it is a single point of failure and a bandwidth limit
# — fine for the beachhead, revisit before scaling node count.
#
# D2 PIN. node-26 gets a `v1` tag (owner cuts it — see the git command in
# the mint panel / node-26#4). DEFAULT_REF below is `v1`. `--track-main`
# stays as an opt-in for the owner's own nodes.
#
# STANDING HAZARD
#
# D3 SOURCE-OVERLAY. docker-compose.yml bind-mounts ./drb-edge-node/app and
# ./op25-container/app OVER the image's /app. So even in prebuilt mode
# the running Python is the CLONED REF's code against the pulled image's
# dependencies. `v1` == the commit CI built the current :latest/:stable
# from, so today they match — but the moment `v1` and the image tags
# diverge this silently mixes them. Fix: drop those two mounts from the
# prod compose path, or always retag images from the same ref as `v1`.
# ---------------------------------------------------------------------------
set -euo pipefail
# ── Defaults ────────────────────────────────────────────────────────────────
DEFAULT_REF="v1" # D2
DEFAULT_REPO_URL="https://git.vpn.cusano.net/logan/node-26.git" # D1 (public)
DEFAULT_INSTALL_DIR="/opt/drb/node-26"
REPO_URL="${DRB_REPO_URL:-$DEFAULT_REPO_URL}"
REF="${DRB_REF:-$DEFAULT_REF}"
INSTALL_DIR="${DRB_INSTALL_DIR:-$DEFAULT_INSTALL_DIR}"
NODE_ID="${DRB_NODE_ID:-}"
NODE_NAME="${DRB_NODE_NAME:-}"
NODE_LAT="${DRB_NODE_LAT:-}"
NODE_LON="${DRB_NODE_LON:-}"
C2_URL="${DRB_C2_URL:-}"
MQTT_BROKER="${DRB_MQTT_BROKER:-}"
MQTT_PORT="${DRB_MQTT_PORT:-8883}"
MQTT_TLS="${DRB_MQTT_TLS:-true}"
ENROLLMENT_TOKEN="${DRB_ENROLLMENT_TOKEN:-}"
DASHBOARD_USER="${DRB_DASHBOARD_USERNAME:-admin}"
DASHBOARD_PASS="${DRB_DASHBOARD_PASSWORD:-}"
REGISTRY="${DRB_IMAGE_REGISTRY:-git.vpn.cusano.net}" # D1 (public host)
DOCKER_ORG="${DRB_DOCKER_ORG:-logan}"
DOCKER_REPO="${DRB_DOCKER_REPO:-node-26}"
REGISTRY_USER="${DRB_REGISTRY_USER:-}"
REGISTRY_PASS="${DRB_REGISTRY_PASS:-}"
DO_BUILD=0 # 0 = pull prebuilt images (default), 1 = build on the Pi (~1h for op25)
DO_START=1
ASSUME_YES=0
ENROLL_WAIT="${DRB_ENROLL_WAIT:-0}" # seconds to block waiting for admin approval; 0 = don't block
C='\033[0;36m'; G='\033[0;32m'; Y='\033[1;33m'; R='\033[0;31m'; N='\033[0m'
say() { printf "${C}==>${N} %s\n" "$*"; }
ok() { printf "${G} ok${N} %s\n" "$*"; }
warn() { printf "${Y} !!${N} %s\n" "$*" >&2; }
die() { printf "${R}error:${N} %s\n" "$*" >&2; exit 1; }
usage() {
cat <<'USAGE'
Usage: install.sh [options]
--token TOKEN Enrollment token (Settings -> Nodes -> New token)
--node-id ID Unique node id, e.g. node-003
--name NAME Display name (default: node id)
--lat N --lon N Decimal degrees for the map
--c2-url URL e.g. https://api.drb.example.net
--mqtt-broker HOST e.g. mqtt.drb.example.net
--mqtt-port N default 8883
--no-tls plaintext MQTT (LAN/dev brokers only)
--dashboard-pass PW local dashboard password (generated if omitted)
--ref REF git ref to install (default: v1)
--track-main install main HEAD instead of the v1 tag
--dir PATH install location (default /opt/drb/node-26)
--build build images locally instead of pulling (~1h for op25)
--no-start configure and enroll, but do not start containers
--wait-approval SEC block up to SEC seconds polling for admin approval
-y, --yes never prompt; fail instead of asking
Every option also has an env var: DRB_NODE_ID, DRB_C2_URL, DRB_ENROLLMENT_TOKEN,
DRB_MQTT_BROKER, DRB_REF, DRB_INSTALL_DIR, DRB_REGISTRY_USER/PASS, ...
Secrets are read from the environment or prompted on the TTY, never from a pipe.
USAGE
}
while [ $# -gt 0 ]; do
case "$1" in
--token) ENROLLMENT_TOKEN="$2"; shift 2 ;;
--node-id) NODE_ID="$2"; shift 2 ;;
--name) NODE_NAME="$2"; shift 2 ;;
--lat) NODE_LAT="$2"; shift 2 ;;
--lon) NODE_LON="$2"; shift 2 ;;
--c2-url) C2_URL="$2"; shift 2 ;;
--mqtt-broker) MQTT_BROKER="$2"; shift 2 ;;
--mqtt-port) MQTT_PORT="$2"; shift 2 ;;
--no-tls) MQTT_TLS=false; [ "$MQTT_PORT" = 8883 ] && MQTT_PORT=1883; shift ;;
--dashboard-pass) DASHBOARD_PASS="$2"; shift 2 ;;
--ref) REF="$2"; shift 2 ;;
--track-main) REF="main"; shift ;;
--dir) INSTALL_DIR="$2"; shift 2 ;;
--build) DO_BUILD=1; shift ;;
--no-start) DO_START=0; shift ;;
--wait-approval) ENROLL_WAIT="$2"; shift 2 ;;
-y|--yes) ASSUME_YES=1; shift ;;
-h|--help) usage; exit 0 ;;
*) die "unknown option: $1 (try --help)" ;;
esac
done
# Prompts must come from the terminal, not from the `curl |` pipe on stdin.
ask() { # ask VAR "prompt" "default"
local __v="$1" __p="$2" __d="${3:-}" __r=""
if [ "$ASSUME_YES" = 1 ] || [ ! -r /dev/tty ]; then
[ -n "$__d" ] || die "$__p is required (non-interactive: pass the flag or env var)"
printf -v "$__v" '%s' "$__d"; return
fi
read -rp "$__p${__d:+ [$__d]}: " __r </dev/tty
printf -v "$__v" '%s' "${__r:-$__d}"
}
ask_secret() {
local __v="$1" __p="$2" __r=""
if [ "$ASSUME_YES" = 1 ] || [ ! -r /dev/tty ]; then printf -v "$__v" '%s' ''; return; fi
read -rsp "$__p: " __r </dev/tty; echo >/dev/tty
printf -v "$__v" '%s' "$__r"
}
genpw() { head -c 24 /dev/urandom | base64 | tr -d '\n=' ; }
# ── 1. Preflight ────────────────────────────────────────────────────────────
say "Preflight"
[ "$(id -u)" -eq 0 ] || die "run as root: curl -fsSL <url> | sudo bash -s -- ..."
command -v apt-get >/dev/null || die "apt-get not found — this script targets Raspberry Pi OS / Debian"
ARCH="$(dpkg --print-architecture)"
case "$ARCH" in
arm64|aarch64) ok "arch $ARCH" ;;
*) warn "arch is $ARCH — CI only builds linux/arm64 images. Prebuilt pull will fail; use --build." ;;
esac
ok "running as root"
# Not fatal: the dongle can be plugged in after install.
if command -v lsusb >/dev/null 2>&1 && lsusb | grep -qiE 'rtl2838|realtek.*283[28]|sdr'; then
ok "SDR dongle detected on USB"
else
warn "no RTL-SDR dongle detected on USB — plug one in before expecting audio"
fi
# The kernel's DVB-TV driver auto-binds RTL2838 dongles. op25 detaches it from
# the one it opens, but any other dongle stays claimed and the secondary-sdr
# decoder fails with "usb_claim_interface error -6" (node-26#9).
echo "blacklist dvb_usb_rtl28xxu" > /etc/modprobe.d/blacklist-rtl-sdr.conf
modprobe -r rtl2832_sdr dvb_usb_rtl28xxu 2>/dev/null || true
ok "DVB-TV kernel driver blacklisted for RTL-SDR"
# ── 2. Dependencies ─────────────────────────────────────────────────────────
say "Installing dependencies"
export DEBIAN_FRONTEND=noninteractive
apt-get update -qq
apt-get install -y -qq git curl ca-certificates jq usbutils openssl >/dev/null
ok "git curl jq openssl"
if ! command -v docker >/dev/null 2>&1; then
say "Installing Docker (get.docker.com)"
curl -fsSL https://get.docker.com | sh
fi
docker --version >/dev/null || die "docker install failed"
ok "$(docker --version)"
if ! docker compose version >/dev/null 2>&1; then
apt-get install -y -qq docker-compose-plugin >/dev/null
fi
docker compose version >/dev/null 2>&1 || die "docker compose plugin missing"
ok "$(docker compose version --short 2>/dev/null || echo 'compose plugin')"
systemctl enable --now docker >/dev/null 2>&1 || true
# The docker group only matters for the human who logs in LATER — this script
# is already root, so nothing below needs a re-login. setup.sh's bug (add the
# group, then immediately run compose in the same unprivileged shell) does not
# apply here.
TARGET_USER="${SUDO_USER:-}"
if [ -n "$TARGET_USER" ] && [ "$TARGET_USER" != root ]; then
usermod -aG docker "$TARGET_USER" || true
ok "added '$TARGET_USER' to the docker group (takes effect at their next login)"
fi
# ── 3. Fetch the repo at a pinned ref ───────────────────────────────────────
say "Fetching node-26 @ ${REF}"
mkdir -p "$(dirname "$INSTALL_DIR")"
if [ -d "$INSTALL_DIR/.git" ]; then
ok "existing install at $INSTALL_DIR — updating in place (.env is preserved)"
git -C "$INSTALL_DIR" remote set-url origin "$REPO_URL"
git -C "$INSTALL_DIR" fetch --tags --prune origin
else
git clone --no-checkout "$REPO_URL" "$INSTALL_DIR"
fi
git -C "$INSTALL_DIR" -c advice.detachedHead=false checkout --force "$REF"
RESOLVED_SHA="$(git -C "$INSTALL_DIR" rev-parse HEAD)"
ok "checked out $REF ($RESOLVED_SHA)"
[ "$REF" = main ] && warn "tracking main — two nodes installed on different days run different software"
cd "$INSTALL_DIR"
mkdir -p configs recordings
chmod 700 configs
# ── 4. Configuration ────────────────────────────────────────────────────────
say "Configuring"
if [ -f .env ]; then
ok ".env already present — keeping it (delete it to reconfigure)"
# Re-read the three values section 5 needs. Not `source .env` — that would
# execute whatever is in the file.
envget() { grep -E "^$1=" .env | head -1 | cut -d= -f2- | tr -d '"'; }
NODE_ID="$(envget NODE_ID)"
C2_URL="${C2_URL:-$(envget C2_URL)}"; C2_URL="${C2_URL%/}"
NODE_NAME="${NODE_NAME:-$(envget NODE_NAME)}"
NODE_LAT="${NODE_LAT:-$(envget NODE_LAT)}"
NODE_LON="${NODE_LON:-$(envget NODE_LON)}"
DASHBOARD_USER="$(envget DASHBOARD_USERNAME)"
[ -n "$NODE_ID" ] || die ".env exists but has no NODE_ID — fix or delete it"
else
[ -n "$NODE_ID" ] || ask NODE_ID "Node ID (e.g. node-003)"
[[ "$NODE_ID" =~ ^[A-Za-z0-9_-]+$ ]] || die "NODE_ID must be letters/numbers/dash/underscore only"
[ -n "$NODE_NAME" ] || ask NODE_NAME "Display name" "$NODE_ID"
[ -n "$NODE_LAT" ] || ask NODE_LAT "Latitude" "0.0"
[ -n "$NODE_LON" ] || ask NODE_LON "Longitude" "0.0"
[ -n "$C2_URL" ] || ask C2_URL "C2 API base URL (https://api.<domain>)"
C2_URL="${C2_URL%/}"
[ -n "$MQTT_BROKER" ] || ask MQTT_BROKER "MQTT broker host (mqtt.<domain>)"
if [ -z "$DASHBOARD_PASS" ]; then
ask_secret DASHBOARD_PASS "Local dashboard password (Enter to generate)"
[ -n "$DASHBOARD_PASS" ] || { DASHBOARD_PASS="$(genpw)"; GENERATED_DASH=1; }
fi
ICE_SRC="$(genpw)"; ICE_ADM="$(genpw)"
umask 077
cat > .env <<EOF
# Written by install.sh on $(date -Is) from ref ${RESOLVED_SHA}
NODE_ID=${NODE_ID}
NODE_NAME="${NODE_NAME}"
NODE_LAT=${NODE_LAT}
NODE_LON=${NODE_LON}
# MQTT — post-cutover auth. There is NO shared node login: the node
# authenticates as username=NODE_ID, password=<its C2-issued api_key>, which
# section 5 below fetches into configs/credentials.json. Deliberately no
# MQTT_USER/MQTT_PASS here; a dynsec broker rejects them.
MQTT_BROKER=${MQTT_BROKER}
MQTT_PORT=${MQTT_PORT}
MQTT_TLS=${MQTT_TLS}
C2_URL=${C2_URL}
ICECAST_SOURCE_PASSWORD=${ICE_SRC}
ICECAST_ADMIN_PASSWORD=${ICE_ADM}
ICECAST_HOST=localhost
ICECAST_PORT=8000
ICECAST_MOUNT=/radio
DASHBOARD_USERNAME=${DASHBOARD_USER}
DASHBOARD_PASSWORD=${DASHBOARD_PASS}
PULSE_SOURCE=drb_sink.monitor
OP25_API_URL=http://localhost:8001
OP25_TERMINAL_URL=http://localhost:8081
OP25_DEBUG_EXPOSE=false
IMAGE_REGISTRY=${REGISTRY}
DOCKER_ORG=${DOCKER_ORG}
DOCKER_REPO=${DOCKER_REPO}
EOF
umask 022
chmod 600 .env
[ -n "$TARGET_USER" ] && chown "$TARGET_USER" .env 2>/dev/null || true
ok ".env written for '$NODE_ID'"
fi
# ── 5. Enrollment ───────────────────────────────────────────────────────────
# Client half of Server/drb-c2-core/app/routers/enrollment.py. It does NOT
# exist in the edge-node app today (mqtt_manager.py:73-85 says so explicitly),
# so without this section a fresh node can never obtain an api_key against a
# dynsec broker: MQTT needs the key, and the legacy key-over-MQTT delivery
# needs MQTT. Doing it here breaks that loop.
#
# Two server endpoints, and the exact response shapes verified against
# enrollment.py @ v1:
#
# POST /nodes/enroll (X-Enrollment-Token)
# 200 -> {node_id, pickup_secret, approval_status}
# 403 -> node_id is ALREADY APPROVED. The CRITICAL GUARD in enrollment.py
# refuses to mint a fresh pickup_secret off the shared fleet token.
# So we must only ever POST this for a node we have not enrolled
# from this machine before — i.e. when configs/pickup_secret is
# absent. Re-running the installer must NOT re-POST here.
# 401 bad/revoked token · 400 missing node_id · 429 rate limited
#
# GET /nodes/{id}/credentials (X-Pickup-Secret)
# Always HTTP 200 with {approval_status, api_key} unless the secret or
# node is bad. api_key is null until an admin approves the node in the UI
# (nodes.py approve_node() mints node_keys/{id}.api_key synchronously in
# the same call — approve is enough; assigning a system is independent and
# NOT required for a key). This endpoint has NO already-approved guard, so
# it is the correct — and only working — re-run path after approval.
# 401 -> missing/invalid/rotated pickup secret
# 404 -> node unknown to C2 (deleted server-side, or never enrolled)
CREDS="$INSTALL_DIR/configs/credentials.json"
PICKUP_FILE="$INSTALL_DIR/configs/pickup_secret"
# GET /nodes/{id}/credentials. Sets CRED_HTTP + CRED_BODY (no -f: we need the
# body and status on a 4xx). One implementation so first-run and re-run agree.
creds_pickup() { # creds_pickup PICKUP_SECRET
local _tmp; _tmp="$(mktemp)"
CRED_HTTP="$(curl -sS -o "$_tmp" -w '%{http_code}' \
"$C2_URL/nodes/$NODE_ID/credentials" -H "X-Pickup-Secret: $1" 2>/dev/null || echo 000)"
CRED_BODY="$(cat "$_tmp" 2>/dev/null || true)"
rm -f "$_tmp"
}
cred_field() { printf '%s' "${CRED_BODY:-}" | jq -r "$1 // empty" 2>/dev/null || true; }
write_creds() { # write_creds API_KEY
umask 077; jq -n --arg k "$1" '{api_key:$k}' > "$CREDS"; umask 022
ok "api_key received and written to configs/credentials.json"
}
say_pending() { # say_pending APPROVAL_STATUS — not an error: node is enrolled, key not minted yet
warn "not approved yet — C2 reports approval_status=${1:-pending}, no api_key minted."
warn "an admin must, at <app-url>/settings/nodes : Approve '$NODE_ID' (assigning a system is separate)."
warn "then re-run this installer — it reuses configs/pickup_secret — or fetch it directly:"
warn " curl -fsS $C2_URL/nodes/$NODE_ID/credentials -H \"X-Pickup-Secret: \$(cat $PICKUP_FILE)\" | jq -r .api_key"
}
# Poll creds_pickup for up to ENROLL_WAIT seconds while still pending. Result
# left in CRED_HTTP/CRED_BODY. --wait-approval is what would have avoided the
# original prod bug; the default (0) does not block, so the re-run path below
# must stand on its own.
wait_for_key() { # wait_for_key PICKUP_SECRET
[ "${ENROLL_WAIT:-0}" -gt 0 ] || return 0
local _end; _end=$(( $(date +%s) + ENROLL_WAIT ))
say "Waiting up to ${ENROLL_WAIT}s for an admin to approve '$NODE_ID'"
while [ "$(date +%s)" -lt "$_end" ]; do
sleep 10
creds_pickup "$1"
[ "$CRED_HTTP" = 200 ] || return 0
[ -z "$(cred_field '.api_key')" ] || return 0
done
}
do_fresh_enroll() {
say "Enrolling '$NODE_ID' with $C2_URL"
[ -n "$ENROLLMENT_TOKEN" ] || ask_secret ENROLLMENT_TOKEN "Enrollment token"
[ -n "$ENROLLMENT_TOKEN" ] || die "no enrollment token — mint one at Settings -> Nodes, then re-run with --token"
local BODY RESP PICKUP STATUS KEY
BODY="$(jq -nc --arg id "$NODE_ID" --arg n "${NODE_NAME:-$NODE_ID}" \
--argjson lat "${NODE_LAT:-0}" --argjson lon "${NODE_LON:-0}" \
'{node_id:$id,name:$n,lat:$lat,lon:$lon}')"
# Token goes in a header from a shell variable — never on a process command
# line, never echoed.
if ! RESP="$(curl -fsS -X POST "$C2_URL/nodes/enroll" \
-H "Content-Type: application/json" \
-H "X-Enrollment-Token: $ENROLLMENT_TOKEN" \
--data "$BODY" 2>&1)"; then
case "$RESP" in
*403*) die "enroll refused (403): node '$NODE_ID' is already approved on C2, and
this machine has no configs/pickup_secret to collect its key with. A token
alone cannot re-issue an approved node's key (enrollment.py CRITICAL GUARD).
Recover by restoring this node's original configs/pickup_secret and re-running,
or have an admin Reissue key (Settings -> Nodes) and write the key into
$CREDS by hand (see node-26#6)." ;;
*401*) die "enroll refused (401): enrollment token missing/invalid/revoked —
mint a fresh one at Settings -> Nodes and re-run with --token." ;;
*429*) die "enroll refused (429): rate limited. Wait ~1 minute, then re-run." ;;
*) die "enroll failed: $RESP" ;;
esac
fi
PICKUP="$(printf '%s' "$RESP" | jq -r '.pickup_secret')"
STATUS="$(printf '%s' "$RESP" | jq -r '.approval_status')"
[ -n "$PICKUP" ] && [ "$PICKUP" != null ] || die "enroll returned no pickup_secret: $RESP"
umask 077; printf '%s' "$PICKUP" > "$PICKUP_FILE"; umask 022
ok "enrolled — approval_status=$STATUS (pickup secret saved to configs/pickup_secret)"
creds_pickup "$PICKUP"
[ "$CRED_HTTP" = 200 ] || die "post-enroll credential fetch failed (HTTP $CRED_HTTP): ${CRED_BODY:-<no body>}"
KEY="$(cred_field '.api_key')"
if [ -z "$KEY" ]; then
wait_for_key "$PICKUP" || true
KEY="$(cred_field '.api_key')"
fi
if [ -n "$KEY" ]; then
write_creds "$KEY"
else
say_pending "$(cred_field '.approval_status')"
fi
}
if [ -s "$CREDS" ] && jq -e '.api_key // empty' "$CREDS" >/dev/null 2>&1; then
say "Enrollment"
ok "api_key already on disk — skipping enrollment"
elif [ -z "${C2_URL:-}" ]; then
say "Enrollment"
warn "no C2_URL — skipping enrollment"
elif [ -s "$PICKUP_FILE" ]; then
# RE-RUN. This node already enrolled from this machine. Do NOT POST
# /nodes/enroll again — an approved node_id gets 403 there, and the v1
# installer's own "re-run to pick up the key" advice then dead-ends on
# "use Reissue key". The pickup endpoint has no such guard: use it.
say "Enrollment — collecting credentials for '$NODE_ID' (pickup secret from a previous run)"
PICKUP="$(cat "$PICKUP_FILE")"
creds_pickup "$PICKUP"
case "$CRED_HTTP" in
200)
API_KEY="$(cred_field '.api_key')"
if [ -z "$API_KEY" ]; then
wait_for_key "$PICKUP" || true
API_KEY="$(cred_field '.api_key')"
fi
if [ -n "$API_KEY" ]; then
write_creds "$API_KEY"
else
# Still pending. Enrolled and idempotent — next run collects the key.
# Clean exit, fall through to start the stack. NOT a failure.
say_pending "$(cred_field '.approval_status')"
fi
;;
401)
if [ -n "$ENROLLMENT_TOKEN" ]; then
warn "saved pickup secret rejected (401) — likely rotated by a re-enroll elsewhere. Re-enrolling with --token."
rm -f "$PICKUP_FILE"
do_fresh_enroll
else
die "saved pickup secret is stale (401) and no --token was given. Re-run with
--token DRB-… to re-enroll (only works while the node is still pending), or
have an admin Reissue key for an approved node and write $CREDS by hand."
fi
;;
404)
if [ -n "$ENROLLMENT_TOKEN" ]; then
warn "C2 does not know node '$NODE_ID' (404) — deleted server-side or never fully enrolled. Re-enrolling with --token."
rm -f "$PICKUP_FILE"
do_fresh_enroll
else
die "C2 does not know node '$NODE_ID' (404) and no --token was given.
Re-run with --token DRB-… to enroll it again."
fi
;;
000)
die "could not reach $C2_URL/nodes/$NODE_ID/credentials — check --c2-url and connectivity." ;;
*)
die "credential pickup failed (HTTP $CRED_HTTP): ${CRED_BODY:-<no body>}" ;;
esac
else
say "Enrollment"
do_fresh_enroll
fi
# ── 6. Images + start ───────────────────────────────────────────────────────
if [ "$DO_START" = 1 ]; then
if [ -n "$REGISTRY_USER" ] && [ -n "$REGISTRY_PASS" ]; then
printf '%s' "$REGISTRY_PASS" | docker login "$REGISTRY" -u "$REGISTRY_USER" --password-stdin >/dev/null
ok "logged in to $REGISTRY"
fi
if [ "$DO_BUILD" = 1 ]; then
say "Building images locally — op25 takes roughly an hour on a Pi"
docker compose build
docker compose up -d
else
say "Pulling prebuilt images from $REGISTRY/$DOCKER_ORG/$DOCKER_REPO"
if ! docker compose pull; then
die "pull failed. $REGISTRY is public, so this is most likely a login
requirement or a transient network error: re-run with DRB_REGISTRY_USER /
DRB_REGISTRY_PASS set, or with --build to compile on the Pi (~1h for op25)."
fi
docker compose up --no-build -d
fi
ok "stack started"
else
say "Skipping start (--no-start). Run: cd $INSTALL_DIR && make up-prebuilt"
fi
# ── 7. What the operator does next ──────────────────────────────────────────
IP="$(hostname -I 2>/dev/null | awk '{print $1}')"
cat <<EOF
$(printf "${G}Node '%s' installed at %s${N}" "$NODE_ID" "$INSTALL_DIR")
ref ${RESOLVED_SHA}
images $([ "$DO_BUILD" = 1 ] && echo "built locally" || echo "pulled from $REGISTRY")
dashboard http://${IP:-<node-ip>}/ (user: ${DASHBOARD_USER})
logs cd $INSTALL_DIR && docker compose logs -f edge-node
NEXT — an admin must approve this node before it can do anything:
1. Open <app-url>/settings/nodes
2. Approve "$NODE_ID"
3. Assign it a radio system
EOF
if [ "${GENERATED_DASH:-0}" = 1 ]; then
printf "${Y}Generated dashboard password (shown once): %s${N}\n\n" "$DASHBOARD_PASS"
fi
if [ ! -s "$CREDS" ]; then
if [ -s "$PICKUP_FILE" ]; then
printf "${Y}This node has no api_key yet. After an admin approves it, re-run the same\ninstall command — it reuses configs/pickup_secret and will collect the key\n(no --token needed for the re-run).${N}\n\n"
else
printf "${Y}This node has no api_key and no saved pickup secret, so a plain re-run cannot\nfix it. Re-run with --token DRB-… to enroll; or, if the node is already\napproved, have an admin Reissue key and write it into\n%s by hand.${N}\n\n" "$CREDS"
fi
fi
@@ -1,57 +0,0 @@
name: release-tag
on:
push:
branches:
- dev
jobs:
release-image:
runs-on: ubuntu-latest
env:
DOCKER_LATEST: stable
CONTAINER_NAME: drb-client-discord-bot
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Set up QEMU
uses: docker/setup-qemu-action@v3
- name: Set up Docker BuildX
uses: docker/setup-buildx-action@v3
with: # replace it with your local IP
config-inline: |
[registry."git.vpn.cusano.net"]
http = false
insecure = false
- name: Login to DockerHub
uses: docker/login-action@v3
with:
registry: git.vpn.cusano.net # replace it with your local IP
username: ${{ secrets.GIT_REPO_USERNAME }}
password: ${{ secrets.GIT_REPO_PASSWORD }}
- name: Get Meta
id: meta
run: |
echo REPO_NAME=$(echo ${GITHUB_REPOSITORY} | awk -F"/" '{print $2}') >> $GITHUB_OUTPUT
echo REPO_VERSION=$(git describe --tags --always | sed 's/^v//') >> $GITHUB_OUTPUT
- name: Validate build configuration
uses: docker/build-push-action@v6
with:
call: check
- name: Build and push
uses: docker/build-push-action@v6
with:
context: .
file: ./Dockerfile
platforms: |
linux/arm64
push: true
tags: | # replace it with your local IP and tags
git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ env.CONTAINER_NAME }}:${{ steps.meta.outputs.REPO_VERSION }}
git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ env.CONTAINER_NAME }}:${{ env.DOCKER_LATEST }}
@@ -1,60 +0,0 @@
name: release-tag
on:
push:
branches:
- master
jobs:
release-image:
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
env:
DOCKER_LATEST: stable
CONTAINER_NAME: op25-client
steps:
- name: Checkout
uses: actions/checkout@v5
- name: Set up QEMU
uses: docker/setup-qemu-action@v3
- name: Set up Docker BuildX
uses: docker/setup-buildx-action@v3
with:
config-inline: |
[registry."git.vpn.cusano.net"]
http = false
insecure = false
- name: Login to Gitea Container Registry
uses: docker/login-action@v3
with:
registry: git.vpn.cusano.net
username: ${{ gitea.actor }} # Uses the user or bot that triggered the workflow
password: ${{ secrets.GITHUB_COM_TOKEN }} # The built-in, temporary token
- name: Get Meta
id: meta
run: |
echo REPO_NAME=$(echo ${GITHUB_REPOSITORY} | awk -F"/" '{print $2}') >> $GITHUB_OUTPUT
echo REPO_VERSION=$(git describe --tags --always | sed 's/^v//') >> $GITHUB_OUTPUT
- name: Validate build configuration
uses: docker/build-push-action@v6
with:
call: check
- name: Build and push
uses: docker/build-push-action@v6
with:
context: .
file: ./Dockerfile
platforms: |
linux/arm64
push: true
tags: |
git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ env.CONTAINER_NAME }}:${{ steps.meta.outputs.REPO_VERSION }}
git.vpn.cusano.net/${{ vars.DOCKER_ORG }}/${{ steps.meta.outputs.REPO_NAME }}/${{ env.CONTAINER_NAME }}:${{ env.DOCKER_LATEST }}
-30
View File
@@ -1,30 +0,0 @@
name: Lint
on:
push:
branches:
- master
pull_request:
branches:
- "*"
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.13'
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install flake8
- name: Run Lint
run: |
flake8 --max-line-length=88 --ignore=E203,E302,E501 .
+12 -4
View File
@@ -1,5 +1,10 @@
# OP25 Core Container
FROM python:slim-trixie
# Pinned to a Python major version deliberately. The bare `slim-trixie` tag
# carries no version at all, so a rebuild could move the interpreter across a
# major release -- which this repo has already been bitten by once, when
# app/models.py only ran because trixie happened to ship 3.14 and PEP 649
# defers annotation evaluation. Matches drb-edge-node, which is already 3.14.
FROM python:3.14-slim
# Set environment variables
ENV DEBIAN_FRONTEND=noninteractive
@@ -7,7 +12,7 @@ ENV DEBIAN_FRONTEND=noninteractive
# Install system dependencies
RUN apt-get update && \
apt-get upgrade -y && \
apt-get install git pulseaudio pulseaudio-utils liquidsoap -y
apt-get install git pulseaudio pulseaudio-utils liquidsoap usbutils -y
# Install custom PulseAudio system config (enables anonymous access for edge-node)
COPY system.pa /etc/pulse/system.pa
@@ -50,5 +55,8 @@ RUN sed -i 's/\r$//' /usr/local/bin/docker-entrypoint.sh && \
# 2. Update ENTRYPOINT to use the wrapper script
ENTRYPOINT ["/usr/local/bin/docker-entrypoint.sh"]
# 3. Use CMD to pass the uvicorn command as arguments to the ENTRYPOINT script
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8001", "--reload"]
# 3. Use CMD to pass the launch command as arguments to the ENTRYPOINT script.
# main.py starts uvicorn itself (see `if __name__ == "__main__"`) so the bind
# address can be driven by OP25_DEBUG_EXPOSE at runtime instead of being baked
# into this image at build time.
CMD ["python", "main.py"]
+33
View File
@@ -0,0 +1,33 @@
from pydantic_settings import BaseSettings
class Settings(BaseSettings):
# ------------------------------------------------------------------
# OP25_DEBUG_EXPOSE — debugging aid, NOT a deployment mode.
#
# False (default): the op25 FastAPI control API (:8001, start/stop/
# generate-config) and OP25's own HTTP terminal (:8081, live talkgroup
# metadata) both bind 127.0.0.1. All three Client containers share the
# host network namespace (network_mode: host), so edge-node still reaches
# both over localhost with no functional change — nothing off-box can.
# Neither surface has authentication, so this is the only thing closing
# that hole.
#
# True: both bind 0.0.0.0 — reachable by anything on the node's LAN with
# NO authentication (start/stop OP25, rewrite its config, raw terminal
# access). Only ever set this for local development off a real deployed
# node. A loud warning naming both ports is logged at startup whenever
# this is true.
# ------------------------------------------------------------------
op25_debug_expose: bool = False
class Config:
env_file = ".env"
settings = Settings()
def bind_host() -> str:
"""Resolve the single bind address for both :8001 and :8081 from the flag."""
return "0.0.0.0" if settings.op25_debug_expose else "127.0.0.1"
+18 -1
View File
@@ -5,18 +5,26 @@ import routers.op25_controller as op25_controller
from internal.logger import create_logger
from internal.liquidsoap_config_utils import generate_liquid_script
from models import IcecastConfig
from config import settings, bind_host
LOGGER = create_logger(__name__)
@asynccontextmanager
async def lifespan(app: FastAPI):
if settings.op25_debug_expose:
LOGGER.warning(
"OP25_DEBUG_EXPOSE=true — the op25 control API (:8001) and OP25's "
"HTTP terminal (:8081) are bound to 0.0.0.0 and reachable by "
"ANYTHING on this node's LAN with NO authentication. This is a "
"debugging aid only; do not leave it set on a deployed node."
)
try:
config = IcecastConfig(
icecast_host=os.getenv("ICECAST_HOST", "localhost"),
icecast_port=int(os.getenv("ICECAST_PORT", "8000")),
icecast_mountpoint=os.getenv("ICECAST_MOUNT", "/radio"),
icecast_password=os.getenv("ICECAST_SOURCE_PASSWORD", "hackme"),
icecast_password=os.getenv("ICECAST_SOURCE_PASSWORD", ""),
)
generate_liquid_script(config)
LOGGER.info("op25.liq generated from environment variables.")
@@ -28,3 +36,12 @@ async def lifespan(app: FastAPI):
app = FastAPI(lifespan=lifespan)
app.include_router(op25_controller.create_op25_router(), prefix="/op25")
if __name__ == "__main__":
# Launched directly (see Dockerfile CMD) instead of via `uvicorn main:app
# --host ...` so the bind address is driven by OP25_DEBUG_EXPOSE (config.py)
# rather than a value baked into the image at build time.
import uvicorn
uvicorn.run("main:app", host=bind_host(), port=8001, reload=True)
+19 -8
View File
@@ -1,6 +1,7 @@
from pydantic import BaseModel
from typing import List, Optional, Union
from enum import Enum
from config import bind_host
# Preset device settings for common RTL-SDR hardware.
# gains: OP25 gain string passed to the device block.
@@ -34,6 +35,20 @@ class TalkgroupTag(BaseModel):
talkgroup: str
tagDec: int
# Defined before ConfigGenerator, which annotates a field with it. Under
# Python 3.14 (PEP 649) annotations are evaluated lazily, so the original
# order happened to work in the container; on 3.13 or earlier it is a hard
# NameError at import. Keep the definition above its first use so this file
# does not depend on the base image's Python version.
class IcecastConfig(BaseModel):
icecast_host: str
icecast_port: int
icecast_mountpoint: str
icecast_password: str
icecast_description: Optional[str] = "OP25"
icecast_genre: Optional[str] = "Public Safety"
class ConfigGenerator(BaseModel):
type: DecodeMode
systemName: str
@@ -115,7 +130,9 @@ class MetadataConfig(BaseModel):
class TerminalConfig(BaseModel):
module: Optional[str] = "terminal.py"
terminal_type: Optional[str] = "http:0.0.0.0:8081"
# Bind address comes from OP25_DEBUG_EXPOSE (config.py) — 127.0.0.1 unless
# that flag is set. See config.py for why.
terminal_type: Optional[str] = f"http:{bind_host()}:8081"
terminal_timeout: Optional[float] = 5.0
curses_plot_interval: Optional[float] = 0.2
http_plot_interval: Optional[float] = 1.0
@@ -127,10 +144,4 @@ class TerminalConfig(BaseModel):
### ======================================================
# Icecast models
class IcecastConfig(BaseModel):
icecast_host: str
icecast_port: int
icecast_mountpoint: str
icecast_password: str
icecast_description: Optional[str] = "OP25"
icecast_genre: Optional[str] = "Public Safety"
# (IcecastConfig itself is defined above ConfigGenerator, which references it.)
@@ -1,6 +1,7 @@
from fastapi import HTTPException, APIRouter
import subprocess
import os
import re
import signal
import json
from models import ConfigGenerator, DecodeMode, ChannelConfig, DeviceConfig, TrunkingConfig, TrunkingChannelConfig, TerminalConfig, MetadataConfig, MetadataStreamConfig, HARDWARE_PRESETS
@@ -69,6 +70,24 @@ def create_op25_router():
async def get_status():
return {"status": "running" if _is_running() else "stopped"}
@router.get("/devices")
async def list_sdr_devices():
"""Enumerate connected SDR-looking USB devices via lsusb.
Same match heuristic as install.sh's one-shot host check — good enough
to answer "is a second SDR plugged in", not a serial-level device
binding (op25's DeviceConfig.args has no serial concept yet either).
"""
devices = []
try:
out = subprocess.run(["lsusb"], capture_output=True, text=True, timeout=5).stdout
for line in out.splitlines():
if re.search(r"rtl2838|realtek.*283[28]|sdr", line, re.IGNORECASE):
devices.append(line.strip())
except Exception as e:
LOGGER.warning(f"SDR device enumeration failed: {e}")
return {"count": len(devices), "devices": devices}
@router.post("/generate-config")
async def generate_config(generator: ConfigGenerator):
try:
+65 -32
View File
@@ -1,32 +1,65 @@
#!/bin/bash
# --- Start PulseAudio Daemon ---
# -n: skip default config (load modules inline — avoids system.pa parsing issues)
# --system: run as system-wide daemon
# --log-target=stderr: makes errors visible in Docker logs
# &: background so this script continues; output still captured by Docker
echo "Starting PulseAudio daemon..."
mkdir -p /run/pulse
chmod 777 /run/pulse
pulseaudio --exit-idle-time=-1 -n --system \
--load="module-native-protocol-unix socket=/run/pulse/native auth-anonymous=1" \
--load="module-null-sink sink_name=drb_sink sink_properties=device.description=DRB-Sink" \
--log-target=stderr &
# Wait for the socket to actually exist before continuing
echo "Waiting for PulseAudio socket..."
for i in $(seq 1 20); do
if [ -S /run/pulse/native ]; then
echo "PulseAudio socket ready."
break
fi
sleep 0.5
done
if [ ! -S /run/pulse/native ]; then
echo "WARNING: PulseAudio socket not found after 10s — edge-node audio will fail."
fi
ls -la /run/pulse/
# --- Execute the main command (uvicorn) ---
echo "Starting FastAPI application..."
exec "$@"
#!/bin/bash
PULSE_SOCKET=/run/pulse/native
PULSE_PIDFILE=/run/pulse/pid
mkdir -p /run/pulse
chmod 777 /run/pulse
# Returns 0 (true) only when a PulseAudio daemon actually answers on
# $PULSE_SOCKET. A socket/pid FILE existing proves nothing by itself — that
# is exactly the bug this script works around (see stale-state check below).
pulse_daemon_alive() {
PULSE_SERVER="unix:${PULSE_SOCKET}" timeout 2 pactl info >/dev/null 2>&1
}
# --- Clear stale PulseAudio state left behind by a killed daemon ---
# The `pulse_socket` named volume survives container recreation, but the
# PulseAudio process that owned it does not. If the previous container was
# recreated (not gracefully stopped), its pid file and native socket are
# still sitting in the volume; pulseaudio's pid.c sees the pid file and
# refuses to start ("Daemon already running") even though nothing is
# listening. Only remove these when nothing actually answers on the socket —
# never delete a socket a live daemon is using.
if [ -S "$PULSE_SOCKET" ] || [ -f "$PULSE_PIDFILE" ]; then
if pulse_daemon_alive; then
echo "PulseAudio daemon already alive and responding at ${PULSE_SOCKET} — leaving state as-is."
else
echo "STALE STATE: found ${PULSE_PIDFILE} / ${PULSE_SOCKET} from a previous container, but no daemon answers — clearing before start."
rm -f "$PULSE_SOCKET" "$PULSE_PIDFILE"
fi
fi
# --- Start PulseAudio Daemon ---
# -n: skip default config (load modules inline — avoids system.pa parsing issues)
# --system: run as system-wide daemon
# --log-target=stderr: makes errors visible in Docker logs
# &: background so this script continues; output still captured by Docker
echo "Starting PulseAudio daemon..."
pulseaudio --exit-idle-time=-1 -n --system \
--load="module-native-protocol-unix socket=${PULSE_SOCKET} auth-anonymous=1" \
--load="module-null-sink sink_name=drb_sink sink_properties=device.description=DRB-Sink" \
--log-target=stderr &
# Wait for the daemon to actually answer — NOT just for the socket file to
# exist. A stale socket file from a killed daemon exists but nothing is
# listening on it; a file-existence check reports "ready" against a dead
# daemon, which is exactly how this class of bug slipped through before.
echo "Waiting for PulseAudio to become live..."
PULSE_LIVE=0
for i in $(seq 1 20); do
if pulse_daemon_alive; then
echo "PulseAudio daemon is live (pactl info succeeded)."
PULSE_LIVE=1
break
fi
sleep 0.5
done
if [ "$PULSE_LIVE" -ne 1 ]; then
echo "WARNING: PulseAudio daemon not responding after 10s — edge-node audio will fail until it recovers."
fi
ls -la /run/pulse/
# --- Execute the main command (uvicorn) ---
echo "Starting FastAPI application..."
exec "$@"
+2 -1
View File
@@ -1,2 +1,3 @@
uvicorn
fastapi
fastapi
pydantic-settings
+54
View File
@@ -0,0 +1,54 @@
# Secondary-SDR Container — node-26#9
#
# Claims the node's SECOND physical SDR (the first is always op25's). Mode is
# chosen at runtime via the control API, not baked in: dump1090 for ADS-B,
# AIS-catcher for AIS (op25_2 mode is not handled here yet — see node-26#9).
#
# Device claiming is by RTL-SDR index, not serial (op25's DeviceConfig.args
# has no serial concept either — see op25-container/app/models.py). Index 0
# is reserved for op25; this container always addresses index 1. That's a
# real limitation once serial-stable device binding matters (hot-unplug /
# replug can swap indices) — tracked in node-26#9, not fixed here.
#
# UNVERIFIED: this image has not been built or run against real hardware in
# this session (sandboxed authoring machine, no docker). dump1090 and
# AIS-catcher's exact CLI flags below are believed correct from their
# published docs but not confirmed against a real capture — the CTO/QA
# review before this ships to a real node should build and smoke-test it.
FROM python:3.14-slim
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && \
apt-get upgrade -y && \
apt-get install -y --no-install-recommends \
git build-essential cmake pkg-config \
librtlsdr-dev libusb-1.0-0-dev libssl-dev zlib1g-dev libzstd-dev libncurses-dev usbutils
# readsb (wiedehopf) — ADS-B decoder. antirez/dump1090 was used here first but
# has no --write-json at all (it only serves /data.json over --net), so the
# decoder exited on an unknown flag. readsb writes the dump1090-fa style
# aircraft.json that _read_adsb_snapshot() parses.
RUN git clone --depth 1 https://github.com/wiedehopf/readsb /opt/readsb && \
cd /opt/readsb && make RTLSDR=yes
# AIS-catcher — AIS decoder.
RUN git clone https://github.com/jvde-github/AIS-catcher /opt/AIS-catcher && \
cd /opt/AIS-catcher && mkdir build && cd build && cmake .. && make
EXPOSE 8002
VOLUME ["/configs"]
WORKDIR /app
COPY ./app /app
COPY docker-entrypoint.sh /usr/local/bin/
RUN sed -i 's/\r$//' /usr/local/bin/docker-entrypoint.sh && \
chmod +x /usr/local/bin/docker-entrypoint.sh
COPY requirements.txt /tmp/requirements.txt
RUN pip3 install --no-cache-dir -r /tmp/requirements.txt
ENTRYPOINT ["/usr/local/bin/docker-entrypoint.sh"]
CMD ["python", "main.py"]
+20
View File
@@ -0,0 +1,20 @@
from pydantic_settings import BaseSettings
class Settings(BaseSettings):
# Same rationale as op25-container's OP25_DEBUG_EXPOSE: both containers
# share the host network namespace (network_mode: host), so edge-node
# reaches this control API over localhost regardless of this flag. False
# (default) binds 127.0.0.1; true exposes unauthenticated start/stop to
# the node's LAN and should only ever be set for local development.
secondary_sdr_debug_expose: bool = False
class Config:
env_file = ".env"
settings = Settings()
def bind_host() -> str:
return "0.0.0.0" if settings.secondary_sdr_debug_expose else "127.0.0.1"
@@ -0,0 +1,353 @@
import ctypes
import ctypes.util
import json
import os
import signal
import subprocess
import threading
from pathlib import Path
from typing import Any, Dict, List, Optional
from internal.logger import create_logger
LOGGER = create_logger(__name__)
# One decoder per mode, each on its own SDR (node-26#9). The node's secondary
# SDR *priority* (e.g. ["adsb", "ais"]) is applied by apply(): decoders start
# down the list until no free dongle is left, so every SDR the node has gets
# used and the ones beyond the list's reach simply aren't started. OP25 always
# keeps its own dongle — it is started first and never part of this list.
MODES = ("adsb", "ais")
# Which RTL-SDR index op25 holds is NOT fixed — on radio-box op25 had index 1
# and index 0 was free, so "op25 is always 0" was wrong. rtlsdr can't open a
# dongle another process has claimed, and the decoders exit within ~50ms when
# that happens, so _start_one() tries each index and keeps the first that
# stays up. op25 is never disturbed: a failed claim doesn't touch its dongle.
MAX_SDR_INDEX = 4
_STARTUP_GRACE_S = 2.0
_STATE_DIR = Path("/tmp/secondary_sdr")
ADSB_JSON_DIR = Path("/tmp/adsb")
# Live decoder handles. poll() is the only reliable liveness check: a decoder
# that dies on startup stays an unreaped zombie, and killpg(pgid, 0) still
# succeeds on a zombie.
_procs: Dict[str, subprocess.Popen] = {}
_indices: Dict[str, int] = {}
_lock = threading.Lock()
# AIS-catcher streams one JSON object per received message on stdout rather
# than writing a periodic snapshot file (readsb's approach) — so the
# current-vessel snapshot lives in memory, keyed by mmsi, kept warm by a
# background reader thread for as long as the decoder is alive.
_ais_vessels: Dict[str, Dict[str, Any]] = {}
_ais_lock = threading.Lock()
def _pgid_file(mode: str) -> Path:
return _STATE_DIR / f"{mode}.pgid"
def _reap_orphans() -> None:
"""Kill decoders left behind by a previous API process (uvicorn --reload
restarts this process, but decoders run in their own session and would
otherwise keep holding their SDRs)."""
if not _STATE_DIR.exists():
return
for f in _STATE_DIR.glob("*.pgid"):
try:
os.killpg(int(f.read_text().strip()), signal.SIGTERM)
LOGGER.info(f"Stopped orphaned secondary decoder from {f.name}")
except Exception:
pass
f.unlink(missing_ok=True)
_reap_orphans()
def _adsb_command(index: int) -> List[str]:
ADSB_JSON_DIR.mkdir(parents=True, exist_ok=True)
return [
"/opt/readsb/readsb",
"--net",
"--device-type", "rtlsdr",
"--device", str(index),
"--write-json", str(ADSB_JSON_DIR),
"--write-json-every", "1",
]
def _ais_command(index: int) -> List[str]:
return [
"/opt/AIS-catcher/build/AIS-catcher",
f"-d:{index}", # "-d <x>" would select by serial, not index
"-o", "5", # JSON Full: decoded fields (4 = sparse, "JSON" is rejected)
]
_COMMANDS = {"adsb": _adsb_command, "ais": _ais_command}
def _ais_reader(proc: subprocess.Popen) -> None:
"""
Consume AIS-catcher's stdout, one JSON message per line, and keep the
latest report per mmsi. Field names (mmsi/lat/lon/speed/course or
heading/shipname or name) are believed correct from AIS-catcher's
published JSON output docs but UNVERIFIED against a real capture in
this session — same caveat as dump1090's aircraft.json mapping.
Malformed/partial lines (e.g. static-data-only messages with no
position) are skipped rather than raising, since dropping one line must
never kill the reader thread.
"""
if not proc.stdout:
return
for line in proc.stdout:
try:
msg = json.loads(line)
except Exception:
continue
mmsi = msg.get("mmsi")
if not mmsi:
continue
# AIS-catcher emits separate message TYPES per mmsi — static data
# (name, no position) and position reports (lat/lon, no name) arrive
# as distinct lines. Merge onto the existing entry, only overwriting
# a field the new message actually carries, so a position-only
# report doesn't blank out a name learned from an earlier message.
name = (msg.get("shipname") or msg.get("name") or "").strip() or None
heading = msg.get("heading") if msg.get("heading") is not None else msg.get("course")
updates = {
"mmsi": str(mmsi),
"name": name,
"lat": msg.get("lat"),
"lon": msg.get("lon"),
"speed_kt": msg.get("speed"),
"heading_deg": heading,
}
with _ais_lock:
existing = _ais_vessels.get(str(mmsi), {})
for key, value in updates.items():
if value is not None:
existing[key] = value
_ais_vessels[str(mmsi)] = existing
def is_running(mode: str) -> bool:
proc = _procs.get(mode)
return proc is not None and proc.poll() is None
def running() -> List[str]:
return [m for m in MODES if is_running(m)]
def _start_one(mode: str, candidates: Optional[List[int]] = None) -> bool:
"""Start one decoder on the first free SDR among `candidates` (default:
every index). False when none is free."""
if mode not in _COMMANDS:
raise ValueError(f"Unknown secondary SDR mode: {mode!r}")
if is_running(mode):
return True
if mode == "ais":
with _ais_lock:
_ais_vessels.clear()
needs_stdout = mode == "ais"
for index in candidates if candidates is not None else range(MAX_SDR_INDEX):
if index in {_indices[m] for m in running() if m in _indices}:
continue
try:
proc = subprocess.Popen(
_COMMANDS[mode](index),
preexec_fn=os.setsid,
stdout=subprocess.PIPE if needs_stdout else None,
text=True if needs_stdout else None,
bufsize=1 if needs_stdout else -1,
)
except Exception as e:
LOGGER.error(f"Failed to start secondary SDR decoder mode={mode!r}: {e}")
return False
try:
proc.wait(timeout=_STARTUP_GRACE_S)
LOGGER.info(f"Secondary SDR decoder mode={mode!r} could not use SDR index {index}, trying next")
continue
except subprocess.TimeoutExpired:
pass
if needs_stdout:
threading.Thread(target=_ais_reader, args=(proc,), daemon=True).start()
_procs[mode] = proc
_indices[mode] = index
_STATE_DIR.mkdir(parents=True, exist_ok=True)
_pgid_file(mode).write_text(str(proc.pid))
LOGGER.info(f"Started secondary SDR decoder mode={mode!r} on SDR index {index} pid={proc.pid}")
return True
LOGGER.info(f"Secondary SDR decoder mode={mode!r}: no free SDR left")
return False
def _stop_one(mode: str) -> None:
proc = _procs.pop(mode, None)
_indices.pop(mode, None)
if proc is not None:
try:
os.killpg(proc.pid, signal.SIGTERM)
except OSError:
pass
try:
proc.wait(timeout=5)
except subprocess.TimeoutExpired:
pass
_pgid_file(mode).unlink(missing_ok=True)
def start(mode: str) -> bool:
with _lock:
return _start_one(mode)
def stop(mode: Optional[str] = None) -> None:
with _lock:
for m in [mode] if mode else list(_procs):
_stop_one(m)
def _candidates(mode: str, pins: Dict[str, str], reserved: List[str], devs: List[Dict[str, Any]]) -> List[int]:
"""SDR indices `mode` may use. A pinned mode gets exactly its dongle; an
unpinned one gets any dongle that isn't op25's (reserved) or pinned to
another service. Without enumeration, fall back to probing every index."""
if not devs:
return [] if pins.get(mode) else list(range(MAX_SDR_INDEX))
if pins.get(mode):
idx = _index_of(pins[mode], devs)
return [] if idx is None else [idx]
taken = set(reserved) | {s for m, s in pins.items() if m != mode and s}
return [d["index"] for d in devs if d["serial"] not in taken]
def apply(priority: List[str], pins: Optional[Dict[str, str]] = None,
reserved: Optional[List[str]] = None) -> List[str]:
"""Run decoders in priority order until SDRs run out; stop everything else.
`pins` maps a mode to the serial of the dongle carrying its antenna;
`reserved` lists serials no decoder may touch (op25's). Pins only bind
enabled modes — a disabled service's dongle is free for the others.
Unchanged when the right decoders already run on allowed dongles, so
re-applying the same settings is a no-op rather than a restart.
"""
for m in priority:
if m not in _COMMANDS:
raise ValueError(f"Unknown secondary SDR mode: {m!r}")
pins = {m: s for m, s in (pins or {}).items() if m in priority and s}
reserved = [s for s in (reserved or []) if s]
devs = devices()
with _lock:
for m in running():
if m not in priority or _indices.get(m) not in _candidates(m, pins, reserved, devs):
_stop_one(m)
for i, m in enumerate(priority):
if m in running():
continue
cands = _candidates(m, pins, reserved, devs)
if _start_one(m, cands):
continue
# Out of free dongles: take one from the lowest-priority decoder
# holding a dongle this mode may use, which then gets its own turn.
lower = [x for x in priority[i + 1:] if x in running() and _indices.get(x) in cands]
if lower:
_stop_one(lower[-1])
_start_one(m, cands)
return running()
def devices() -> List[Dict[str, Any]]:
"""Every RTL-SDR on the node with its USB serial — readable even while a
dongle is claimed (op25's included), since it doesn't open the device.
Cheap dongles often ship with the same serial (00000001); those are
flagged, because pinning a service to a shared serial is ambiguous."""
try:
lib = ctypes.CDLL(ctypes.util.find_library("rtlsdr") or "librtlsdr.so.0")
lib.rtlsdr_get_device_name.restype = ctypes.c_char_p
out = []
for i in range(lib.rtlsdr_get_device_count()):
manufact, product, serial = (ctypes.create_string_buffer(256) for _ in range(3))
lib.rtlsdr_get_device_usb_strings(i, manufact, product, serial)
out.append({
"index": i,
"serial": serial.value.decode(errors="replace") or None,
"name": (lib.rtlsdr_get_device_name(i) or b"").decode(errors="replace"),
})
except Exception as e:
LOGGER.warning(f"SDR enumeration failed: {e}")
return []
serials = [d["serial"] for d in out]
for d in out:
d["duplicate_serial"] = d["serial"] is not None and serials.count(d["serial"]) > 1
return out
def _index_of(serial: str, devs: List[Dict[str, Any]]) -> Optional[int]:
"""Index of a uniquely-identified serial; None if absent or ambiguous."""
matches = [d["index"] for d in devs if d["serial"] == serial]
return matches[0] if len(matches) == 1 else None
def status() -> Dict[str, Any]:
devs = devices()
by_index = {d["index"]: d["serial"] for d in devs}
return {
"running": [
{"mode": m, "sdr_index": _indices.get(m), "serial": by_index.get(_indices.get(m))} for m in running()
],
"sdr_count": len(devs) if devs else None,
"devices": devs,
}
def _altitude(a: Dict[str, Any]) -> Optional[int]:
alt = a.get("alt_baro", a.get("altitude"))
if alt == "ground":
return 0
return alt if isinstance(alt, (int, float)) else None
def _read_adsb_snapshot() -> List[Dict[str, Any]]:
"""
Map readsb's aircraft.json (--write-json output) to the server's
telemetry schema. readsb uses the dump1090-fa field names: alt_baro (int,
or the string "ground"), gs, track. Older dump1090 forks used
altitude/speed, kept as a fallback.
"""
path = ADSB_JSON_DIR / "aircraft.json"
try:
raw = json.loads(path.read_text())
except Exception:
return []
out = []
for a in raw.get("aircraft", []):
icao = a.get("hex")
if not icao:
continue
out.append({
"icao": icao.upper(),
"callsign": (a.get("flight") or "").strip() or None,
"lat": a.get("lat"),
"lon": a.get("lon"),
"altitude_ft": _altitude(a),
"ground_speed_kt": a.get("gs", a.get("speed")),
"track_deg": a.get("track"),
})
return out
def data() -> Dict[str, Any]:
with _ais_lock:
vessels = list(_ais_vessels.values()) if is_running("ais") else []
return {
"running": running(),
"aircraft": _read_adsb_snapshot() if is_running("adsb") else [],
"vessels": vessels,
}
@@ -0,0 +1,31 @@
import logging
from logging.handlers import RotatingFileHandler
def create_logger(name, level=logging.DEBUG, max_bytes=10485760, backup_count=2):
debug_log_file = "./secondary-sdr.debug.log"
info_log_file = "./secondary-sdr.log"
logger = logging.getLogger(name)
logger.setLevel(level)
if not logger.hasHandlers():
console_handler = logging.StreamHandler()
console_handler.setLevel(level)
debug_file_handler = RotatingFileHandler(debug_log_file, maxBytes=max_bytes, backupCount=backup_count)
debug_file_handler.setLevel(logging.DEBUG)
info_file_handler = RotatingFileHandler(info_log_file, maxBytes=max_bytes, backupCount=backup_count)
info_file_handler.setLevel(logging.INFO)
formatter = logging.Formatter('%(asctime)s - %(name)s - %(levelname)s - %(message)s')
console_handler.setFormatter(formatter)
debug_file_handler.setFormatter(formatter)
info_file_handler.setFormatter(formatter)
logger.addHandler(console_handler)
logger.addHandler(debug_file_handler)
logger.addHandler(info_file_handler)
return logger
+13
View File
@@ -0,0 +1,13 @@
from fastapi import FastAPI
import routers.secondary_controller as secondary_controller
from config import bind_host
app = FastAPI()
app.include_router(secondary_controller.create_secondary_router(), prefix="/secondary")
if __name__ == "__main__":
import uvicorn
uvicorn.run("main:app", host=bind_host(), port=8002, reload=True)
@@ -0,0 +1,64 @@
from typing import Dict, List, Optional
from fastapi import APIRouter, HTTPException
from pydantic import BaseModel
from internal import decoder_control
from internal.logger import create_logger
LOGGER = create_logger(__name__)
class StartBody(BaseModel):
mode: str # adsb | ais
class StopBody(BaseModel):
mode: Optional[str] = None # omit to stop every decoder
class ApplyBody(BaseModel):
priority: List[str] # ordered, e.g. ["adsb", "ais"]
pins: Dict[str, str] = {} # mode -> serial of the dongle with its antenna
reserved: List[str] = [] # serials no decoder may touch (op25's)
def create_secondary_router():
router = APIRouter()
@router.post("/apply")
async def apply(body: ApplyBody):
try:
live = decoder_control.apply(body.priority, body.pins, body.reserved)
except ValueError as e:
raise HTTPException(status_code=400, detail=str(e))
return {"running": live}
@router.post("/start")
async def start(body: StartBody):
try:
ok = decoder_control.start(body.mode)
except ValueError as e:
raise HTTPException(status_code=400, detail=str(e))
if not ok:
raise HTTPException(status_code=409, detail="No free SDR for this decoder")
return {"status": f"secondary SDR started ({body.mode})"}
@router.post("/stop")
async def stop(body: Optional[StopBody] = None):
decoder_control.stop(body.mode if body else None)
return {"status": "secondary SDR stopped"}
@router.get("/status")
async def get_status():
return decoder_control.status()
@router.get("/devices")
async def get_devices():
return {"devices": decoder_control.devices()}
@router.get("/data")
async def get_data():
return decoder_control.data()
return router
@@ -0,0 +1,3 @@
#!/bin/bash
mkdir -p /tmp/adsb
exec "$@"
+3
View File
@@ -0,0 +1,3 @@
uvicorn
fastapi
pydantic-settings
@@ -0,0 +1,84 @@
"""
node-26#11 — which dongle each secondary decoder may use.
Run from secondary-sdr-container/: PYTHONPATH=app python -m pytest -q tests
"""
from unittest.mock import patch
import pytest
from internal import decoder_control as dc
# radio-box's real pair: op25's dongle and the one with the 1090 antenna.
DEVS = [
{"index": 0, "serial": "69420", "name": "RTL", "duplicate_serial": False},
{"index": 1, "serial": "00000001", "name": "RTL", "duplicate_serial": False},
]
def test_unpinned_mode_never_gets_op25s_dongle():
assert dc._candidates("adsb", {}, ["00000001"], DEVS) == [0]
def test_pinned_mode_gets_exactly_its_dongle():
assert dc._candidates("ais", {"ais": "69420"}, ["00000001"], DEVS) == [0]
def test_unpinned_mode_skips_a_dongle_pinned_to_another_service():
assert dc._candidates("ais", {"adsb": "69420"}, ["00000001"], DEVS) == []
def test_missing_or_ambiguous_pin_gets_nothing():
assert dc._candidates("adsb", {"adsb": "nope"}, [], DEVS) == []
dup = [dict(d, serial="00000001") for d in DEVS]
assert dc._candidates("adsb", {"adsb": "00000001"}, [], dup) == []
class FakeDecoders:
"""Stands in for real processes: one decoder per free index."""
def __init__(self):
self.live = {}
def start(self, mode, candidates=None):
free = [i for i in candidates if i not in self.live.values()]
if not free:
return False
self.live[mode] = free[0]
return True
def stop(self, mode):
self.live.pop(mode, None)
@pytest.fixture
def fake():
f = FakeDecoders()
with patch.object(dc, "devices", return_value=DEVS), \
patch.object(dc, "_start_one", side_effect=f.start), \
patch.object(dc, "_stop_one", side_effect=f.stop), \
patch.object(dc, "running", side_effect=lambda: [m for m in dc.MODES if m in f.live]), \
patch.dict(dc._indices, clear=True):
dc._indices.update(f.live)
yield f
def _apply(fake, priority, pins=None):
dc.apply(priority, pins, ["00000001"])
dc._indices.clear()
dc._indices.update(fake.live)
return fake.live
def test_one_spare_goes_to_the_top_pick(fake):
assert _apply(fake, ["ais", "adsb"]) == {"ais": 0}
def test_reordering_hands_the_spare_to_the_new_top_pick(fake):
_apply(fake, ["adsb", "ais"])
assert _apply(fake, ["ais", "adsb"]) == {"ais": 0}
def test_pinned_lower_priority_still_runs_when_top_pick_has_no_dongle(fake):
# AIS ranked first but its only candidate is ADS-B's pinned antenna dongle.
assert _apply(fake, ["ais", "adsb"], {"adsb": "69420"}) == {"adsb": 0}
-148
View File
@@ -1,148 +0,0 @@
#!/usr/bin/env bash
# Interactive first-time setup for a DRB edge node.
# Installs system dependencies (Docker, make, curl) then writes .env
# and optionally builds + starts the stack.
set -e
GREEN='\033[0;32m'; YELLOW='\033[1;33m'; CYAN='\033[0;36m'; RED='\033[0;31m'; NC='\033[0m'
cd "$(dirname "$0")"
echo -e "${CYAN}DRB Edge Node Setup${NC}"
echo "-------------------"
# ── Dependency installation ──────────────────────────────────────────────────
install_deps() {
if ! command -v apt-get &>/dev/null; then
echo -e "${YELLOW}⚠ apt-get not found — skipping auto-install. Ensure docker, make, and curl are installed.${NC}"
return
fi
echo ""
echo -e "${CYAN}Installing system dependencies…${NC}"
sudo apt-get update -qq
local pkgs=()
command -v make &>/dev/null || pkgs+=(make)
command -v curl &>/dev/null || pkgs+=(curl)
command -v git &>/dev/null || pkgs+=(git)
if [ ${#pkgs[@]} -gt 0 ]; then
echo " Installing: ${pkgs[*]}"
sudo apt-get install -y -qq "${pkgs[@]}"
fi
# Docker — use get.docker.com if not present
if ! command -v docker &>/dev/null; then
echo " Installing Docker via get.docker.com…"
curl -fsSL https://get.docker.com | sudo sh
# Allow current user to run docker without sudo
sudo usermod -aG docker "$USER"
echo -e "${YELLOW} ⚠ Docker group added. You may need to log out and back in for it to take effect.${NC}"
echo -e "${YELLOW} If 'docker compose' fails below, run: newgrp docker${NC}"
else
echo -e "${GREEN} ✓ docker$(docker --version | grep -oP ' \d+\.\d+\.\d+' | head -1)${NC}"
fi
# Docker Compose plugin check (comes with Docker Engine ≥ 20.10)
if ! docker compose version &>/dev/null 2>&1; then
echo -e "${RED} docker compose plugin not found. Installing…${NC}"
sudo apt-get install -y -qq docker-compose-plugin
else
echo -e "${GREEN} ✓ docker compose $(docker compose version --short 2>/dev/null || true)${NC}"
fi
echo -e "${GREEN}✓ Dependencies ready${NC}"
}
install_deps
if [ -f .env ]; then
echo -e "${YELLOW}Warning: .env already exists.${NC}"
read -rp "Overwrite? [y/N] " yn
[[ "$yn" =~ ^[Yy]$ ]] || { echo "Aborted."; exit 0; }
fi
# --- Node identity ---
echo ""
echo "Unique node ID — no spaces (e.g. node-ossining, node-002)"
read -rp "NODE_ID: " NODE_ID
while [[ ! "$NODE_ID" =~ ^[a-zA-Z0-9_-]+$ ]]; do
echo " Use letters, numbers, dashes, underscores only."
read -rp "NODE_ID: " NODE_ID
done
echo ""
read -rp "Node display name [$NODE_ID]: " NODE_NAME
NODE_NAME="${NODE_NAME:-$NODE_ID}"
# --- GPS ---
echo ""
echo "GPS coordinates (decimal degrees — used for the map)"
read -rp "Latitude [0.0]: " NODE_LAT; NODE_LAT="${NODE_LAT:-0.0}"
read -rp "Longitude [0.0]: " NODE_LON; NODE_LON="${NODE_LON:-0.0}"
# --- C2 server ---
echo ""
echo "C2 server — hostname or IP of the machine running the server stack"
read -rp "C2 server host: " C2_HOST; C2_HOST="${C2_HOST:-localhost}"
read -rp "C2 API port [8888]: " C2_PORT; C2_PORT="${C2_PORT:-8888}"
# --- MQTT ---
echo ""
echo "MQTT credentials (must match MQTT_NODE_USER/PASS in the server .env)"
read -rp "MQTT port [1883]: " MQTT_PORT; MQTT_PORT="${MQTT_PORT:-1883}"
read -rp "MQTT username [drb-node]: " MQTT_USER; MQTT_USER="${MQTT_USER:-drb-node}"
read -rsp "MQTT password: " MQTT_PASS; echo ""; MQTT_PASS="${MQTT_PASS:-change-me-node}"
# --- Icecast ---
echo ""
echo "Icecast passwords (local container)"
read -rsp "Source password [hackme]: " ICECAST_SOURCE; echo ""; ICECAST_SOURCE="${ICECAST_SOURCE:-hackme}"
read -rsp "Admin password [admin]: " ICECAST_ADMIN; echo ""; ICECAST_ADMIN="${ICECAST_ADMIN:-admin}"
# --- Write .env ---
cat > .env <<EOF
# Node Identity
NODE_ID=${NODE_ID}
NODE_NAME="${NODE_NAME}"
NODE_LAT=${NODE_LAT}
NODE_LON=${NODE_LON}
# MQTT — point to your C2 server
MQTT_BROKER=${C2_HOST}
MQTT_PORT=${MQTT_PORT}
MQTT_USER=${MQTT_USER}
MQTT_PASS=${MQTT_PASS}
# C2 server for audio upload
C2_URL=http://${C2_HOST}:${C2_PORT}
# API key is provisioned automatically via MQTT after admin approves the node
# Icecast (local container — usually no need to change)
ICECAST_SOURCE_PASSWORD=${ICECAST_SOURCE}
ICECAST_ADMIN_PASSWORD=${ICECAST_ADMIN}
ICECAST_HOST=localhost
ICECAST_PORT=8000
ICECAST_MOUNT=/radio
# OP25 container (usually no need to change)
OP25_API_URL=http://localhost:8001
OP25_TERMINAL_URL=http://localhost:8081
EOF
echo ""
echo -e "${GREEN}✓ .env written for node '${NODE_ID}'${NC}"
echo ""
read -rp "Build and start now? [Y/n] " start
if [[ ! "$start" =~ ^[Nn]$ ]]; then
echo ""
echo "Building images (op25 takes ~10 min on first run)…"
docker compose build
docker compose up -d
echo ""
echo -e "${GREEN}✓ Node '${NODE_ID}' started.${NC}"
echo " → Check the dashboard — it will appear as pending approval."
else
echo "Run 'make up' when ready."
fi