Files
server-26/drb-c2-core/app/config.py
T
Logan Cusano 8fbfe7d6de
Build & Deploy / Build & push images (push) Successful in 4m4s
Build & Deploy / Deploy to VM (push) Failing after 2m16s
Build & Deploy / Report a failed deploy (push) Successful in 1s
Make a failed deploy impossible to miss, and a wildcard CORS harmless
Two unrelated-looking problems with the same shape: a dangerous state that
looked fine from the outside.

DEPLOY (server-26#21). The Deploy job failed on fifteen consecutive pushes
between 2026-08-18 and 08-20 and nobody noticed for two days, because the
build job was green and a red run is only visible to someone who opens Gitea.
Production served 08-18 code the whole time -- including the entire frontend
redesign, chunks 2 through 8. Three changes:

  * The health check now asserts WHICH build answered, not just that something
    did. CI bakes the commit into the image (Dockerfile ARG/ENV GIT_SHA) and
    /health reports it, so a deploy that "succeeds" while the previous
    container keeps running now fails. Liveness alone could never have caught
    this.
  * The image pull retries once after a prune. The actual failure was
    containerd unable to extract a layer -- "failed to Lchown ... no such file
    or directory" -- a corrupted entry in the snapshot store, which a prune
    clears. A second failure after pruning is a real problem (check the VM's
    disk) and still stops the deploy.
  * A notify-failure job POSTs to DEPLOY_ALERT_WEBHOOK when anything in the
    workflow fails. Unset means skip quietly, not fail.

CORS (server-26#20). allow_origins=["*"] with allow_credentials=True is not
the permissive-but-harmless setting it reads as. Starlette does not reject the
pair -- it reflects the caller's Origin back and still sends
Access-Control-Allow-Credentials: true, so the effective policy is "any
origin, WITH credentials", the opposite of what a wildcard normally means.

Rather than trust every deployment to remember CORS_ORIGINS, the pair is now
unrepresentable: a wildcard forces allow_credentials off and logs an ERROR
naming the variable to set. Correctly configured deployments that name their
origins are unaffected and keep credentialed requests.

Severity honestly: low today. c2-core is bearer-auth, and browsers do not
attach bearer tokens cross-origin the way they attach cookies. This is a
misconfiguration waiting for the day something starts trusting a cookie.

Also adds firebase_admin.auth.UserRecord and the list/update/create/delete_user
names to the conftest stub. routers/users.py annotates with UserRecord at
import time, so without it importing app.main failed at collection -- which is
why nothing had ever tested anything wired at app level, CORS included.

Tests: 5 new in test_cors_policy.py, covering the pure policy function, the
middleware actually mounted on the app (so re-hardcoding allow_credentials=True
fails here), and the presence of the build stamp.

Closes logan/server-26#20
Closes logan/server-26#21
2026-08-23 01:26:15 -04:00

157 lines
8.8 KiB
Python

from pydantic_settings import BaseSettings
from typing import Optional
class Settings(BaseSettings):
# MQTT
mqtt_broker: str = "localhost"
mqtt_port: int = 1883
mqtt_user: Optional[str] = None
mqtt_pass: Optional[str] = None
# mosquitto's built-in dynamic-security plugin (see app/internal/dynsec.py).
# "admin" is hardcoded by the plugin itself on first boot — not actually
# configurable — kept as a named setting rather than a literal for
# readability. mqtt_dynsec_admin_pass must equal the mosquitto
# container's own MOSQUITTO_DYNSEC_PASSWORD env var (root .env /
# root.env.j2) or c2-core can't administer node credentials at all.
mqtt_dynsec_admin_user: str = "admin"
mqtt_dynsec_admin_pass: Optional[str] = None
# GCP
gcp_credentials_path: Optional[str] = None # None → uses ADC
gcs_bucket: Optional[str] = None # None → audio upload disabled
firestore_database: str = "(default)"
# Node health
node_offline_threshold: int = 90 # seconds without checkin before marking offline
# OpenAI (STT + intelligence)
openai_api_key: Optional[str] = None
stt_model: str = "whisper-1" # whisper-1 | gpt-4o-mini-transcribe | gpt-4o-transcribe
# Google Maps (geocoding)
google_maps_api_key: Optional[str] = None
# Gemini (intelligence extraction, embeddings, incident summaries)
gemini_api_key: Optional[str] = None
# Correlation consensus models
# corr_cheap_model — first-pass LLM correlator (runs on every call)
# corr_smart_model — tiebreaker (only fires when rules and cheap LLM disagree)
# Both IDs below were retired by Google and returned 404 on every call from
# some point before 2026-08-18 until they were corrected. Because a failed
# LLM call falls back to the rules decision, nothing broke loudly -- the
# entire LLM tier and the consensus tiebreak were simply dead in production
# while correlation behaviour was being tuned against rules-only output.
# Verify against https://ai.google.dev/gemini-api/docs/models before changing.
corr_cheap_model: str = "gemini-3.6-flash" # was gemini-2.0-flash (shut down)
corr_smart_model: str = "gemini-2.5-pro" # was gemini-1.5-pro (shut down)
summary_interval_minutes: int = 2 # how often the summary loop runs
correlation_window_hours: int = 2 # slow/location path: max hours since last call
embedding_similarity_threshold: float = 0.93 # slow-path: requires location corroboration
embedding_no_location_threshold: float = 0.97 # slow-path: match without location (very high bar)
embedding_cross_tg_threshold: float = 0.85 # cross-TG path: same dept + 2+ shared units
location_proximity_km: float = 0.5 # radius for location-proximity matching
geocode_max_km: float = 40.0 # reject geocode results farther than this from the node
incident_auto_resolve_minutes: int = 90 # auto-resolve after N minutes with no new calls
unit_continuity_max_idle_minutes: int = 20 # unit-continuity path: skip if incident idle > this
recorrelation_scan_minutes: int = 60 # re-examine orphaned calls ended within this window
tg_fast_path_idle_minutes: int = 90 # fast path: max minutes since incident last updated
# Dispatch channels only: tier-2 thin calls attach to a lone candidate idle < this.
# Was 10, which is long enough for the channel to have moved on to something else:
# on 2026-08-16 a "72 at Holland Station" incident absorbed a Grand Central train
# meet 9.6 min later, and a status check absorbed a records lookup at 9.7 min.
# Across that dump every correct thin attach was <= 3.4 min idle and every wrong
# one was >= 8.2, so 5 separates them with room on both sides. Genuine
# back-and-forth is handled by the 30-second tier-1 path above this.
tg_dispatch_thin_idle_minutes: int = 5
# Every other channel: tier-2 thin calls attach to a lone candidate idle < this.
# Non-dispatch talkgroups previously had NO tier-2 bound at all — they used the
# whole 90-minute tg_fast_path_idle_minutes window with no single-candidate
# requirement and no fit test, which is the widest version of the 2026-08-20
# over-merge. A tactical channel really is dedicated to one scene, so it earns
# a longer window than a dispatch backbone, but not an unbounded one: 15 sits
# inside the 20-minute tactical-default window in _call_fits_incident, so the
# no-evidence thin path is never more permissive than the fit-tested path on
# the same channel.
tg_thin_idle_minutes: int = 15
# ── Hard caps: an incident past either of these stops accepting calls ──────
# Enforced on every correlation path (see _incident_at_capacity). Pairwise fit
# tests judge one call against one incident and cannot see the shape of the
# chain they are building, so these are the only guard against a "work shift"
# incident regardless of how individually plausible each link looked.
#
# 120 minutes: the one incident in the 2026-08-20 dump that was genuinely a
# single event ran 63 minutes (06:15 wrong-way driver → 07:18 closeout), so
# the cap has to clear an hour with real headroom. The four junk chains ran
# 3h41m, 3h43m, 4h05m and 4h09m, so it has to sit well under three hours.
# 120 also equals correlation_window_hours: the location and slow paths
# already refuse to consider a candidate older than that, and the fast path
# was the only one exempt. Making it agree removes that inconsistency rather
# than inventing a new number.
incident_max_duration_minutes: int = 120
# 40 calls: a backstop for a burst that fills up inside the duration cap
# rather than the primary bound. The worst observed chain averaged ~16
# calls/hour while absorbing an ENTIRE dispatch backbone, so 40 calls in
# under two hours means the incident is eating most of the channel — that is
# a chain, not an event. Set deliberately above any plausible single-incident
# call volume (a multi-alarm fire on its own tactical channel) so this cap
# errs toward keeping real incidents whole and lets the duration cap do the
# cutting.
incident_max_calls: int = 40
# Vocabulary learning
vocabulary_induction_interval_hours: int = 24 # how often the induction loop runs
vocabulary_induction_sample_tokens: int = 4000 # ~tokens of transcript text sampled per system
# Internal service key — allows server-side services (discord bot) to call C2 without Firebase
service_key: Optional[str] = None
# Fleet-wide token edge nodes present to POST /nodes/enroll on first boot.
# Not a per-node secret — see routers/enrollment.py for why a leaked copy
# of this alone can't steal an already-approved node's key.
enrollment_token: Optional[str] = None
# Upload size limit — reject audio files larger than this (bytes). Default 100 MB.
upload_max_bytes: int = 100 * 1024 * 1024
# Public origin this API is reachable on, e.g. "https://api.drb.example.com".
# Only used to build absolute call-audio playback links: an <audio src> is
# fetched by the browser directly, so a relative path would resolve against
# the frontend origin, not this one.
public_api_url: Optional[str] = None
# How long a minted call-audio playback link stays valid. Long enough for a
# browsing session, short enough that a copied link isn't durable access.
audio_link_ttl_seconds: int = 6 * 60 * 60
# Two nodes hearing the same transmission start recording within about a
# second of each other (measured across node-002/node-PI-2 on TG 9048).
# 10s is generous against clock skew while staying well under the gap
# between genuinely separate transmissions on a busy dispatch channel.
duplicate_window_seconds: int = 10
# CORS — set to your frontend origin(s) in production, e.g. ["https://app.example.com"]
# Defaults to "*" for local development only.
#
# Leaving this as "*" is not merely permissive: main.py turns OFF
# allow_credentials when it sees a wildcard, because Starlette would
# otherwise reflect each caller's origin back WITH
# Access-Control-Allow-Credentials. So a production deployment that
# forgets to set this gets a loud ERROR at startup and loses credentialed
# cross-origin requests, rather than silently accepting every origin.
cors_origins: list[str] = ["*"]
# Discord webhook URL that app/internal/ai_health.py posts to when an AI
# tier (transcription/correlation) transitions into or out of degraded
# state. Empty disables the POST entirely — not every self-hosted
# deployment will set this up, and skipping it must be silent.
ai_alert_webhook_url: str = ""
class Config:
env_file = ".env"
settings = Settings()