Correction existed, but as a line in intelligence.py's EXTRACTION_PROMPT --
which put it in the wrong place twice over. The same model call that extracted
units, location and severity emitted the correction afterwards, so extraction
reasoned over text already known to be wrong; and it sat behind
correlation_enabled, so during a cost-controlled STT-only window nothing was
ever corrected at all. That is the normal state during development.
internal/transcript_correction.py is now its own pass, between the degenerate
filter and the Firestore write. It receives an already-produced transcript plus
a reference list, so unlike a Whisper prompt it has no series to extend -- the
distinction that keeps vocabulary out of the recogniser's prompt, where an
enumerated ten-code list once made it hallucinate ten-code runs.
Reference data is merged from the talkgroup and the system, TALKGROUP FIRST. A
system spanning several counties can have a talkgroup covering one
municipality, and that municipality's streets must not be buried under a
county-wide list. A single-municipality system is the degenerate case: populate
the system level and every talkgroup inherits it. Area context is now SET --
municipality, county, roads, landmarks, on both scopes -- rather than guessed
from talkgroup names, which is what vocabulary_learner did and which is close
to useless across multiple counties.
Segments are corrected too, not just the joined text. extract_scenes builds its
prompt from numbered segments whenever there is more than one, so a correction
that only fixed the transcript would have been discarded on exactly the
multi-transmission calls carrying the most content. Alignment is enforced: an
array of the wrong length or type is dropped whole, because scenes map back to
transmissions by index and a shifted array would misattribute audio silently.
Whisper is also retried once on degenerate output. Call e49ea32c produced a
56-word ten-code counting run on one attempt and ordinary speech on the next --
same clip, same temperature=0 -- so a hallucination is a coin-flip, and
discarding on the first bad roll threw away a recoverable transcript.
Two things found on the way:
PUT /systems/{id} wiped ten_codes on every save. The systems form sends only
{name, type, config}, and model_dump() wrote every omitted field as its default
over the top. Now exclude_unset. area_context would have been the next victim,
which is why it gets its own route alongside ten-codes rather than a field on
that payload.
Closes server-26#36.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
168 lines
9.5 KiB
Python
168 lines
9.5 KiB
Python
from pydantic_settings import BaseSettings
|
|
from typing import Optional
|
|
|
|
|
|
class Settings(BaseSettings):
|
|
# MQTT
|
|
mqtt_broker: str = "localhost"
|
|
mqtt_port: int = 1883
|
|
mqtt_user: Optional[str] = None
|
|
mqtt_pass: Optional[str] = None
|
|
|
|
# mosquitto's built-in dynamic-security plugin (see app/internal/dynsec.py).
|
|
# "admin" is hardcoded by the plugin itself on first boot — not actually
|
|
# configurable — kept as a named setting rather than a literal for
|
|
# readability. mqtt_dynsec_admin_pass must equal the mosquitto
|
|
# container's own MOSQUITTO_DYNSEC_PASSWORD env var (root .env /
|
|
# root.env.j2) or c2-core can't administer node credentials at all.
|
|
mqtt_dynsec_admin_user: str = "admin"
|
|
mqtt_dynsec_admin_pass: Optional[str] = None
|
|
|
|
# GCP
|
|
gcp_credentials_path: Optional[str] = None # None → uses ADC
|
|
gcs_bucket: Optional[str] = None # None → audio upload disabled
|
|
firestore_database: str = "(default)"
|
|
|
|
# Node health
|
|
node_offline_threshold: int = 90 # seconds without checkin before marking offline
|
|
|
|
# OpenAI (STT + intelligence)
|
|
openai_api_key: Optional[str] = None
|
|
stt_model: str = "whisper-1" # whisper-1 | gpt-4o-mini-transcribe | gpt-4o-transcribe
|
|
|
|
# Google Maps (geocoding)
|
|
google_maps_api_key: Optional[str] = None
|
|
|
|
# Gemini (intelligence extraction, embeddings, incident summaries)
|
|
gemini_api_key: Optional[str] = None
|
|
# Correlation consensus models
|
|
# corr_cheap_model — first-pass LLM correlator (runs on every call)
|
|
# corr_smart_model — tiebreaker (only fires when rules and cheap LLM disagree)
|
|
# Both IDs below were retired by Google and returned 404 on every call from
|
|
# some point before 2026-08-18 until they were corrected. Because a failed
|
|
# LLM call falls back to the rules decision, nothing broke loudly -- the
|
|
# entire LLM tier and the consensus tiebreak were simply dead in production
|
|
# while correlation behaviour was being tuned against rules-only output.
|
|
# Verify against https://ai.google.dev/gemini-api/docs/models before changing.
|
|
corr_cheap_model: str = "gemini-3.6-flash" # was gemini-2.0-flash (shut down)
|
|
corr_smart_model: str = "gemini-2.5-pro" # was gemini-1.5-pro (shut down)
|
|
# Transcript correction (server-26#36). Runs inside transcription, once per
|
|
# transcribed call above MIN_WORDS_FOR_CORRECTION, so it is priced like STT
|
|
# rather than like the correlation tier — cheap model on purpose.
|
|
transcript_correction_enabled: bool = True
|
|
transcript_correction_model: str = "gemini-3.6-flash"
|
|
# Retry Whisper once when its output is degenerate. The same clip produced a
|
|
# 56-word ten-code counting run on one attempt and real speech on the next
|
|
# (2026-08-23, call e49ea32c), so a hallucination is a coin-flip rather than
|
|
# a property of the audio, and discarding on the first bad roll threw away a
|
|
# recoverable transcript.
|
|
stt_retry_on_degenerate: bool = True
|
|
summary_interval_minutes: int = 2 # how often the summary loop runs
|
|
correlation_window_hours: int = 2 # slow/location path: max hours since last call
|
|
embedding_similarity_threshold: float = 0.93 # slow-path: requires location corroboration
|
|
embedding_no_location_threshold: float = 0.97 # slow-path: match without location (very high bar)
|
|
embedding_cross_tg_threshold: float = 0.85 # cross-TG path: same dept + 2+ shared units
|
|
location_proximity_km: float = 0.5 # radius for location-proximity matching
|
|
geocode_max_km: float = 40.0 # reject geocode results farther than this from the node
|
|
incident_auto_resolve_minutes: int = 90 # auto-resolve after N minutes with no new calls
|
|
unit_continuity_max_idle_minutes: int = 20 # unit-continuity path: skip if incident idle > this
|
|
recorrelation_scan_minutes: int = 60 # re-examine orphaned calls ended within this window
|
|
tg_fast_path_idle_minutes: int = 90 # fast path: max minutes since incident last updated
|
|
# Dispatch channels only: tier-2 thin calls attach to a lone candidate idle < this.
|
|
# Was 10, which is long enough for the channel to have moved on to something else:
|
|
# on 2026-08-16 a "72 at Holland Station" incident absorbed a Grand Central train
|
|
# meet 9.6 min later, and a status check absorbed a records lookup at 9.7 min.
|
|
# Across that dump every correct thin attach was <= 3.4 min idle and every wrong
|
|
# one was >= 8.2, so 5 separates them with room on both sides. Genuine
|
|
# back-and-forth is handled by the 30-second tier-1 path above this.
|
|
tg_dispatch_thin_idle_minutes: int = 5
|
|
# Every other channel: tier-2 thin calls attach to a lone candidate idle < this.
|
|
# Non-dispatch talkgroups previously had NO tier-2 bound at all — they used the
|
|
# whole 90-minute tg_fast_path_idle_minutes window with no single-candidate
|
|
# requirement and no fit test, which is the widest version of the 2026-08-20
|
|
# over-merge. A tactical channel really is dedicated to one scene, so it earns
|
|
# a longer window than a dispatch backbone, but not an unbounded one: 15 sits
|
|
# inside the 20-minute tactical-default window in _call_fits_incident, so the
|
|
# no-evidence thin path is never more permissive than the fit-tested path on
|
|
# the same channel.
|
|
tg_thin_idle_minutes: int = 15
|
|
|
|
# ── Hard caps: an incident past either of these stops accepting calls ──────
|
|
# Enforced on every correlation path (see _incident_at_capacity). Pairwise fit
|
|
# tests judge one call against one incident and cannot see the shape of the
|
|
# chain they are building, so these are the only guard against a "work shift"
|
|
# incident regardless of how individually plausible each link looked.
|
|
#
|
|
# 120 minutes: the one incident in the 2026-08-20 dump that was genuinely a
|
|
# single event ran 63 minutes (06:15 wrong-way driver → 07:18 closeout), so
|
|
# the cap has to clear an hour with real headroom. The four junk chains ran
|
|
# 3h41m, 3h43m, 4h05m and 4h09m, so it has to sit well under three hours.
|
|
# 120 also equals correlation_window_hours: the location and slow paths
|
|
# already refuse to consider a candidate older than that, and the fast path
|
|
# was the only one exempt. Making it agree removes that inconsistency rather
|
|
# than inventing a new number.
|
|
incident_max_duration_minutes: int = 120
|
|
# 40 calls: a backstop for a burst that fills up inside the duration cap
|
|
# rather than the primary bound. The worst observed chain averaged ~16
|
|
# calls/hour while absorbing an ENTIRE dispatch backbone, so 40 calls in
|
|
# under two hours means the incident is eating most of the channel — that is
|
|
# a chain, not an event. Set deliberately above any plausible single-incident
|
|
# call volume (a multi-alarm fire on its own tactical channel) so this cap
|
|
# errs toward keeping real incidents whole and lets the duration cap do the
|
|
# cutting.
|
|
incident_max_calls: int = 40
|
|
|
|
# Vocabulary learning
|
|
vocabulary_induction_interval_hours: int = 24 # how often the induction loop runs
|
|
vocabulary_induction_sample_tokens: int = 4000 # ~tokens of transcript text sampled per system
|
|
|
|
# Internal service key — allows server-side services (discord bot) to call C2 without Firebase
|
|
service_key: Optional[str] = None
|
|
|
|
# Fleet-wide token edge nodes present to POST /nodes/enroll on first boot.
|
|
# Not a per-node secret — see routers/enrollment.py for why a leaked copy
|
|
# of this alone can't steal an already-approved node's key.
|
|
enrollment_token: Optional[str] = None
|
|
|
|
# Upload size limit — reject audio files larger than this (bytes). Default 100 MB.
|
|
upload_max_bytes: int = 100 * 1024 * 1024
|
|
|
|
# Public origin this API is reachable on, e.g. "https://api.drb.example.com".
|
|
# Only used to build absolute call-audio playback links: an <audio src> is
|
|
# fetched by the browser directly, so a relative path would resolve against
|
|
# the frontend origin, not this one.
|
|
public_api_url: Optional[str] = None
|
|
|
|
# How long a minted call-audio playback link stays valid. Long enough for a
|
|
# browsing session, short enough that a copied link isn't durable access.
|
|
audio_link_ttl_seconds: int = 6 * 60 * 60
|
|
|
|
# Two nodes hearing the same transmission start recording within about a
|
|
# second of each other (measured across node-002/node-PI-2 on TG 9048).
|
|
# 10s is generous against clock skew while staying well under the gap
|
|
# between genuinely separate transmissions on a busy dispatch channel.
|
|
duplicate_window_seconds: int = 10
|
|
|
|
# CORS — set to your frontend origin(s) in production, e.g. ["https://app.example.com"]
|
|
# Defaults to "*" for local development only.
|
|
#
|
|
# Leaving this as "*" is not merely permissive: main.py turns OFF
|
|
# allow_credentials when it sees a wildcard, because Starlette would
|
|
# otherwise reflect each caller's origin back WITH
|
|
# Access-Control-Allow-Credentials. So a production deployment that
|
|
# forgets to set this gets a loud ERROR at startup and loses credentialed
|
|
# cross-origin requests, rather than silently accepting every origin.
|
|
cors_origins: list[str] = ["*"]
|
|
|
|
# Discord webhook URL that app/internal/ai_health.py posts to when an AI
|
|
# tier (transcription/correlation) transitions into or out of degraded
|
|
# state. Empty disables the POST entirely — not every self-hosted
|
|
# deployment will set this up, and skipping it must be silent.
|
|
ai_alert_webhook_url: str = ""
|
|
|
|
class Config:
|
|
env_file = ".env"
|
|
|
|
|
|
settings = Settings()
|