Both Gemini model IDs had been retired by Google. Production logs show every correlation call 404ing -- "models/gemini-2.0-flash is no longer available" -- and gemini-1.5-pro is gone from the model list as well. Because a failed LLM call falls back to the rules decision by design, nothing surfaced: the pipeline kept producing incidents, so the LLM tier and the consensus tiebreak were dead in production for an unknown number of days while correlation was being tuned. Some of what recent tuning was reacting to was rules-only behaviour that was never meant to run alone. Cheap model becomes gemini-3.6-flash, which is the migration target named in Google's own 404. Smart model becomes gemini-2.5-pro, the only stable Pro-tier text model left; the tiebreak fires rarely and its value comes from being a different, stronger model than the first pass, so a second Flash was not worth the consensus it would give up. Model list checked against https://ai.google.dev/gemini-api/docs/models on 2026-08-18. The more important half is the logging. A per-call WARNING was the only signal, and it is indistinguishable from an ordinary API hiccup, so a permanent misconfiguration read as noise. Failures that look like a missing model (404, "not found", "no longer available") now log once per model at ERROR, name the config keys to change, and say plainly that correlation is running rules-only. Transient errors keep the old per-call WARNING. Once per model, not once per call, so the alert stays readable at radio traffic volume. Gemini is used nowhere else in c2-core -- extraction, embeddings and summaries all run on OpenAI -- so the blast radius was exactly the correlation LLM tier. 38 correlator tests still pass. No new environment variables: both model IDs are config.py defaults and are not templated into any .env, so CI deploys this without an ansible run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
109 lines
5.7 KiB
Python
109 lines
5.7 KiB
Python
from pydantic_settings import BaseSettings
|
|
from typing import Optional
|
|
|
|
|
|
class Settings(BaseSettings):
|
|
# MQTT
|
|
mqtt_broker: str = "localhost"
|
|
mqtt_port: int = 1883
|
|
mqtt_user: Optional[str] = None
|
|
mqtt_pass: Optional[str] = None
|
|
|
|
# mosquitto's built-in dynamic-security plugin (see app/internal/dynsec.py).
|
|
# "admin" is hardcoded by the plugin itself on first boot — not actually
|
|
# configurable — kept as a named setting rather than a literal for
|
|
# readability. mqtt_dynsec_admin_pass must equal the mosquitto
|
|
# container's own MOSQUITTO_DYNSEC_PASSWORD env var (root .env /
|
|
# root.env.j2) or c2-core can't administer node credentials at all.
|
|
mqtt_dynsec_admin_user: str = "admin"
|
|
mqtt_dynsec_admin_pass: Optional[str] = None
|
|
|
|
# GCP
|
|
gcp_credentials_path: Optional[str] = None # None → uses ADC
|
|
gcs_bucket: Optional[str] = None # None → audio upload disabled
|
|
firestore_database: str = "(default)"
|
|
|
|
# Node health
|
|
node_offline_threshold: int = 90 # seconds without checkin before marking offline
|
|
|
|
# OpenAI (STT + intelligence)
|
|
openai_api_key: Optional[str] = None
|
|
stt_model: str = "whisper-1" # whisper-1 | gpt-4o-mini-transcribe | gpt-4o-transcribe
|
|
|
|
# Google Maps (geocoding)
|
|
google_maps_api_key: Optional[str] = None
|
|
|
|
# Gemini (intelligence extraction, embeddings, incident summaries)
|
|
gemini_api_key: Optional[str] = None
|
|
# Correlation consensus models
|
|
# corr_cheap_model — first-pass LLM correlator (runs on every call)
|
|
# corr_smart_model — tiebreaker (only fires when rules and cheap LLM disagree)
|
|
# Both IDs below were retired by Google and returned 404 on every call from
|
|
# some point before 2026-08-18 until they were corrected. Because a failed
|
|
# LLM call falls back to the rules decision, nothing broke loudly -- the
|
|
# entire LLM tier and the consensus tiebreak were simply dead in production
|
|
# while correlation behaviour was being tuned against rules-only output.
|
|
# Verify against https://ai.google.dev/gemini-api/docs/models before changing.
|
|
corr_cheap_model: str = "gemini-3.6-flash" # was gemini-2.0-flash (shut down)
|
|
corr_smart_model: str = "gemini-2.5-pro" # was gemini-1.5-pro (shut down)
|
|
summary_interval_minutes: int = 2 # how often the summary loop runs
|
|
correlation_window_hours: int = 2 # slow/location path: max hours since last call
|
|
embedding_similarity_threshold: float = 0.93 # slow-path: requires location corroboration
|
|
embedding_no_location_threshold: float = 0.97 # slow-path: match without location (very high bar)
|
|
embedding_cross_tg_threshold: float = 0.85 # cross-TG path: same dept + 2+ shared units
|
|
location_proximity_km: float = 0.5 # radius for location-proximity matching
|
|
geocode_max_km: float = 40.0 # reject geocode results farther than this from the node
|
|
incident_auto_resolve_minutes: int = 90 # auto-resolve after N minutes with no new calls
|
|
unit_continuity_max_idle_minutes: int = 20 # unit-continuity path: skip if incident idle > this
|
|
recorrelation_scan_minutes: int = 60 # re-examine orphaned calls ended within this window
|
|
tg_fast_path_idle_minutes: int = 90 # fast path: max minutes since incident last updated
|
|
# Dispatch channels only: tier-2 thin calls attach to a lone candidate idle < this.
|
|
# Was 10, which is long enough for the channel to have moved on to something else:
|
|
# on 2026-08-16 a "72 at Holland Station" incident absorbed a Grand Central train
|
|
# meet 9.6 min later, and a status check absorbed a records lookup at 9.7 min.
|
|
# Across that dump every correct thin attach was <= 3.4 min idle and every wrong
|
|
# one was >= 8.2, so 5 separates them with room on both sides. Genuine
|
|
# back-and-forth is handled by the 30-second tier-1 path above this.
|
|
tg_dispatch_thin_idle_minutes: int = 5
|
|
|
|
# Vocabulary learning
|
|
vocabulary_induction_interval_hours: int = 24 # how often the induction loop runs
|
|
vocabulary_induction_sample_tokens: int = 4000 # ~tokens of transcript text sampled per system
|
|
|
|
# Internal service key — allows server-side services (discord bot) to call C2 without Firebase
|
|
service_key: Optional[str] = None
|
|
|
|
# Fleet-wide token edge nodes present to POST /nodes/enroll on first boot.
|
|
# Not a per-node secret — see routers/enrollment.py for why a leaked copy
|
|
# of this alone can't steal an already-approved node's key.
|
|
enrollment_token: Optional[str] = None
|
|
|
|
# Upload size limit — reject audio files larger than this (bytes). Default 100 MB.
|
|
upload_max_bytes: int = 100 * 1024 * 1024
|
|
|
|
# Public origin this API is reachable on, e.g. "https://api.drb.example.com".
|
|
# Only used to build absolute call-audio playback links: an <audio src> is
|
|
# fetched by the browser directly, so a relative path would resolve against
|
|
# the frontend origin, not this one.
|
|
public_api_url: Optional[str] = None
|
|
|
|
# How long a minted call-audio playback link stays valid. Long enough for a
|
|
# browsing session, short enough that a copied link isn't durable access.
|
|
audio_link_ttl_seconds: int = 6 * 60 * 60
|
|
|
|
# Two nodes hearing the same transmission start recording within about a
|
|
# second of each other (measured across node-002/node-PI-2 on TG 9048).
|
|
# 10s is generous against clock skew while staying well under the gap
|
|
# between genuinely separate transmissions on a busy dispatch channel.
|
|
duplicate_window_seconds: int = 10
|
|
|
|
# CORS — set to your frontend origin(s) in production, e.g. ["https://app.example.com"]
|
|
# Defaults to "*" for local development only.
|
|
cors_origins: list[str] = ["*"]
|
|
|
|
class Config:
|
|
env_file = ".env"
|
|
|
|
|
|
settings = Settings()
|