e30d594eeab08a7d23303e39f40305bbc4cd2fd9
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e30d594eea |
Stop a back-dated call from silently disabling every recency gate (server-26#74)
_call_fits_incident measured incident idle with the signed helper while every other recency gate in the file uses the unsigned one. On the re-correlation sweep, `now` is the call's own started_at, which can precede the incident's last activity, so the value went negative. Negative idle made `idle_min >= 15` false, which meant the content-divergence veto never ran and unit overlap was accepted unconditionally -- on a shared dispatch backbone that is the feedback loop that lets one incident absorb a whole talkgroup. It also made `idle_min < 20.0` true at any back-dating, so a tactical channel returned tactical_default for every swept orphan out to the 90-minute bound. One variable feeds all four gates in the function, so this is a one-line change at the source. The signed value is untouched where it belongs: callers still compute corr_incident_idle_min themselves, so debug output keeps its meaning. Direction is toward more splitting, on the sweep path only, which is the point -- the bug was suppressing an over-merge veto. Forward-dated calls and anything inside the thresholds behave exactly as before. Two tests added alongside the existing idle-gate cases; both fail on the old line and pass on the new one. 266 pass, 0 fail. Refs server-26#74, #5, #80. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
187b8c1500 |
Declare and create the calls(system_id, started_at) index dedup needs
Dedup was failing on essentially every inbound call in production. dedup.py queries system_id == X with a started_at range; that composite index was neither declared in firestore.indexes.json nor present in the live c2-server database, so the query returned FAILED_PRECONDITION, dedup swallowed it as a warning, and every duplicate check degraded to "not a duplicate". While AI is off that only cost duplicate call documents. With a window open it would have paid Whisper and Gemini twice for every double-heard transmission, and fed Gate B5's cost measurement a figure that is wrong for a reason unrelated to the pipeline being measured. Two documents for one transmission is also the exact input shape that produces a spurious second incident. The index is created on c2-server and building. This declares it in source so the file and the live database agree. Refs server-26#84, #33, #45. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d18e4f0743 |
Make "AI is off" true, and stop the transcript PATCH from destroying calls
config/ai_features was not the switch it was documented to be. Three paths spent money with it off, and one path read it wrong, so per-system opt-outs did not opt anything out. - Correlation in the ingest pipeline tested the raw global flag instead of the per-system resolution. With a system opted out, extraction was skipped but the no-scenes fallback still correlated the call with empty tags, taking the thin/recency path and attaching it to whatever incident was most recent on that system. The opt-out did not disable correlation, it disabled good correlation and left the worst kind running. (#75) - Transcript correction ran on every transcribed call gated only by an env var, spending Gemini tokens and a Places lookup per proposed location. An "STT-only" window was never STT-only and its cost could not be attributed. Now behind transcript_correction_enabled. (#76) - _run_extraction_pipeline and the vocabulary learner, both reachable from PATCH /calls/{id}/transcript, checked no flags at all. (#76, #81) The flag resolver now lives in feature_flags.resolve_flags() rather than as a local helper in upload.py. Three copies of that logic is how #75 happened. PATCH /calls/{id}/transcript now refuses with 409 when correlation is off. That route wipes tags, severity, location, units, embedding and unlinks the call from every incident before queueing re-extraction. Gating extraction alone would have made it destructive-only in the standing flags-off configuration: the call left blank and orphaned forever, with the route still answering 200. The wipe and the rebuild are one transaction in intent, so it refuses before the first write. Also: the summarizer's stale-incident sweep is no longer behind summaries_enabled. It is pure Firestore with no model call in it, and gating it meant nothing auto-resolved while AI was off - so every incident stayed active forever and the candidate set every correlation reads kept growing. transcript_correction_enabled is documented as NOT a pure cost lever. The corrector is also the noise gate that sets not_speech; with it off, recogniser noise reaches extraction as a real transcript, comes back thin, and auto-attaches. Never open an evaluation window with correction off and correlation on. 14 tests added covering flag precedence, both pipeline paths, the 409, the correction gate and the summarizer no-op. Suite: 264 passed. Refs #75, #76, #81, #45. |
||
|
|
5fc4e2c57b |
Roll back a bad deploy instead of leaving it live (server-26#65)
deploy.yml ran `compose up -d` before the health check and never reverted on failure. A build that passes tests, returns 200 on /health with the right git_sha, but has a live logic bug (exactly the class of bug the correlator instrumentation exists to catch) would stay live indefinitely - notify-failure would even claim production was "still running the previous build", which is false in that scenario. Deploy step now reads /opt/drb/.last_good_tag (written only after a prior deploy's own health check confirmed its SHA) to capture the previously- verified tag before switching, and emits it as a step output. Health check is unchanged in shape (bounded 20x5s retry, still requires the polled git_sha to match) but now persists the new SHA as the rollback target only once confirmed live. A new Rollback step runs on any failure above, re-deploys the previous tag, and re-verifies via the same git_sha check rather than trusting mere liveness - then fails the job loudly either way, since the push itself was still bad. notify-failure now reports what actually happened (rollback succeeded/failed/skipped and to which SHA) instead of the old unconditional claim. This unblocks #62 decision 9: autonomous pushes to incident_correlator.py, llm_correlator.py, intelligence.py and routers/upload.py were frozen until this rollback path landed. Refs #65, #62, #60, #57. |
||
|
|
cc038e6326 |
A unit call-sign is not a place
"Post 1-2" reached the geocoder, resolved against its talkgroup anchor and produced a confident pin in the right town for an event with no known location — while sitting in the same incident's `units` list the whole time. A plausible wrong pin is worse than no pin: nothing downstream can tell it is wrong. Extraction returns `location` and `units` from one pass, so a string in both is a misclassification, not two facts. Drop it before the geocoder sees it. Closes server-26#52. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
964343c819 |
area_context v2 + Maps place verification (server-26#36, #37)
#36 — the correction pass shipped in
|
||
|
|
58efdbd6eb |
Correct the transcript before anything reads it
Correction existed, but as a line in intelligence.py's EXTRACTION_PROMPT --
which put it in the wrong place twice over. The same model call that extracted
units, location and severity emitted the correction afterwards, so extraction
reasoned over text already known to be wrong; and it sat behind
correlation_enabled, so during a cost-controlled STT-only window nothing was
ever corrected at all. That is the normal state during development.
internal/transcript_correction.py is now its own pass, between the degenerate
filter and the Firestore write. It receives an already-produced transcript plus
a reference list, so unlike a Whisper prompt it has no series to extend -- the
distinction that keeps vocabulary out of the recogniser's prompt, where an
enumerated ten-code list once made it hallucinate ten-code runs.
Reference data is merged from the talkgroup and the system, TALKGROUP FIRST. A
system spanning several counties can have a talkgroup covering one
municipality, and that municipality's streets must not be buried under a
county-wide list. A single-municipality system is the degenerate case: populate
the system level and every talkgroup inherits it. Area context is now SET --
municipality, county, roads, landmarks, on both scopes -- rather than guessed
from talkgroup names, which is what vocabulary_learner did and which is close
to useless across multiple counties.
Segments are corrected too, not just the joined text. extract_scenes builds its
prompt from numbered segments whenever there is more than one, so a correction
that only fixed the transcript would have been discarded on exactly the
multi-transmission calls carrying the most content. Alignment is enforced: an
array of the wrong length or type is dropped whole, because scenes map back to
transmissions by index and a shifted array would misattribute audio silently.
Whisper is also retried once on degenerate output. Call e49ea32c produced a
56-word ten-code counting run on one attempt and ordinary speech on the next --
same clip, same temperature=0 -- so a hallucination is a coin-flip, and
discarding on the first bad roll threw away a recoverable transcript.
Two things found on the way:
PUT /systems/{id} wiped ten_codes on every save. The systems form sends only
{name, type, config}, and model_dump() wrote every omitted field as its default
over the top. Now exclude_unset. area_context would have been the next victim,
which is why it gets its own route alongside ten-codes rather than a field on
that payload.
Closes server-26#36.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
1bfa856d1b |
Serve audio as whatever it actually is
audio/mpeg was hardcoded at both points call audio is written and served, from back when the node produced nothing but 16 kbps MP3. It now uploads FLAC, and a browser will not play a FLAC body labelled audio/mpeg. storage.py grows one extension -> Content-Type map, used by the GCS upload and by /media. Keyed off the object's real extension, so every existing .mp3 recording keeps working with no migration -- and _safe_audio_filename already accepted .flac, so object naming needed nothing. Also flags what this costs: /media sends the whole body with Accept-Ranges: none, which was fine at ~60 KB per call and is not fine at ~1.3 MB/min. Noted at the header and in DEFERRED.md, whose stated reason for deferring Range support was the old file size. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
457e6d7e0f |
Make Archive a real page instead of a redirect
/calls was a ten-line stub that redirected to /incidents, so there was nowhere in the app to look at a call. The nav's "Archive" link led to the incident list, and a call that never correlated was invisible entirely -- which is backwards when correlation quality is the thing under development, because the orphans are the evidence. Its stated blocker (Gitea #17/#18) closed weeks ago. The page browses the org's calls newest-first over the new /calls/search route, filtered by link state (all / orphans / linked), transcript presence, and system, with a transcript substring search and cursor paging. A row expands to the full transcript, a playback link minted on demand, and the correlation path that decided it. The counts line -- how many of the loaded calls are orphaned, how many have no transcript at all -- is the number worth watching during an AI window. Attribution is the point of it: attach an orphan to the incident it belongs to, or detach one the correlator got wrong. Both go through the routes fixed in the previous commit, so a manual attachment now actually shows up on the incident. Admin-only. It exposes every call in the org regardless of node ownership and carries controls that rewrite incident membership. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
140dfbfc74 |
Give the archive a real read, and the debug view a verdict
Three backend pieces the /calls page needs, plus the fix for a debug view that
hid its data exactly when it was wanted.
GET /calls/search — paged, filterable call archive. GET /calls returns every
call in one unordered shot: fine for a node's handful of active calls, useless
as an archive. Only the org scope and the started_at ordering go to Firestore,
since that pair is the one composite index that exists; the rest filters in
Python over a bounded window, the same shape admin.py's debug route uses. The
cursor advances over the scanned window rather than the returned page, or a
sparse filter would re-scan from the same place forever.
Manual attribution. POST /incidents/{id}/calls/{id} only ever wrote the legacy
scalar incident_id, never incident_ids -- which is what the correlator writes
and what the frontend queries with array-contains. A manually attached call was
therefore invisible on the incident page it had just been attached to. It now
maintains both and marks the summary stale. DELETE is new: there was no way to
undo an attachment at all, so a wrong link was permanent.
The debug view no longer filters to AI-enabled systems by default. That filter
emptied the view the moment the flags went off, which is precisely when a
window gets reviewed -- on 2026-08-23 it fell from 100 incidents to 6 between
switching correlation off and opening the tab. ai_systems_only=true restores it.
It also returns a summary block now: corr_path / fit_signal / consensus /
llm_action tallies, transcript coverage on both linked and orphaned calls,
single-call and median-calls-per-incident for fragmentation, max span and
anything past the server-26#22 caps for merging, and the count of incidents
still carrying a fallback "— TGID" title. All of it was being recomputed by
hand from the raw payload on every review.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
7ef5704be2 |
Make firestore.indexes.json describe the database again
The file had drifted four indexes behind c2-server, so the 2026-08-23 deploy offered to delete four live indexes and then added ASC copies of two that already existed as DESC. Reconciled against gcloud's actual list. Adds the two backend indexes that were live but undeclared and are genuinely in use -- calls(status, ended_at) for recorrelation_sweep's ended-call scan and calls(system_id, ended_at) for vocabulary_learner. Deleting either would have broken a background loop with no frontend symptom. Declares every index ASCENDING. Firestore scans an index in either direction, so org_id+started_at ASC already serves the orderBy(started_at, 'desc') that every frontend hook actually asks for; a matched ASC/DESC pair is one index of pure write amplification on every call document. The three duplicates now left undeclared are named in the file header so the next deploy's interactive delete prompt has a documented answer instead of a guess. Refs server-26#33. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
039a06dc72 |
Let C2 name a talkgroup it already knows
84 of the 100 incidents in the 2026-08-23 dump were titled "Ems — TGID 9048"
or "Other — TGID 9600" -- the fallback, not a description. The title is the
incident's name everywhere it appears: list rows, map pins, Discord alerts.
_create_incident builds it from a content tag and a talkgroup label, and the
label was collapsing to "TGID {id}" because talkgroup_name arrived as None.
It is a plain form field on /upload, forwarded untouched into correlation, and
the node only sends it when OP25 had the name in its loaded tags file -- which
is exactly the case C2 can cover from its own systems collection, where all 125
talkgroup definitions live.
The lookup already existed, on the other path: mqtt_handler resolved it from
the system config on call_start. So the call document held the right name while
the pipeline that titles the incident ignored it. That asymmetry is the bug.
internal/talkgroups.py is now the one implementation -- caller's hint, then the
call document, then the system config -- and both paths use it.
_run_intelligence_pipeline resolves once at the funnel /upload and
/calls/{id}/reprocess share, so the dispatch-channel test, scene extraction and
the title all see a real name. When the call document was the thing missing it,
the resolved name is written back, so the archive and the orphan panel stop
showing a bare TGID too.
Also gives fast/thin a corr_fit_signal. It is 63% of all links and was the only
path writing none, so corr_fit_signal was absent on 295 of 309 calls and the
admin debug view's distribution panel read empty -- looking broken when it was
faithfully reporting that the dominant path records nothing. It now says
thin_recency, which is what actually decided it.
Closes server-26#34. Refs server-26#35 -- the tier's 3.5% invocation rate is a
cost/benefit question, not a bug, and stays open.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a278e2215a |
Pin the Firestore deploy target to the c2-server database
firebase.json declared rules and indexes with no database key, so the CLI deploys them to (default). This project does not use (default) -- c2-core reads FIRESTORE_DATABASE and the frontend reads NEXT_PUBLIC_FIRESTORE_DATABASE, both c2-server in production, and the index-required errors the browser prints name /databases/c2-server/ outright. A deploy without this key reports success and changes nothing the app can see, which is a bad way to find out. Refs server-26#13. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
be79499635 |
Give the nav's dead links somewhere to land
Three of the app's routes were referenced but never existed, so the redesign's
navigation pointed at 404s from several directions.
/dashboard was the post-login and fallback redirect target in nine places --
login, onboarding, middleware, the admin/nodes/systems/tokens/settings guards,
and the marketing header -- but app/dashboard/ was never created. Signing in
normally dropped the user on a 404. The real signed-in home is "/", which
app/page.tsx already renders as LiveView for an authed user with an org, and
which the nav labels "Live"; all nine now point there.
Nav also linked /watch and /network, neither of which existed. /watch is the
alerts screen under its redesign name, so it re-exports app/alerts/page.tsx
and /alerts stays reachable for old links. /network is new: the "my equipment"
hub the redesign moved /nodes, /systems and /tokens behind and then never
built, which had left /systems and /tokens with no entry point in the UI at
all. Its hooks all run before the admin/operator guard, per
|
||
|
|
861ea41cec |
Deploy the commit's own images instead of :latest
The build-stamp health check added in
|
||
|
|
82c88379d4 |
Stop an incident lying about what it is and where it is
An incident header had two independently last-write-wins halves, and in the
2026-08-20 dump both were wrong at once. `b9b4f392` opened on a suspect search
at 80 Grasslands Road; it was labelled "100 South Mosher" (its third call),
pinned at `Westmed` (its second), and titled after the label. Five of six
incidents were pinned somewhere other than the place they claimed to be.
Location and pin are now one value
----------------------------------
`_resolve_location_pair()` computes `location`, `location_coords` and the new
`location_coords_source` together, and `_update_incident`/`_create_incident`/
`_create_master_incident` always write all three. There is no longer a code
path that can move one and leave another behind — including the cross-system
master, which used to take its label from the parent and its pin from the call.
The pin now carries the label it was geocoded from. `_verified_pin()` returns
it only when that source still matches the incident's current label; anything
else is dropped. That includes every pre-existing incident, whose pin has no
recorded source and therefore cannot be reconciled — which is the right
outcome, since the dump says 5 in 6 of those are wrong. A missing pin reads as
missing data; a wrong pin reads as fact, and this is a map people may act on.
An incident also keeps the first place it was given rather than the latest.
Later mentions still accumulate in `location_mentions` (what the map path is
drawn from); they just don't rename the incident's own location. The one
permitted change is filling in a pin the incident never had, from a later call
naming the exact same label — geocoding needs the node position, a quota and a
response, so the same address genuinely does fail once and resolve later.
"49" is not a place
-------------------
`clean_location()` rejects any string with no two-letter word in it, applied at
extraction (intelligence.py, before the geocoder and before the call document)
and again at the correlator's context boundary. `9d376ffe` carried
`location: "49"` from "Fire received. Flames from 49." — a box number — and its
summary asserted "A fire incident was reported at location 49". Nothing
validated that field at all, so it would have recurred.
Title: the founding event, escalation only
------------------------------------------
The title was re-derived from the newest classified call, so `f5190670` was
named after the thirteenth of its thirteen events. It now names the call that
opened the incident, recorded in `title_tag`/`title_severity`, and can only be
replaced by a call of strictly higher severity.
Three candidates were considered:
* Newest call (status quo) — rejected. The same incident has a different name
at different times, so a user who saw it in the rail cannot find it again,
and the name is decided by radio timing rather than by the event.
* Highest severity alone — rejected as the sole rule. Severity has four
levels and most traffic sits on one of them, so ties are the common case
and the tiebreak degrades to "newest" — the defect it was meant to fix.
* Founding event, escalated by strictly-greater severity — chosen. An
incident's identity is the event that opened it, so that is its default
name and it is stable for the incident's whole life. The single case where
the header MUST change is the one where the situation got worse: a check
condition that becomes a structure fire is a structure fire, and the
worst-first rail, the "Major only" filter and the map colour all exist so
that is never missed. Requiring strictly-greater makes it monotonic, the
same contract `_max_severity` already gives the severity field: routine
chatter can never take the name back.
A summary-level title regenerated as a whole was rejected outright: it needs an
LLM call per incident, AI flags are off in production, and every incident today
would have no title at all.
Two renames survive, because neither replaces an event name: filling in the
placeholder title of an incident that opened on a call with no content tags
("Police — Ch 1"), and re-rendering the same event once the incident learns its
address. Incidents created before this change have no `title_tag`, so their
existing title is treated as the founding one rather than handed to whichever
call links next.
Interaction with the caps from
|
||
|
|
c7be6416f2 |
Surface LLM correlation fields in debug view; fix unit-continuity path
/admin/debug/correlation stripped corr_consensus and the corr_llm_* fields
that upload.py's consensus correlator writes onto the call doc, making it
the one tool built to answer "is the LLM correlation tier alive" unable to
answer it (2026-08-19 dump had to infer LLM state from commit dates instead
of reading it off the data). admin.py's _call_summary() now includes
corr_consensus, corr_llm_reasoning, corr_llm_action, corr_rules_action.
The unit-continuity correlation path never wrote corr_matched_units, unlike
fast/single and fast/disambig, so the debug view showed null for a match
that was in fact unit-driven by construction. Now populated unconditionally
on that path (server-26#16).
Also traced the negative corr_incident_idle_min (-4.1 observed) to its root
cause: the re-correlation sweep anchors `now` to the linking call's own
started_at, and that back-dated value was being written straight into the
incident's updated_at, letting it land before the incident's own
started_at. Added _floor_at_started_at() so updated_at can never precede
started_at. (commit
|
||
|
|
8fbfe7d6de |
Make a failed deploy impossible to miss, and a wildcard CORS harmless
Two unrelated-looking problems with the same shape: a dangerous state that
looked fine from the outside.
DEPLOY (server-26#21). The Deploy job failed on fifteen consecutive pushes
between 2026-08-18 and 08-20 and nobody noticed for two days, because the
build job was green and a red run is only visible to someone who opens Gitea.
Production served 08-18 code the whole time -- including the entire frontend
redesign, chunks 2 through 8. Three changes:
* The health check now asserts WHICH build answered, not just that something
did. CI bakes the commit into the image (Dockerfile ARG/ENV GIT_SHA) and
/health reports it, so a deploy that "succeeds" while the previous
container keeps running now fails. Liveness alone could never have caught
this.
* The image pull retries once after a prune. The actual failure was
containerd unable to extract a layer -- "failed to Lchown ... no such file
or directory" -- a corrupted entry in the snapshot store, which a prune
clears. A second failure after pruning is a real problem (check the VM's
disk) and still stops the deploy.
* A notify-failure job POSTs to DEPLOY_ALERT_WEBHOOK when anything in the
workflow fails. Unset means skip quietly, not fail.
CORS (server-26#20). allow_origins=["*"] with allow_credentials=True is not
the permissive-but-harmless setting it reads as. Starlette does not reject the
pair -- it reflects the caller's Origin back and still sends
Access-Control-Allow-Credentials: true, so the effective policy is "any
origin, WITH credentials", the opposite of what a wildcard normally means.
Rather than trust every deployment to remember CORS_ORIGINS, the pair is now
unrepresentable: a wildcard forces allow_credentials off and logs an ERROR
naming the variable to set. Correctly configured deployments that name their
origins are unaffected and keep credentialed requests.
Severity honestly: low today. c2-core is bearer-auth, and browsers do not
attach bearer tokens cross-origin the way they attach cookies. This is a
misconfiguration waiting for the day something starts trusting a cookie.
Also adds firebase_admin.auth.UserRecord and the list/update/create/delete_user
names to the conftest stub. routers/users.py annotates with UserRecord at
import time, so without it importing app.main failed at collection -- which is
why nothing had ever tested anything wired at app level, CORS included.
Tests: 5 new in test_cors_policy.py, covering the pure policy function, the
middleware actually mounted on the app (so re-hardcoding allow_credentials=True
fails here), and the presence of the build stamp.
Closes logan/server-26#20
Closes logan/server-26#21
|
||
|
|
33a247d306 |
Stop thin calls fusing a work shift into one incident (server-26#22)
The 2026-08-20 production dump had 4 of 6 sampled incidents as junk chains,
the worst being f5190670: 68 calls over 4h09m, 44 units, 12 tags, at least
13 genuinely distinct events. 58 of 133 linked calls took the fast/thin
path, which is the one path that attaches a call with no fit test at all.
Three defects combined to produce that, and all three are fixed here.
1. What counted as thin was wrong.
is_thin_call was "not units and not vehicles and not coords". A real
dispatch qualified as thin whenever no unit ID parsed and the geocode
failed - six of them did in that dump, including "All units head over to
the powerhouse, 55 Hyman Hills Road ... she's 87 years old", a brand new
job that attached to the four-hour chain and then overwrote its location
and its title. A call is now substantive if it carries tags, a location
string, a severity above routine, or is a reassignment; only genuinely
content-free housekeeping ("10-4", "Copy") stays thin. Those calls now go
through _call_fits_incident like everything else, which on a dispatch
backbone with no positive signal means they open their own incident or
orphan rather than merging.
The reassignment clause closes a self-defeating guard: upload.py blanks
units when dispatch pulls a unit onto a NEW job, specifically to stop
unit-overlap chaining - and blanking units made the call thin, routing it
to the only path with no fit check. The guard produced the merge it
existed to prevent.
2. The thin path was bounded on dispatch channels only.
Every other talkgroup fell through to "thin_pool = tg_recent": any
incident idle up to tg_fast_path_idle_minutes (90), no single-candidate
requirement, no fit test. The 30-second tier-1 / single-candidate tier-2
structure now applies to all channels. Non-dispatch gets its own window,
TG_THIN_IDLE_MINUTES=15, rather than sharing the dispatch value: a
tactical channel really is dedicated to one scene so it earns longer, but
15 sits inside the 20-minute tactical-default window already used in
_call_fits_incident, so the no-evidence path is never more permissive than
the fit-tested path on the same channel.
Recency gates now compare the magnitude of the idle, not the signed value.
The re-correlation sweep anchors "now" to the call's own started_at, so
idle goes negative routinely - incident 9d376ffe recorded
corr_incident_idle_min: -4.1 - and every "idle <= window" test in this
module reads True for a negative number. Those gates had silently stopped
bounding anything for exactly the calls the sweep re-examines.
3. Nothing capped an incident's total size.
Every fit test in the correlator is pairwise: does this call belong with
that incident. Each of f5190670's 68 links was individually arguable; the
mistake was the accumulated shape, which no pairwise rule can see. Two
hard caps now remove an incident from the candidate pool entirely, before
any path can choose it - including the LLM tier, which reads the same
ctx lists.
INCIDENT_MAX_DURATION_MINUTES=120. The one incident in that dump that was
genuinely a single event ran 63 minutes (06:15 wrong-way driver to 07:18
closeout), so the cap has to clear an hour with real headroom. The four
junk chains ran 3h41m, 3h43m, 4h05m and 4h09m, so it has to sit well under
three hours. 120 also equals correlation_window_hours: the location and
slow paths already refuse a candidate older than that, and the fast path
was the only one exempt, so this removes an inconsistency rather than
inventing a number.
INCIDENT_MAX_CALLS=40. A backstop for a burst that fills up inside the
duration cap, not the primary bound. The worst chain averaged ~16
calls/hour while absorbing an entire dispatch backbone, so 40 calls in
under two hours means one incident is eating most of the channel. Set
deliberately above any plausible single-incident call volume (a
multi-alarm fire on its own tactical channel) so this cap errs toward
keeping real incidents whole and lets the duration cap do the cutting.
Capping is not truncation: the incident keeps every call it has and still
auto-resolves on the normal idle sweep. It just stops being a candidate.
Every ambiguous call here was resolved toward a separate incident rather
than a merge. A wrongly-separate incident is visibly wrong and can be
merged later; a wrongly-merged one silently corrupts every unit, tag,
severity and map pin on the incident it joined, and poisons the AI
summary written from them. The cost is some acknowledgements orphaning
instead of riding along on an incident, which is a small, visible loss.
Deliberately NOT changed, since both push toward more merging while the
current failure mode is entirely over-merging (every incident in the dump
has exactly one "new" call; there is no over-splitting left to trade
against):
- unit-overlap positive feedback on shared dispatch channels, which is
now bounded by the caps rather than fixed at its root
- the sweep retry budget expiring before the target incident exists
Tests: 31 new cases in tests/test_correlator_merge_caps.py, including a
replay of the f5190670 night - 13 unrelated jobs at their real offsets,
plus roster unit traffic and acknowledgements every two minutes. Without
the caps that traffic still builds a 125-call incident spanning 244
minutes; with the old thinness test on top, 153 calls over 247 minutes in
3 incidents. With this commit it is 13 incidents, largest 40 calls over 80
minutes. Each new case was checked to fail when the behaviour it covers is
reverted. Suite: 138 passed.
No AI feature flag was touched; correlation stays off in production.
Closes logan/server-26#22
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
baa9d1811f | Pin *.sh to LF so Windows checkouts cannot ship a CRLF shebang | ||
|
|
a250c29e3c |
Add AI provider degradation registry and alerting (server-26#14)
Three AI dependency failures in one night (retired Gemini model IDs, depleted Gemini balance, unpayable OpenAI account) each surfaced only as a single ERROR log line that nobody was watching. Add app/internal/ai_health.py, a shared in-memory registry that transcription.py and llm_correlator.py report into on every call (success and failure), distinguishing permanent conditions (dead model, dead billing) which alert immediately from transient ones (rate limits, network blips) which only alert after they persist. Alerts POST once per degradation episode and once on recovery to an optional Discord webhook (AI_ALERT_WEBHOOK_URL), reusing alerter.py's httpx pattern. State is exposed unauthenticated at GET /health/ai alongside the existing /health. Closes logan/server-26#14 |
||
|
|
5355095c48 |
Compare node API keys in constant time on /upload
/upload compared the per-node API key with a plain !=, which short-circuits on the first differing byte and so leaks a little information about how much of a guess was correct. The reason to fix it is less the timing channel itself -- an HTTP round trip is noisy -- than the inconsistency: enrollment.py and dynsec.py both went out of their way to use secrets.compare_digest for the same class of credential, so the codebase contradicted itself on whether this mattered. Now it does not. Also coalesces a missing api_key field to "" so compare_digest is never handed None, which would raise TypeError and turn a malformed node_keys document into a 500 instead of a 401. Closes logan/server-26#12 |
||
|
|
6dfa5bc66d |
fix: repair 10 stale tests in test_mqtt_handler.py and test_node_sweeper.py
All 10 failures were tests that had drifted behind the product code, not
regressions in it. Diagnosed each individually:
test_mqtt_handler.py:
- test_checkin_creates_new_node, test_checkin_new_node_defaults_lat_lon:
unpacked 4 positional args from doc_set.call_args[0], but
fstore.doc_set(collection, doc_id, data, merge=False) always passes
merge as a kwarg, so only 3 positional args are ever recorded. Fixed
the unpack to 3.
- test_call_start_creates_call_doc, test_call_start_uses_now_when_started_at_missing:
mocked fstore.doc_get, but _on_call_start looks the node up via the
cached fstore.doc_get_cached (added when Firestore reads were cut to
stay in the free tier). The unmocked doc_get_cached returned a bare
MagicMock, which isn't awaitable. Mocked doc_get_cached instead; also
fixed the same 4-vs-3 positional-arg unpack on doc_set's merge=False call.
- test_call_end_updates_status_and_times, test_call_end_sets_audio_url_when_present:
mocked fstore.doc_update, but _on_call_end now writes via
fstore.doc_set(merge=True) (see the "Fix Upload 404 warning" commit —
doc_update raised "No document to update" when call_end arrived before
call_start). Also calls doc_get_cached to stamp org_id. Mocked
doc_get_cached and asserted against doc_set instead of doc_update.
test_node_sweeper.py:
- test_stale_online_node_marked_offline, test_stale_recording_node_marked_offline,
test_tz_naive_last_seen_is_handled, test_only_stale_nodes_updated_in_batch:
_sweep() now calls app.routers.tokens.release_token(node_id) for every
node it marks offline (added in
|
||
|
|
4919b02238 |
Frontend redesign chunk 8: incidents browse
Rewrite app/incidents/page.tsx per UI_REDESIGN.md chunk 8. Replaces the old active/resolved two-table split with a single timeline-grouped list (Today / Yesterday / date), each row using the same rail-card anatomy as Live's incident panel — severity spine + type glyph + severity chip + ON AIR pill (from useActiveCalls, matching a call's incident_ids against the row) + title + location + on-scene unit chips + age/call-count — so status is a chip on the row instead of a section boundary, and Live/Incidents visibly read as the same object at two densities. Severity filter and sort are unchanged. The create-incident modal and resolve action are unchanged. Per UI_REDESIGN.md chunk 8. |
||
|
|
4b5cf1971e |
Frontend redesign chunk 7: incident detail rebuild
Rewrite app/incidents/[id]/page.tsx to UI_REDESIGN.md §5.2. Header is now type glyph + SeverityMark + active/resolved chip + a 27px title, with elapsed time, path length (haversine sum over geocoded calls) and call count as a single subline. Summary is promoted out of the old tab into a first-class prose block (16.5px/1.58) — it's the artifact the product sells, so it gets the best position instead of competing with Units/ Details behind a click. Units/Details tabs are gone; On scene / Cleared render directly from units_active/units_cleared (chunk 3), Vehicles below. New components/CallSpineEntry.tsx replaces CallRow for this page (CallRow stays for the Archive table until chunk 12): time-ordered entries with a numbered stop marker that matches the map's path stops via the same sort-by-started_at-over-geocoded-calls index MapView's IncidentPathLayer uses — the "shared index" from §2.4. Includes an inline play/scrub audio player (lazy-fetches the signed URL on first play, same pattern CallRow already used), transcript in sans prose instead of a font-mono <pre>, unit/ cleared-unit chips, and a paginating "N earlier calls" control. Thin/ status-only calls collapse to one line. The incident map keeps the location_coords guard and now passes `calls` through to MapView so its path polyline (chunk 5) renders here too. Per UI_REDESIGN.md chunk 7. |
||
|
|
bc636c00ce |
Frontend redesign chunk 6: Live view
New components/LiveView.tsx renders the default landing at "/": full-bleed MapView (rail + legend from chunk 5) plus a new TimeScrubber strip below it — real call-density bars over the selected 1h/6h/24h/7d window, tinted by the worst severity in each bucket, playhead pinned to NOW. The playhead doesn't scrub yet; that needs `resolved_at` on incidents, which doesn't exist server-side (blocked chunk 13, in DEFERRED.md) — the density data itself is live, not a fixture. Distinguishes the two empty states UI_REDESIGN.md §4 calls out: a configured-but-quiet org (nodes online, zero active incidents) now shows "Listening — last check-in Xm ago" instead of rendering nothing, separate from the zero-node case (chunk 10's Activation screen). app/page.tsx's HomePage now renders LiveView directly for a signed-in, provisioned user instead of the chunk-4 interim redirect to /incidents. Per UI_REDESIGN.md chunk 6. |
||
|
|
31c0b3addf |
Frontend redesign chunk 5: MapView rewrite — draw the incident path
The flagship feature: a police pursuit has never been drawn as a path. Add an IncidentPathLayer that, for each incident, takes calls with location_coords (now declared on CallRecord as of chunk 3), sorts them by started_at, and draws a <Polyline> with numbered stop markers — first stop hollow, last stop haloed, using the same index the call spine will use in chunk 7 (UI_REDESIGN.md §2.4's "shared index"). Needs no backend; per-call geocodes are already written by intelligence.py. MapView takes a new optional `calls` prop (the caller's already-loaded recent calls) and groups them by incident_id internally, so it stays a pure presentation component. Retheme markers onto the §2.3 encoding: incident pins are a teardrop with the type glyph knocked out (from TypeGlyph's paths, duplicated as raw SVG since Leaflet icons are HTML strings, not React nodes), filled by severity colour and hollow-with-ink-stroke for minor/routine; node markers are NodeMark-style diamonds via a shared nodeDiamondSvg() helper, deleting statusColor() and all its green. Legend rebuilt shape-first (severity glyphs + node diamond weights, never a bare colour swatch) and reads correctly in both themes via the surface/ink tokens instead of the old bg-gray-950/90 that had no light mapping. Removed the three dead placeholder overlays (News Alerts, ADS-B, Meshtastic). Fan-cluster grouping (computeGroups) is unchanged. Per UI_REDESIGN.md chunk 5. |
||
|
|
eaae452d4e |
Frontend redesign chunk 4: navigation and routing
Rewrite Nav.tsx to the five-destination IA from UI_REDESIGN.md §3 (Live,
Incidents, Archive, Watch, Network) on tokens/sans type, with Settings,
Admin, Trips and Profile moved into the avatar dropdown instead of sitting
as nav peers. Network stays gated to admin/operator, matching the write
boundary its constituent pages (nodes/systems/tokens) already had.
Delete app/dashboard/page.tsx — its incident cards become the Live rail,
its node cards become Network, its call table becomes Archive; nothing on
it is unique. Add app/map/page.tsx -> redirect('/') and rewrite
app/calls/page.tsx -> redirect('/incidents') (Archive/search is blocked on
backend work, chunk 12).
ChromeSwitcher now gives a signed-in user at "/" the app shell instead of
marketing chrome; app/page.tsx branches the same way, sending a signed-in
provisioned user to /incidents as an honest interim until the Live screen
itself lands (chunk 6) — marketing content and behavior for signed-out
visitors is unchanged.
Left the light-mode !important overrides in globals.css in place past this
chunk (deviating from the chunk 4 acceptance criteria) — they still back
every page outside this redesign's 11-chunk scope (settings, admin,
profile, marketing). Deleting them now would break light mode on all of
those. Logged in DEFERRED.md.
Per UI_REDESIGN.md chunk 4.
|
||
|
|
8fdedee25b |
Frontend redesign chunk 3: type layer honesty and the duplicate fix
Declare the fields the backend already writes and the UI was discarding: CallRecord gains location_coords, units, vehicles, cleared_units, duplicate_of, srcaddr (intelligence.py ~315-327); IncidentRecord gains units_active, units_cleared, location_mentions, last_thin_at (incident_correlator.py _attach, ~1270-1300). Filter duplicate_of client-side in useCalls.ts's three hooks (useCalls, useCallsByIncident, useActiveCalls) so a call flagged as a second node's recording of the same transmission no longer renders twice. Client-side rather than a where() clause to avoid a new composite index. Removes the two now-resolved DEFERRED.md entries (dedup/useCalls, lib/types.ts field gaps). Per UI_REDESIGN.md chunk 3. |
||
|
|
70d63abeaa |
Re-evaluate incident severity on link, stamp resolved_at at every resolution site
#17: severity was written once at _create_incident and never touched again, so an incident that opened routine and escalated to a working fire stayed routine forever. _update_incident now merges call_severity into the incident via _max_severity() on every link. Severity is monotonic: it only ever rises, never falls. An incident briefly assessed "major" genuinely was major at that moment; a later, calmer-sounding call is evidence the situation is winding down, not that the earlier read was wrong. status/resolved_at exist to retire an incident — severity should stay as the high-water mark so the worst-first rail, "Major only" filter, and map colouring never bury a call that was genuinely major. See _max_severity's docstring in incident_correlator.py for the full argument. #18: none of the resolution sites wrote resolved_at, so an incident's lifespan couldn't be reconstructed for the history-scrub feature. Added resolved_at alongside status="resolved" at all six sites that flip it: - incident_correlator.py _update_incident (signal-based: units all cleared) - incident_correlator.py maybe_resolve_parent (master auto-resolve) - summarizer.py _stale_sweep (90-minute auto-resolve) - upload.py, both scene-resolution loops (single- and multi-scene) - calls.py reprocess/correction path (_update_incident's signal-resolve and maybe_resolve_parent's master-resolve weren't named in the issue's four call sites, but they set status the same way and were missing resolved_at too.) No backfill: existing resolved incidents keep resolved_at = null, which means "resolved before this field existed," not "never resolved." Backfilling from updated_at would be a guess dressed up as data. Tests: added to tests/test_correlator_gate.py, which needs no Firestore for the pure _max_severity cases and patches fstore for the _update_incident/ maybe_resolve_parent writes. Covers the escalation case (routine -> major), the no-downgrade case, and resolved_at on both the signal-resolve and master-resolve paths. 52/52 passing in that file; 83 passed / 10 pre-existing failures for drb-c2-core overall (baseline was 69/10 — the +14 is exactly the new tests, no regressions). Fixes #17, #18. |
||
|
|
3a786bc227 |
Frontend redesign chunk 2: primitives and marks
Rewrite components/ui/* (Button, Card, Badge, PageHeader, EmptyState, Skeleton) against the chunk-1 tokens instead of hardcoded gray-9xx classes, and drop the remaining font-mono from label/heading text. Add the three colour-blindness-validated encoding components from UI_REDESIGN.md §2.3: - components/marks/SeverityMark.tsx — glyph (filled triangle / outline triangle / outline circle) + optional spine + optional label, from the four-level severity ladder. Colour is never the only channel. - components/marks/TypeGlyph.tsx — five stroked SVG glyphs (fire, police, ems, collision, other) in currentColor. Incident type is now shape, not hue, since five hues can't clear an all-pairs CVD gate. - components/marks/NodeMark.tsx — diamond at four weights (filled+ring / filled / hollow / hollow-dashed). Green is gone from node state entirely. lib/severity.tsx now renders through SeverityMark; SEVERITY_COLORS reads the validated sev-moderate/sev-major tokens with routine/minor neutral. IncidentBadges.tsx's TypeBadge is reimplemented on TypeGlyph instead of a coloured pill. Per UI_REDESIGN.md chunk 2. |
||
|
|
c6bc712b54 |
Frontend redesign chunk 1: design tokens and type
Replace hardcoded dark-palette Tailwind classes with semantic CSS custom properties (page/surface/raised/line/ink/accent/sev-moderate/sev-major/ map-*) defined on :root (light) and .dark (dark), wired through tailwind.config.ts theme.extend.colors. Add IBM Plex Sans/Mono via next/font/google: sans for everything a person reads, mono reserved for machine identifiers only. Drop font-mono from body and Button's base classes. Existing !important light-mode overrides kept temporarily so nothing goes unreadable mid-migration (removed in chunk 4). Per UI_REDESIGN.md chunk 1. |
||
|
|
d041c8648d |
Run every hook before the admin guard on /nodes and /systems
Both pages crashed to a blank "client-side exception" screen in production. React error #310: the useState calls sat *below* `if (authLoading || (!isAdmin && !isOperator)) return null`, so the first render returned before reaching them and the next render, once auth resolved, ran more hooks than the previous one. React tracks hooks by call order and refuses. The guard itself is fine and stays where it is -- only the hook declarations move above it. Behaviour is unchanged for a user who passes the guard, and a user who fails it still renders nothing before the effect redirects them. Found by walking the deployed site: /nodes and /systems were the only two routes that failed outright rather than merely showing empty data. The empty data everywhere else is the org_id backfill, which is a separate problem. npx tsc --noEmit clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bc191fb59f |
Stop one malformed call document 500ing the whole debug view
/admin/debug/correlation built its call lookup as {doc["call_id"]: doc}, which
raises KeyError on any stored call missing that field -- and at least one in
production is missing it. One bad document took down the entire view rather
than dropping a single call from it.
The document id is authoritative and always present; the call_id *field* is
written by the upload path and evidently has not always been. Keying off the id
we asked for removes the dependency on the field entirely.
Found while generating a correlation dump server-side, because the UI route this
serves has been unusable tonight.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
157be0c049 |
Serve Firebase's auth handler from our own domain
Google sign-in fails in production: the popup opens, flashes, closes, and the page shows a generic failure with nothing in the console or the network tab. The app is served from drb.cusano.net while signInWithPopup opens its handler on the project's firebaseapp.com origin. Chrome partitions third-party storage, so the popup cannot read back the state its opener wrote and dies immediately. Visiting the handler directly says so: "missing initial state ... a storage-partitioned browser environment". Nothing about authorised domains or the build was wrong -- the shipped bundle carries the correct apiKey and authDomain, which is exactly what made this look like a code bug. Caddy now proxies /__/auth/* on the bare domain to the Firebase Hosting origin, rewriting Host so Firebase recognises the request. Same-site again, which is Google's documented fix. The vhost becomes a `route` so the handler matches before the catch-all proxy to Next. The upstream host is a jinja default rather than a group_vars entry because group_vars/all.yml is gitignored; override it there if the project ever moves. Two manual steps remain, and all three parts are required or nothing changes: the CI secret FIREBASE_AUTH_DOMAIN must become drb.cusano.net with a frontend rebuild, and drb.cusano.net must be an authorised domain in the Firebase console. This template also needs an ansible run -- CI alone will not deploy it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4dc3f27ac4 |
Fix the no-org redirect loop and swallowed Google sign-in errors
Redirect chain traced across middleware.ts, ChromeSwitcher.tsx and
AuthProvider.tsx before touching anything, per the ask. Those three were
already correct as of c7f985d/2a1d52b/83416fe (middleware exempts
/onboarding and /signup from the drb_session cookie gate, ChromeSwitcher
sends any signed-in no-org user to /onboarding, AuthProvider only sets the
cookie once an org_id claim exists). The actual loop was one file upstream
of all three: app/login/page.tsx hardcoded `router.push("/dashboard")`
after both the email/password and Google handlers resolved. That push
races AuthProvider's async onAuthStateChanged -> getIdTokenResult ->
cookie decision. For a no-org account the cookie never gets set, so
middleware bounces the very next request back to /login with no
explanation — the ping-pong the coordinator saw live.
Fix: login page no longer navigates from the handlers. It waits on
AuthProvider's own `loading`/`orgId` and redirects once claims are
settled (/dashboard with org_id, /onboarding without). This also fixes a
second case: a user who lands on /login already signed in (e.g. bounced
there by middleware while their Firebase session was still valid) now
gets routed the same way instead of sitting inert on a login form with no
feedback. /onboarding itself (org-name form, single action) was already
adequate as the "explain the state" screen once the loop stopped
recreating it.
Also, live tonight: Google sign-in was failing outright in prod with no
console/network trace. app/login/page.tsx's Google handler did
`catch { setError("Google sign-in failed. Try again.") }` — no binding,
error discarded. Added lib/authErrors.ts: logs the raw error, and maps
Firebase codes to messages that distinguish two categories — the user's
own situation (popup blocked/closed, bad password, network) says "try
again"; deployment misconfiguration (auth/unauthorized-domain,
auth/operation-not-allowed) says so explicitly and does not suggest
retrying, since retrying can't fix a missing authorized-domain entry or a
disabled provider. Applied to both handlers in login/page.tsx and both
in signup/page.tsx (same swallowing pattern, same fix). Per the
coordinator's steer: this is diagnosis only — no popup-to-redirect
fallback, no auth method change. If production is hitting
auth/unauthorized-domain, that's a Firebase Console fix
(drb.cusano.net -> Authorized domains), not a code fix.
Nav.tsx: sign-out was only reachable from /profile. Added a profile
dropdown (desktop) and drawer entries (mobile) with Profile / Refresh
access / Sign out, so sign-out is reachable from anywhere in the app.
"Refresh access" calls AuthProvider.refreshClaims() (already existed,
already used by /onboarding after signup) so a user whose role or org
was just changed server-side can pick it up without a full logout.
Decision on unknown Google accounts (point 4): kept self-serve org
creation via /onboarding rather than a "request access" pending state.
BUSINESS_MODEL.md #2.1 already answers this for the owner: "a limited
free public tier *and* full paid access without contributing... cash is
the primary revenue line from day one." A pending-approval gate would
contradict that — it would make org creation itself the thing being
gated, when the model explicitly does not want contribution (or approval)
to be the only door. Self-serve org provisioning via POST /auth/signup
was already built for this (
|
||
|
|
90a0412066 |
Bound the correlation debug reads so the view stops hanging
/admin/debug/correlation read every incident ever created, sorted them in Python and kept 20, and separately pulled every call in the orphan window with no cap. That worked while the collections were small. They are not small now: Firestore kills an unbounded scan with a 503 and the request never returns, so the debug view simply spins -- which is also what made the org backfill script fail earlier tonight, same cause, different caller. Incidents now come back pre-sorted from Firestore with a limit, and the orphan scan is capped at 3000 documents. Both queries order on the single field they already filter or sort by (updated_at, ended_at), so neither needs a composite index -- worth preserving, since the index file from the tenancy work has not been deployed. Capping introduces a way to be wrong quietly: a truncated window looks exactly like a quiet night. The payload now carries incidents_window_exhausted and orphan_scan_truncated so a short result announces itself instead of being read as a correlation improvement. The AI-system filter still runs in Python, so the incident window is 10x the requested limit rather than the limit itself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c3fe2a3466 |
Page the org backfill instead of streaming whole collections
A bare .stream() over calls/ returned 503 "Query timed out. Please try either limiting the entities scanned", which the Firestore client then re-raised as an AttributeError from its own retry path -- so the real cause was only visible in the chained traceback. The collection has simply outgrown a single scan. Both passes now walk each collection in 500-document pages ordered by document id, which needs no composite index. The counting pass also stops building a list of every document just to count the ones missing org_id. Documents written mid-run may be missed or seen twice; neither matters, since post-tenancy code stamps org_id at write time and the update is idempotent, so a second run cleans up anything the first skipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1a563c995c |
Point the org backfill at the database the app actually uses
The script initialised firebase_admin from a hardcoded gcp-key.json path and then called a bare firestore.client(). Production has neither: the server is a GCE instance using Application Default Credentials, and the app talks to FIRESTORE_DATABASE=c2-server, not "(default)". The credentials half failed loudly. The database half would not have: the script would have scanned an empty (default) database, found nothing to backfill, created the founding org there, and printed a clean success while the real data stayed untenanted and invisible. Both now read the same environment the app reads, and the chosen database is printed before any work so a wrong one is visible in the dry run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
427d2a9f37 |
Say so loudly when the OpenAI account can no longer be billed
Transcription is the top of the pipeline and it fails soft: any exception logs a WARNING, returns None, and upload.py carries on. That is the right behaviour for a network blip and exactly the wrong behaviour for an unpayable account, because with no transcript there is no extraction, no correlation and no incident -- the system keeps accepting calls and quietly stores empty ones, which looks like quiet radio traffic rather than an outage. This is the third instance of the same failure mode today. The Gemini correlator was down first on a retired model ID and then on a depleted balance, and in both cases the only signal was a per-call WARNING that read as noise. The OpenAI balance is low enough that this one is a matter of when. Billing-shaped errors (insufficient_quota, billing, credit, quota exceeded) now log once at ERROR, name what is dead downstream, and link the top-up page. Everything else keeps the existing per-call WARNING. No new environment variables, so CI deploys this without an ansible run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
83416fe169 |
Split platform-admin from org-owner, hide Trips from non-founding orgs
SAAS_PLAN.md B7. "admin" meant two different things before this: platform operator (SAAS_PLAN.md's own framing) and, by accident of how app/settings/layout.tsx was gated, the only role that could ever reach an org's own billing/members/node-ownership settings. A paying customer who is their own org's owner couldn't reach their own Settings page - the gate checked isAdmin, which only platform admins ever have. settings/layout.tsx now admits org_role === "owner" as well as platform admins (isAdmin stays valid too, for support access to any org's settings). Nav.tsx shows the Settings link on the same condition, and moves Admin (the platform-operator screens: feature flags, users, audit, correlation debug) out of the customer-facing link group entirely - it was already gated server-side, this is just the nav no longer implying it's part of the product. Trips - an internal utility feature riding along on this stack, not a tenant-scoped product surface (see [[trips-feature-intentional]]) - drops out of the customer-facing viewer link group and only shows for the founding org (new lib/tenancy.ts mirrors app/internal/tenancy.py's FOUNDING_ORG_ID) or a platform admin, matching the mutation-route gating routers/trips.py already got in the backend tenancy commit. Reads stay open to any signed-in user, same as before - trips' own visibility model (public/private per trip) predates and is unrelated to org tenancy, and restricting it further wasn't asked for. Also closes two DEFERRED.md items now that they have somewhere to write to: app/settings/organization's "Save changes" button now actually calls c2api.getOrg()/updateOrg() (routers/org.py, shipped in the backend tenancy commit) instead of being permanently disabled. app/settings/nodes gained an EnrollmentTokensPanel (mint/list/revoke against the same commit's /org/enrollment-tokens routes) - without this, B2b's whole point (a customer enrolls their own node with their own token instead of an admin-issued key) had no way to actually be used outside a raw API call. Left alone, and written up as new DEFERRED.md entries instead of guessed at: node/system *write* routes (approve, create, delete) stay platform-admin-only rather than being loosened to org owner/operator - a real gap per SAAS_PLAN.md 2.4, but a separate authorization design that the plan's 12-item build order doesn't enumerate. And settings/members + settings/nodes' ownership table both still call GET /admin/users (platform-admin-only) - a pure org owner who reaches the page via this commit's gate will get 403s from it. Today's only real user is also a platform admin, so this is invisible until a second, non-admin org owner exists. Typecheck: clean (tsc --noEmit via the WSL-native ~/drb-frontend copy). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1b4ed0d09c |
Ship /terms, /privacy, and /waitlist as structure, not finished pages
SAAS_PLAN.md B5/B6, narrowed: no Stripe/pricing/tier work of any kind this pass (a mid-build correction from the business side landed while this was in progress - the commercial model, SAAS_PLAN.md section 6.1, is still undecided), so app/pricing and lib/billing.ts's PLANS are untouched here. What's left of B5/B6 without that - real legal pages and a working waitlist - still ships. app/terms/page.tsx and app/privacy/page.tsx are section scaffolding, not legal text. Every section is a TODO(legal) note describing what that section needs to cover, and the page leads with a "Draft - not yet in force" banner. This isn't caution for its own sake: DRB records, stores, and transcribes public-safety radio traffic, and recording/rebroadcast legality varies by state (SAAS_PLAN.md section 6.3) - an agent-generated draft here would be actively wrong to publish, not just unpolished. Both were pre-added to middleware.ts's PUBLIC_PATHS and ChromeSwitcher's MARKETING_PATHS two commits ago; MarketingFooter now links both. app/waitlist/page.tsx is a real, working form against the already-shipped POST /waitlist - email + optional org name/note, no plan or price mentioned anywhere on it, matching the backend route's own scope (rate limited by source IP, not coupled to any tier). Linked from MarketingFooter as "Request access," not from the pricing page - pricing CTAs stay exactly as they were. Typecheck: clean (tsc --noEmit via the WSL-native ~/drb-frontend copy). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4842725a03 |
Escalate a depleted Gemini balance the same way as a dead model ID
Correcting the model IDs got past the 404s and straight into 429 "Your
prepayment credits are depleted" on every call, so the LLM correlation tier is
still down -- same symptom, different cause, and the previous commit would have
logged it as an ordinary per-call WARNING and buried it exactly like the last
one.
An empty balance shares a status code with an ordinary rate limit but is the
opposite kind of problem: a rate limit clears on its own, a dead account never
does. The match is on the billing wording ("credits are depleted",
"prepayment", "billing") rather than on 429, so a burst of rate limiting still
reads as WARNING while an unpayable account escalates to the once-per-model
ERROR that names the fix.
The two escalation paths now share _log_tier_down, which is also where the
once-per-model suppression lives -- this code runs on every call at radio
traffic volume, so an ERROR per call would be its own kind of noise.
38 correlator tests still pass. No new environment variables, so CI deploys
this without an ansible run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
2a1d52b7af |
Add a real signup path instead of the accidental one
SAAS_PLAN.md 2.2: there was no /signup page. The only self-serve path was Google sign-in on /login, which auto-provisions a Firebase account with no role or org claim at all - previously that meant "viewer role, full read access" the moment the AuthProvider cookie logic (previous commit) let it through. That's closed now regardless; this commit is the other side of it - giving people an actual way in. app/signup/page.tsx: email/password (createUserWithEmailAndPassword) or Google, same visual language as /login. It only creates the Firebase account - org naming is deliberately not on this page, so every path that produces an account with no org (this one, and Google-via-/login) converges on the same next screen. app/onboarding/page.tsx: that screen. Shown to any signed-in user with no orgId (ChromeSwitcher's redirect, previous commit), collects an org name, calls the new c2api.signup() -> POST /auth/signup (routers/links.py, already shipped), then refreshClaims() to force-refetch the ID token so orgId picks up immediately and the same redirect effect sends them on to /dashboard - no manual reload needed. lib/c2api.ts also gained getOrg/updateOrg and the enrollment-token mint/list/revoke calls (routers/org.py, already shipped on the backend) and joinWaitlist (routers/waitlist.py) - none consumed yet, wired in ahead of the settings/legal commits that use them so this stays one add per concept rather than scattering client additions across later commits. /login gained a "Don't have an account? Sign up" link to /signup. This is signup plumbing, not marketing copy - pricing/plan copy (app/pricing, lib/billing.ts) is untouched in this pass, that's a separate, still-open decision (SAAS_PLAN.md section 6). Typecheck: clean (tsc --noEmit via the WSL-native ~/drb-frontend copy). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7f985df42 |
Scope every Firestore hook to org_id and stop granting sessions to nobody's org
Frontend half of SAAS_PLAN.md B2/B3. The backend commits so far (org_id
stamping, Firestore rules) don't protect anything by themselves - every
hook in lib/use*.ts reads Firestore directly from the browser
(onSnapshot(collection(db, ...))), which is why B1's rules commit called
this out as the actual read path in the first place. Until these hooks
filter by org_id, the rules just turn "any signed-in user sees everything"
into "any signed-in user sees nothing" the moment they're deployed, because
nothing supplies the org_id the rules now require.
useCalls (all three exports), useIncidents (useIncidents +
useActiveIncidents), useNodes, useSystems, and useAlerts (both exports) now
pull orgId from AuthProvider and add where("org_id","==",orgId) to their
query. If orgId is falsy - not yet resolved, or the account genuinely has
no org - each hook returns empty rather than falling back to an unfiltered
query, which would silently reopen the exact leak this closes for anyone
whose claim hasn't loaded yet. useIncident/useNodes single-doc-by-id reads
and useTrips are intentionally untouched: single-doc reads are already
covered by the rules directly, and trips has no org_id at all (see the
previous commit's trips.py gating - it's staying founding-org-only via B7,
not becoming tenant-scoped).
AuthProvider grew orgId/orgRole state (read from the org_id/org_role custom
claims POST /auth/signup sets) and a refreshClaims() escape hatch for the
signup flow to force a claims refetch after provisioning. The load-bearing
change is in when it sets the drb_session cookie: only when a claim carries
org_id. A signed-in user with no org - the accidental-signup hole
SAAS_PLAN.md 2.2/2.3 flagged, where Google sign-in on /login auto-creates a
Firebase account with no role or org claim at all - now gets no cookie,
which starts them at "no data, by construction" rather than "viewer role,
full read access" once combined with the rules deployed earlier.
ChromeSwitcher carries the other half of that guard: a signed-in user with
no orgId, anywhere outside the marketing pages, gets redirected to
/onboarding (added to the frontend in the next commit) instead of letting
every page's data hooks just quietly return empty forever. middleware.ts
adds /signup and /onboarding to a new no-cookie-gate list, since
AuthProvider's cookie logic means an unprovisioned user by definition has
no drb_session cookie - gating those two routes on it would bounce exactly
the users who need them back to /login before the client-side redirect
above ever runs. /terms and /privacy (next-next commit) are pre-added to
both PUBLIC_PATHS and ChromeSwitcher's MARKETING_PATHS here so that commit
doesn't need to touch routing files.
Typecheck: clean (tsc --noEmit via the WSL-native ~/drb-frontend copy).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a3681ea698 |
Stamp org_id everywhere and gate every route that leaked across tenants
The previous commit shipped Firestore rules that reference an org_id claim
nothing issues yet, and an org_id filter nothing writes yet - this is the
commit that makes both real. Backend half of SAAS_PLAN.md B2/B2b/B2c.
Data model: organizations/{org_id} and org_members/{uid} are new
collections (models.py OrganizationRecord/OrgMember). org_id is now an
Optional field on NodeRecord, SystemRecord, CallRecord, IncidentRecord,
AlertRule, and AlertEvent - optional because every existing document
predates it; scripts/backfill_org_id.py (written, not run - it touches
production Firestore and Firebase Auth claims) is what closes that gap
later. plan_id/subscription_status/stripe_* on OrganizationRecord are
deliberately None: no billing or pricing model has been decided, so this is
a seam, not a promise. app/internal/tenancy.py holds FOUNDING_ORG_ID, the
org every pre-tenancy document and every legacy enrollment path resolves
into.
Where org_id comes from, end to end: a customer's node enrolls with a
per-org token (new enrollment_tokens/{token_hash} collection, minted via
POST /org/enrollment-tokens - new routers/org.py) instead of the old
fleet-wide ENROLLMENT_TOKEN, which still works as a fallback that resolves
to FOUNDING_ORG_ID so an already-deployed node's .env doesn't start failing
today. The node's org_id then flows onto every call it produces
(mqtt_handler.py's call_start/call_end, upload.py's /upload handler all
resolve it from the node doc), and onto every incident correlated from
those calls (incident_correlator.py's _create_incident/_create_master_incident).
That last one is the part that isn't just a read filter: _build_context's
`all_active = collection_list("incidents", status="active")` fed every
correlation candidate - fast-path talkgroup match, unit-continuity,
disambiguation - from the entire incidents collection, unscoped. Without
scoping it to the call's own org_id, a call from org A could link into an
incident org B already owns, which is a cross-tenant data merge at
correlation time, not just an over-broad read. Same shape of bug in
alerter.py: rule matching pulled every enabled alert_rule regardless of
org, so org A's keyword rule could fire (and POST org A's Discord webhook)
on org B's radio traffic. Both now resolve org_id from the call doc itself
rather than threading a new parameter through every caller.
Every list/get route gained org scoping via a new resolve_caller_org_id()
helper in internal/auth.py, which handles the three credential shapes those
routes accept (service key, node api_key, Firebase user) uniformly and
returns None (unrestricted) for the service key and platform admins -
preserving today's single-org behaviour exactly while closing the leak for
everyone else: GET /nodes, /systems, /calls, /incidents, /alerts,
/alert-rules. Write routes for nodes/systems (approve, create, delete, etc.)
deliberately stay platform-admin-only for now rather than being loosened to
org-owner/operator - that's a real gap called out in SAAS_PLAN.md 2.4's
"should be" column, but it's a separate authorization redesign the 12-item
build order doesn't actually enumerate, and doing it half-considered here
risked being exactly the "half-applied filter is worse than none" failure
mode the plan warns about. Today's founding org keeps working unchanged;
loosening node/system management to org owners is follow-up work, flagged
rather than guessed at.
Also closed the four spend/access-attack routes SAAS_PLAN.md B2c called out
by file and line: POST /calls/{id}/reprocess is now admin-only (was any
signed-in viewer looping the Whisper+Gemini pipeline for free - DEFERRED.md
had this as a live, independent-of-SaaS exploit) plus a per-call rate
limiter as a second guard; POST /alerts/{id}/acknowledge now checks the
alert's org_id; GET /admin/features moved from require_firebase_token to
require_admin_token; and trips.py's four unauthenticated mutation routes
(create_trip, update_trip_tags, create_event, update_event) are now
restricted to the founding org (or the bot's service key, or a platform
admin) - trips has no org_id of its own and isn't getting one, since
[[trips-feature-intentional]] says it's an internal utility riding along on
this stack, not a tenant-scoped product surface.
New public-but-scoped seam: POST /auth/signup (routers/links.py, alongside
the existing /auth/link* routes) provisions an organizations doc and an
owner org_members doc for a just-created Firebase user, then sets their
org_id/org_role claims - idempotent, so a double-submit doesn't create two
orgs. This is the only route that turns "has a Firebase account" into "can
read anything," which is what the frontend AuthProvider no-claim guard
(next commit) is built around.
Also new: GET/PATCH /org for the organization profile (closes the disabled
"Save changes" button noted in DEFERRED.md - there was no organizations
concept to save into before this), and POST /waitlist (public, source-IP
rate-limited, not coupled to any plan or tier - the commercial model is
still an open decision per SAAS_PLAN.md section 6).
Verified: all touched files py_compile clean; c2-core pytest is 69
passed / 10 failed, matching the documented pre-existing baseline exactly
(DEFERRED.md - mqtt_handler/node_sweeper test-vs-code drift, unrelated to
this change) - no new failures. flake8 --max-line-length=120 shows no new
violations in any touched file (checked each new E501/E221/E30x against
`git diff` to confirm it predates this commit); c2-core has no CI lint gate
regardless (CLAUDE.md - flake8 only runs in Client CI).
No new environment variables. Firestore composite indexes for the queries
this introduces were already shipped in the previous commit
(infra/firestore/firestore.indexes.json).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
74faa55396 |
Point correlation at models that still exist, and make a dead one loud
Both Gemini model IDs had been retired by Google. Production logs show every correlation call 404ing -- "models/gemini-2.0-flash is no longer available" -- and gemini-1.5-pro is gone from the model list as well. Because a failed LLM call falls back to the rules decision by design, nothing surfaced: the pipeline kept producing incidents, so the LLM tier and the consensus tiebreak were dead in production for an unknown number of days while correlation was being tuned. Some of what recent tuning was reacting to was rules-only behaviour that was never meant to run alone. Cheap model becomes gemini-3.6-flash, which is the migration target named in Google's own 404. Smart model becomes gemini-2.5-pro, the only stable Pro-tier text model left; the tiebreak fires rarely and its value comes from being a different, stronger model than the first pass, so a second Flash was not worth the consensus it would give up. Model list checked against https://ai.google.dev/gemini-api/docs/models on 2026-08-18. The more important half is the logging. A per-call WARNING was the only signal, and it is indistinguishable from an ordinary API hiccup, so a permanent misconfiguration read as noise. Failures that look like a missing model (404, "not found", "no longer available") now log once per model at ERROR, name the config keys to change, and say plainly that correlation is running rules-only. Transient errors keep the old per-call WARNING. Once per model, not once per call, so the alert stays readable at radio traffic volume. Gemini is used nowhere else in c2-core -- extraction, embeddings and summaries all run on OpenAI -- so the blast radius was exactly the correlation LLM tier. 38 correlator tests still pass. No new environment variables: both model IDs are config.py defaults and are not templated into any .env, so CI deploys this without an ansible run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c09cb72f66 |
Compare unit IDs by normalised key, not exact string
Dispatch audio names the same unit several ways within one conversation, and
every comparison in the correlator used exact string equality, so a follow-up
transmission from a unit already on an incident simply failed to find it. With
the creation gate no longer letting routine traffic open its own incident,
these stopped becoming junk incidents and started becoming orphans instead --
which is how they became visible. In the 01:05Z dump, five of eighteen orphans
were calls belonging to an incident that was open at that moment:
"K-9A2" vs "K-9-A-2" punctuation
"5-1-6" vs "516" digits read out individually
"37" vs "37th Post" ordinal plus role word
"11-Victor" vs "11 Victor" hyphen vs space
_normalize_unit lowercases, drops punctuation and role words (post/unit/car),
strips ordinal suffixes, and joins the remaining tokens, so each pair above
collapses to one key. All six comparison sites now go through it: the two
fast-path debug reporters, unit-continuity candidate selection and its
reassignment check, the cross-talkgroup 2+ shared-unit test, and the
disambiguation scorer.
What it deliberately does NOT do is match a bare district letter -- "Adam" is
not treated as "6-Adam". Every district has an Adam, and collapsing them would
merge unrelated incidents across districts. That leaves a couple of the
observed orphans unlinked, which is the right trade: a missed link leaves an
orphan the re-correlation sweep retries three times, while a false link
corrupts an incident permanently and nothing walks it back.
Two smaller things fall out of the shared helper. Matches are reported as the
original spoken strings rather than the normalised keys, so corr_matched_units
stays readable in the debug view. And a unit made only of role words ("Post")
would normalise to the empty string and then compare equal to every other such
unit, so it falls back to the raw text -- tested, because that failure would be
silent and would merge aggressively.
Adds 13 cases: each observed pair, five pairs that must stay distinct, the
empty-key guard, match reporting, and an end-to-end check that the K-9A2 call
now links where it previously orphaned. 38 pass.
No new environment variables, so CI deploys this without an ansible run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
94ce9d48e2 |
Commit the Firestore security rules that were never in source control
SAAS_PLAN.md's review found the actual finding underneath "no multi-tenancy": drb-frontend reads Firestore directly from the browser (every hook in lib/ does onSnapshot(collection(db, ...))), so drb-c2-core/app/internal/auth.py is never in that read path at all. Whatever rules were protecting calls, incidents, and nodes had been hand-set in the Firebase console - unversioned, unreviewed, and invisible to anyone reading this repo. Added infra/firestore/firestore.rules: deny-by-default, with every tenant-scoped collection (nodes, systems, calls, incidents, alert_events, alert_rules) gated on resource.data.org_id == request.auth.token.org_id, an org_id claim that doesn't exist yet - the next commits add it. All client writes stay denied; c2-core's admin SDK bypasses rules and remains the sole writer, which was already the architecture. Secret-bearing collections (node_keys, the new enrollment_tokens) are denied to clients outright rather than org-scoped, since nothing should ever hand a raw credential to the browser. trips/trip_events keep their current "signed-in users can read" shape rather than being pulled into org scoping - that feature isn't tenant-scoped in this pass (see B7), just hidden from non-founding-org users in the UI. Added infra/firestore/firestore.indexes.json for the composite indexes the org_id-scoped queries will need once the frontend hooks add the equality filter alongside their existing orderBy/range/array-contains clauses - without these, those queries fail at runtime with a FAILED_PRECONDITION "index required" error rather than at review time. Also extended internal/firestore.py's collection_where() with optional order_by/limit_to/start_after params (SAAS_PLAN.md item 1, a stated prerequisite for B2: scoped queries need to stay ordered and bounded, and the existing helper could only do unordered full-collection scans). array_contains needed no new code - it was already a pass-through op string to FieldFilter. None of this is live yet. Deploying rules/indexes is a manual step (firebase deploy --only firestore:rules,firestore:indexes --project <project-id>, from infra/firestore/) - nothing in CI does this. Until it runs, the console-configured rules are still what's actually enforced, and these rules reference an org_id claim no token carries yet. Deploy this alongside (not before) the org_id-stamping commits that follow, or every read breaks for the current single-org deployment. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7b5258cfdf |
Halve the tier-2 thin-call window, from 10 minutes to 5
With over-creation fixed, the incidents that remain are readable enough to judge, and the ones that still do not make sense all fail the same way. A content-free call attaches to the single active incident on its talkgroup if that incident has been idle under tg_dispatch_thin_idle_minutes, and at 10 minutes that is long enough for the channel to have moved on to something else. In the 00:30Z dump a "72 at Holland Station" incident absorbed a Grand Central train-crew meet 9.6 minutes later, and a status check absorbed a records lookup at 9.7. Being the only candidate is not evidence. It means the channel was quiet, which is exactly when guessing is weakest -- the single-candidate rule was meant to avoid picking wrongly among several, not to license a match no other signal supports. Every correct thin attach in that dump was <= 3.4 minutes idle and every wrong one was >= 8.2, so 5 separates them with room on both sides. Real back-and-forth is unaffected: it runs through the 30-second tier-1 path, and the observed conversational replies sit near zero. Tests pin both sides of the new boundary at 4.9 and 5.1 minutes so a later change to this number has to be deliberate. 23 pass. Also corrects a DEFERRED.md entry written earlier today. It claimed nothing ever closes an incident that goes quiet; summarizer.py has run a stale sweep at incident_auto_resolve_minutes (90) the whole time. The 37 open incidents were caused by over-creation, not by a missing sweeper, and 90 minutes may be fine now -- worth rechecking on a fully post-fix dump before changing it. No new environment variables: tg_dispatch_thin_idle_minutes is a config.py default and is not templated into any .env, so CI deploys this without an ansible run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0bd92269d2 |
Stop httpx logging API keys in plaintext
httpx logs every request at INFO as a full URL including the query string, so the Google Maps key appeared in c2-core's container logs on every geocode call -- `?address=Holland+Station&...&key=AIza...`. Anyone who can read the logs, or who is pasted a few lines of them, has the key. It was found exactly that way while checking why the map was empty. Nothing in this service needs per-request client logging; callers already log their own failures with context. httpx and httpcore drop to WARNING, so real transport errors still surface and the URLs stop being printed. This does not un-leak the existing key -- it is in the container's log history and has to be rotated in GCP separately. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
96625fabd0 |
Stop ambient radio chatter from opening incidents, and refill the map
The 23:46Z correlation dump confirmed the severity gate fixed the problem it
was written for -- orphans fell from 69 to 16, and only three of those are
after the deploy boundary, two of them deliberate skips. Nothing on TG 9048
absorbs the channel any more; the largest post-deploy incident is four calls
over nine minutes and is genuinely one event.
It overcorrected. 37 of 50 incidents were open, most a single routine call.
The cause was the gate's own substance test, which counted `units` and
`location`. Radio protocol puts a unit ID in essentially every transmission
and a place name in most of them, so has_substance was true almost always and
the severity check never actually ran -- "11-Victor, 72 at Holland Station"
became its own permanent incident. Substance is now a vehicle, a geocode or a
tag: things the extractor found beyond who was speaking and where they stood.
Severity still opens an incident on its own, so nothing real is lost.
incident_type is now validated against the enum the prompt offers rather than
trusted. It is written straight through to incident.type and rendered as the
title, so a model that answered the severity question in the type field
produced an incident titled "Routine -- TGID 9563". Unrecognised values become
None and fall to the tag/severity path, which is what "unknown" already did.
The map was empty for a separate reason: geocoding accepted only ROOFTOP and
RANGE_INTERPOLATED. Dispatch names places the way people speak, and Google
returns GEOMETRIC_CENTER for exactly those forms -- intersections ("Lake
Street and Veterans Memorial Drive") and named POIs ("Brewster Station").
Requiring a street address discarded nearly every real dispatch location and
left only numbered addresses plotted, which is why the July incidents have
coordinates and none since do. GEOMETRIC_CENTER is now accepted; APPROXIMATE
is still rejected, since a region centroid is what an ungeocodable string
degrades to. Note this is necessary but may not be sufficient -- if
GOOGLE_MAPS_API_KEY is unset on the host the map stays empty regardless, and
that has not been checked from here.
Two things found and deliberately not fixed, both in DEFERRED.md. One call can
still land in two incidents, because upload.py correlates each extracted scene
independently and the model over-split one conversation; multi-scene is
intentional, so that is prompt tuning rather than a code change. And nothing
closes an incident that merely goes quiet -- signal-resolution and master
auto-resolve both exist, but a one-call incident nobody clears stays active
forever. That wanted the over-creation fixed first so a time-based sweeper
would not just paper over it.
Gate tests updated: units and location alone must now orphan, and the case
that matters most is kept explicit -- units with a real severity still open an
incident. 17 pass. No new environment variables, so CI deploys this without an
ansible run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
53965e1a19 |
Rebuild the frontend as a product rather than an internal tool
The UI worked but read as an operator console: no public face, no way to describe or sell the thing, and no account surface beyond the node list. This adds the missing halves and reorganises what was already there around the incident, which is the unit of value the rest of the pipeline is built to produce. A shared design system replaces per-page styling: components/ui (Button, Card, Badge, EmptyState, Skeleton, PageHeader), a type scale and shadow set in the Tailwind config, and light-mode tokens in globals.css. The existing html:not(.dark) remap mechanism is extended rather than replaced -- a parallel theming system would have been two sources of truth for the same colours. Public marketing pages (/, /features, /pricing, /faq) load without a session. middleware.ts gained a PUBLIC_PATHS allowlist to permit that; it remains a UX redirect and is still NOT an authorisation boundary, which the comment there says explicitly. Real enforcement is unchanged and still lives server-side in c2-core's auth.py. Chrome switching is done by pathname in ChromeSwitcher instead of by route group, because a route group would have collided on / and forced most of app/ to move for no behavioural gain. Billing and API keys ship as typed stubs, not integrations. lib/billing.ts and lib/apiKeys.ts define the data model and the screens consume it, but every mutating call throws with a message naming the backend route that has to exist first, and the sample data is labelled as sample. Nothing here can charge anyone or mint a real credential -- picking a payment processor and holding its keys is a decision for a human, and a half-wired checkout is worse than an obviously absent one. The severity work from the c2-core change lands here too. severity is now a filter and sort dimension on the incident list rather than decoration, since a busy dispatch channel is only readable if you can collapse it to moderate and above. routine gets a muted treatment because it is the majority of traffic, legacy "unknown" still renders nothing, and TypeBadge handles the new "other" incident type. Severity rendering moved into lib/severity.tsx so the incident list, incident detail and call rows cannot drift apart. Deliberately not touched: calls, map, alerts, nodes, systems, tokens, trips and admin. They already share the palette and stay coherent, and rewriting them would have buried the parts that actually needed to change. No colour tokens were renamed, so nothing regressed there. Verified with tsc --noEmit (npm run typecheck), clean. No runtime verification was possible and none was done. No new environment variables. |
||
|
|
6d5eb4c5f2 |
Let severity, not incident_type, decide what becomes an incident
The 2026-08-16 correlation dump showed two failures that looked unrelated and were the same bug. TG 9048 held one incident of 28 calls spanning 49 minutes -- a prisoner transport, a drone retrieval, a records lookup and a canvass, glued together -- while 32 other calls on that same channel stayed permanently orphaned. Creating an incident required a concrete incident_type. Nothing on a transit police channel produced one: the extraction prompt said to prefer "other" when uncertain, extraction then collapsed "other" to None, and the tag-based fallback had no tags to work with because administrative traffic carries none. So the channel could never open a SECOND incident. Every later call funnelled into whichever incident happened to exist first, and every call too substantial for the thin path had nowhere to go at all. The two symptoms were the same missing value seen from opposite ends. Severity now decides incident-worthiness. It is a better fit for the question being asked -- "is this a real event?" -- than a service label ever was, and unlike incident_type it is always present. The prompt defines four levels with no escape hatch (routine/minor/moderate/major, "unknown" is gone) and calls skipped for a too-short transcript are still recorded as routine, because downstream code reads a missing severity as "not processed yet" rather than "nothing happened". Anything above routine, or carrying any extracted content, opens an incident under the neutral "other" type. "other" is also kept as a real classification now -- rail operations and public works genuinely are not police, fire or EMS. Separately, thin calls no longer refresh updated_at; they write last_thin_at. updated_at drives every recency gate in the fast path, so each "10-4" was resetting the idle clock on whatever it attached to, keeping that incident inside the gate for as long as anyone kept acknowledging. An incident now ages from its last substantive call. This is what made the 49-minute incident possible even once buckets existed, so it is fixed independently rather than being left to the gate change. The re-correlation sweep also now honours skip_reason. /upload has always refused to correlate garbage and too-short transcripts, but the sweep did not apply the same filter, so those fragments came back minutes later through the thin path and attached to whatever was most recent -- a second, quieter route into the same over-merge. Adds tests/test_correlator_gate.py (15 cases), the first tests against incident_correlator.py in its 1,517-line history. tests/conftest.py stubs firebase-admin only when it is genuinely absent, so the container's real SDK is never shadowed; this is what makes the correlator importable in the dev venv. That stub also made test_mqtt_handler and test_node_sweeper collectable for the first time, revealing 10 pre-existing failures in them -- test-vs-code drift, untouched here and catalogued in DEFERRED.md. No new environment variables, so CI deploys this without an ansible run. |
||
|
|
97013e1505 |
Stop Whisper hallucinations and dedupe recordings across nodes
Two independent sources of garbage in the AI pipeline, both visible in the 2026-08-16 correlation dump. 1. Hallucinated transcripts. The Whisper prompt opened with an enumerated run of ten-codes: 10-4, 10-23, 10-20, 10-97 and so on. Whisper treats prompt text as preceding transcript, so on noisy or silent audio it continued the series, emitting transcripts that count upward from 10-4 to 10-99. The existing no_speech_prob filter could not catch these: the model is highly confident in text it invented by continuing a pattern. The prompt no longer contains a series to extend, and _is_degenerate() rejects the three shapes this failure takes: ascending ten-code runs, one phrase looping, and near-identical segments across a whole recording. Verified against 13 transcripts from production: all four known hallucinations rejected, all nine real ones kept, including terse traffic containing legitimate codes. 2. Duplicate recordings. node-002 and node-PI-2 both cover TG 9048 and both uploaded the same transmissions, ~1.1s apart. Nine pairs appeared in one dump. Each was transcribed, billed and correlated twice, and the resulting incident listed two units where there was one. Canonical selection is by earliest started_at, tie-broken on call_id, NOT by upload order: upload order varies with encode time and network latency, so it would make the authoritative recording non-deterministic. Call documents are created from MQTT call_start before uploads arrive, so both nodes independently reach the same verdict. The loser keeps its audio (it may be the cleaner capture) but is excluded from STT, correlation, the re-correlation sweep and the orphan debug view. Also fixes _sync_transcribe returning a bare None when OPENAI_API_KEY is missing, where the caller unpacks two values. A missing key surfaced as a misleading "Transcription failed" instead of the real warning. Adds tests/test_dedup.py (15 cases). dedup.py reaches Firestore through an injected callable so it stays importable without firebase-admin present. |
||
|
|
a2cd2c57ca |
Serve call audio through c2-core instead of GCS signed URLs
upload_audio() could only sign a URL when GCP_CREDENTIALS_PATH pointed at a service-account key file. The deployed VM runs on Application Default Credentials with no key file, so every upload silently took the fallback branch and returned a bare gs:// URI. That broke two things at once: * Browsers cannot fetch a gs:// URI, so no recording was ever playable. * _public_url_to_gcs_uri() only matched https://storage.googleapis.com/ and returned None for it, so `if gcs_uri:` in the upload path was always false and transcription never ran. Nothing was logged, which is why this looked like an OpenAI credits problem rather than a storage one. The fallback also interpolated the client-supplied filename instead of the call_id-derived safe name, so the URI did not even name the object written. Calls now store only the canonical gs:// location. A short-lived playback link is minted per read as an HMAC over (call_id, expiry) keyed by SERVICE_KEY, and audio is served from the private bucket by the new /media route. An <audio src> cannot carry an Authorization header, so the link has to be the credential; that router is therefore public with the check done inline, as enrollment.py already does. Signing GCS URLs from the VM would have needed a serviceAccountTokenCreator grant on its own service account — this avoids the IAM change entirely and keeps the bucket private. gcs_uri_for_call() reconstructs the object name from call_id, so recordings made before this fix are reachable again without a data migration. Frontend rows come straight from Firestore via onSnapshot and never see a server-minted field, so CallRow fetches the link lazily on expand. Also removes the last long-lived (1 year) signed URL and the log line that printed it. |
||
|
|
a195563da6 |
Let edge nodes read /systems with their own api_key
The node builds its OP25 config from GET /systems, but that router only accepted a Firebase token or the shared service key — a node holds neither. Every fetch returned 401 and the node fell back to its stale offline cache, so a system edited in the UI never reached the field. Confirmed on node-002 against the live server: "Failed to fetch systems from C2: 401 Unauthorized ... Offline cache will be used." The node sends no node_id with the request, only the bearer token, so the key is matched by querying node_keys for the value instead of fetching a known document the way /upload does. Read access only: the mutating routes in this router each carry their own require_admin_token, so widening the router-level gate doesn't let a node create, edit or delete a system. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
55cd1110df |
Give mosquitto's bind-mounted dirs to uid 1883, not root
The broker crash-looped on every deploy: "Unable to load server certificate /mosquitto/certs/mqtt.crt ... Permission denied". The cert-sync script wrote 600 root:root into a 0700 root:root directory, on the assumption that mosquitto runs as root inside its container. It does not — the stock eclipse-mosquitto entrypoint drops privileges to the in-image mosquitto user, confirmed on the server as uid=1883(mosquitto) gid=1883(mosquitto), and the broker's own log says so on every start. Certs dir is now root:1883 0750 with the cert 0644 and the key 0640, and the data dir is 1883:1883 recursively — recursively because mosquitto WRITES dynamic-security.json there, and a root-owned file left by an earlier deploy would still be unwritable after a directory-only chown. Also drops the "unverified Caddy cert path" note: a real issuance confirmed the path, producing CN=mqtt.drb.cusano.net signed by Let's Encrypt. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a0a414ad21 |
Revert the CI full-fetch workaround and drop the dead Caddyfile
Shallow clones were never a Gitea packing bug. An intruder had set uploadpack.packObjectsHook in Gitea's HOME gitconfig, pointing at a non-executable dropper, so every upload-pack died mid-pack. That hook is gone and --depth=1 clones are verified working, so fetch-depth: 0 buys nothing but slower CI. See INCIDENT-2026-08-11.md. infra/Caddyfile was dead: ansible templates Caddyfile.j2 to /etc/caddy/Caddyfile, and nothing ever deployed the static copy. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
518ac46929 |
Use a full fetch in CI: Gitea fails to pack a shallow clone
actions/checkout defaults to depth=1, and Gitea aborted generating that pack with a bad pack header protocol error on all three retries, failing the build before any image was pushed. A full fetch avoids the shallow-pack path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
052dda0b1f |
Point app_url at the bare domain and publish the broker host
app_url advertised https://app.<domain>, which has never had a DNS record — the frontend is served on the bare domain by Caddy. Adds mqtt_host so the broker endpoint nodes connect to is discoverable from terraform output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ee633cbe46 |
Secure the broker for public exposure: TLS and per-node credentials
Edge nodes are deployed to arbitrary locations by arbitrary people, so the
broker has to be reachable from the internet and secured on its own merits
rather than by a VPN.
Three defects made that impossible. The broker only had a plaintext 1883
listener; every node shared one drb-node password; and the ACL pattern used
%c, the client-supplied client id, so any holder of that shared password
could set client_id to another node and take over its namespace. The comment
claiming this cryptographically prevented cross-node access was wrong and is
gone.
Authentication now uses mosquitto 2.x's built-in dynamic-security plugin on
the stock eclipse-mosquitto image. c2-core administers it over the control
topic, creating each node's client on approval with username=<node_id> and
password=<its node_keys api_key>, attached to a role whose ACL is nodes/%u/#
against the authenticated username. One credential, one revocation point.
An HTTP-callback plugin was implemented first and rejected: that project is
archived upstream, which is not an acceptable dependency on an
internet-facing broker.
Because dynsec state is a second source of truth alongside Firestore,
approve/reissue/delete now write to the broker first and surface a 502
rather than drifting, and c2-core reconciles every approved node into dynsec
on startup.
Adds node self-enrollment (POST /nodes/enroll, GET /nodes/{id}/credentials)
so a new node can obtain its key over HTTPS without an operator handling
secrets by hand. Enrolling an already-approved node_id is refused on the
fleet token alone — otherwise a leaked token plus a guessable id would let
an attacker steal a live node's key before the real node asked for it.
Pickup secrets are stored hashed and returned once, and the endpoint is rate
limited per source IP.
Infrastructure: an 8883 TLS listener fed by Caddy's certificate via a
systemd path unit, a firewall rule for it, and Caddy now 404s /internal/*
so the api vhost cannot proxy internal routes.
Also fixes CORS, which allowed https://app.<domain> while the frontend is
served on the bare domain — every call from the portal would have failed —
and widens the vault gitignore to a glob, since ansible-vault leaves
backup siblings that the exact-name rule left committable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
1f5f1fede8 |
Serve the frontend on the bare domain instead of app.<domain>
Only drb.cusano.net and api.drb.cusano.net have public A records, so the app.<domain> vhost had no cert to present and the bare domain — the record that actually exists — matched no site at all, producing ERR_SSL_PROTOCOL_ERROR in the browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
12c9ad73bb |
Document the no-$-in-vault-values rule that caused the MQTT auth failure
A password containing "$fP" was interpolated away by compose, giving mosquitto and c2-core two different passwords and producing "MQTT connect refused: Not authorized" with nothing in the logs pointing at the cause. Recorded next to the values so the next person generating credentials sees it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
971ab74d44 |
Escape $ in the compose-interpolated .env so MQTT passwords survive
Compose interpolates the top-level .env, so a password containing "$fP" was read as the variable $fP and replaced with an empty string — hence the repeated "The \"fP\" variable is not set" warnings on every compose command. The env_file templates are not interpolated, so c2-core kept the literal password while mosquitto's entrypoint received the mangled one. The two sides disagreed and c2-core could not authenticate to the broker. Escaping $ as $$ here (and only here) makes compose collapse it back to the real value. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6140dd7b9c |
Fix prod compose port collision and make ansible deploy re-runnable
docker-compose.prod.yml: compose merges `ports` by appending, so the prod override left the base file's 8888:8000 and 3000:3000 in place next to the 127.0.0.1-scoped ones. Each container tried to bind its port twice and the second bind failed with "address already in use", so c2-core and frontend could never start. It also meant the localhost-only binding never applied — both ports were published on every interface. Marked both `!override`, the same way mosquitto already used `!reset`. infra/ansible: - add the missing "Reload Caddy" handler; the Deploy Caddyfile task notified a handler that did not exist, which aborts the play - guard mkswap/swapon on whether /swapfile is already active, so a second run does not fail on "mounted" / "Device or resource busy" - git task now updates instead of clone-once, otherwise a re-run redeploys whatever code was on the VM at first clone - vault.yml.example: correct the registry token comment to read-only scope Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2e3fde2448 | refactor: Clean checkin override parsing and require node type in frontend configuration modal | ||
|
|
c42bd1902c | feat: Add local system override with 24h timeout support | ||
|
|
c6684ea61b | Update deploy with next vars | ||
|
|
57ff9f8ea3 | Merge remote-tracking branch 'origin/main' into build-infrastructure | ||
|
|
9fdcad1c46 |
deploy via Gitea CI registry; provision GCP infra with Terraform
- Terraform: e2-micro VM (us-east1-b, free tier), static IP, SSH/web
firewall rules, IAM bindings for Firestore + GCS; imports existing
drb-calls bucket and c2-server Firestore database into state
- Gitea CI: build c2-core, discord-bot, frontend images and push to
git.vpn.cusano.net registry; SSH deploy pulls pre-built images (no
build on VM)
- Ansible: first-time setup only — git clone, env files from vault,
Caddyfile, docker login + compose pull + up; no rsync or on-VM builds
- docker-compose: add image: ${REGISTRY}/name:latest alongside build:
so local dev and CI registry both work
- gitignore: add Terraform state, lock, tfvars, ansible secrets
|
||
|
|
33700448bf |
add Terraform + Ansible infrastructure for GCP deployment
Provisions e2-micro VM (us-east1-b, free tier) with static IP, SSH and web firewall rules, Docker + Caddy startup script, and IAM bindings for Firestore and GCS access via ADC. Imports existing drb-calls bucket and c2-server Firestore database into state. Ansible roles handle first-time setup (swap, docker group) and all subsequent deploys via rsync + docker compose, with secrets managed via Ansible Vault. DNS stays on AWS Route 53. |
||
|
|
3defdf18dc | stale calls fix | ||
|
|
1f17b6c0d2 |
feat: add role-based user management, audit log, and session tracking
Introduces a full user management system with three roles (admin, operator,
viewer), an audit log, and per-session login history.
Backend:
- app/internal/audit.py: write_audit() helper → audit_log Firestore collection
- app/internal/auth.py: get_role() helper; require_admin_token accepts both
legacy admin:true claim and new role:"admin" claim for backward compat
- app/routers/users.py: CRUD under /admin/users — list, create (returns
one-time invite link), get (with sessions), patch role/nodes/name,
disable, enable, delete; operator role requires ≥1 owned node
- app/routers/links.py: POST /auth/session records sign-in events to
user_sessions Firestore collection
- app/routers/admin.py: GET /admin/audit paginated endpoint
- app/main.py: register users router
Frontend:
- AuthProvider: exposes role, isAdmin, isOperator, ownedNodeIds from claims
- Nav: role-gated links — viewers get dashboard/calls/incidents/map/alerts/
trips; operators add nodes/systems/tokens; admins add admin
- admin/page.tsx: new Users tab (list table, create modal, inline edit panel
with role/nodes editor, disable/enable/delete, login history) and Audit
Log tab (paginated, color-coded actions)
- login/page.tsx: calls recordSession() on email and Google sign-in
- nodes, systems, tokens pages: role guards redirect viewers to dashboard
- profile/page.tsx: shows accurate role badge and label
- lib/types.ts: UserRole, UserRecord, UserSession, AuditEntry types
- lib/c2api.ts: user management methods + recordSession
Firestore collections added: user_profiles, audit_log, user_sessions
Firebase custom claims schema: { role, owned_node_ids, admin (legacy) }
|
||
|
|
961cc6f36e | add button to clear stale 'active' calls | ||
|
|
d290b89736 |
New /profile page
Avatar (initials) + display name, email, admin badge Account section: email, UID, role, join date, last sign-in Discord section: link status with username/user ID/linked date, or the get-code flow if unlinked, plus unlink button Sign out button at the bottom |
||
|
|
758c6f4115 | discord link banner | ||
|
|
6ae4d398f8 | add trips permissions | ||
|
|
981f03ac06 | allow overlap (note) tags | ||
|
|
47430827d4 | Fix discord trip itinerary | ||
|
|
4dd3343026 | add event editing | ||
|
|
fce189d8c9 | assistant updates | ||
|
|
3fb3bca034 |
add tags
Trip-level tags: admins configure available tags in the trip header (inline add/remove pills). The AI can also create new tags via the add_tag tool. Event tags: selectable in the Add Event modal, shown as colored pills on event cards in the timeline, and on AI suggestion cards. AI integration: sees available tags in its system prompt, applies them when proposing events, can create new ones with add_tag. Discord: tags shown as inline code blocks under each event in /trip view. Colors: auto-assigned from an 8-color palette by tag index, consistent everywhere. |
||
|
|
a0fdf2486e |
chat fixes
Focus: textarea gets refocused via inputRef after the AI response (or error) lands Persistence: chat history saved to localStorage keyed by trip ID, loaded on mount — survives refreshes |
||
|
|
e7622c7e6d | chat box fixes | ||
|
|
21d15d0426 | assistant markdown update | ||
|
|
21268ab477 |
fix: migrate Places and Routes to new GCP APIs
Switch from legacy Places textsearch and Directions APIs (disabled on this project) to Places API (New) and Routes API (New). Both places.py and the assistant's _places_search helper updated. Also fixes uid() recursive self-call in trips page and adds Places API response logging. |
||
|
|
522748f07a | debugging for trips assistant | ||
|
|
af4079d648 | fix build | ||
|
|
39c002d090 | Fix assistant | ||
|
|
4295bdf4d2 | Merge remote-tracking branch 'origin/main' into build-infrastructure | ||
|
|
18d96193ab |
Security fixes
auth.py
secrets.compare_digest replaces == for service key comparison (timing-safe)
Added require_service_key — bot-only endpoints (trip/event join/leave)
Added require_service_key_or_admin — node commands/config (bot via service key OR dashboard admin via Firebase)
Added _RateLimiter with three shared instances: trip_chat_limiter (20/5min per user), summarize_limiter (5/10min per incident), bootstrap_limiter (2/hr per system)
nodes.py
send_command and assign_system now require require_service_key_or_admin — the Discord bot can still call them via service key, but regular Firebase users are blocked
tokens.py
add_token, flush_tokens, set_preferred_system, delete_token all require require_admin_token
Token masking changed from token[:10] + "…" + token[-4:] to "•••" + token[-4:]
systems.py
All write endpoints (create, update, delete, ai-flags, ten-codes, vocabulary writes, bootstrap) now require require_admin_token
bootstrap_vocabulary also calls bootstrap_limiter.check(system_id)
incidents.py
POST /incidents/summarize (bulk) now requires require_admin_token
POST /incidents/{id}/summarize now calls summarize_limiter.check(incident_id)
trips.py
join_trip, leave_trip, join_event, leave_event require require_service_key — only the Discord bot can set Discord attendee identity
delete_trip, delete_event require require_service_key_or_admin
trip_chat rate-limited per caller UID, history stripped to user/assistant roles only, user message truncated to 2000 chars, Maps query strings capped at 200 chars
upload.py
Rejects files larger than settings.upload_max_bytes (default 100MB) with 413
storage.py
_safe_audio_filename() derives GCS object name from call_id + allowlisted extension, completely ignoring the client-supplied filename
config.py
Added upload_max_bytes: int = 100 * 1024 * 1024
Both Dockerfiles — python:3.14-slim → python:3.12-slim
|
||
|
|
a1c91c5ed3 | Initial infra attempt | ||
|
|
f0a0ea508a | adjust assistant height | ||
|
|
d64259bb18 | Fix auth | ||
|
|
7b9aefbcc5 | Add UI to trips | ||
|
|
8edb717dd2 |
Add trips to UI
lib/types.ts — TripRecord and TripEvent types lib/c2api.ts — getTrips, getTrip, createTrip, deleteTrip, createTripEvent, deleteTripEvent lib/useTrips.ts — Firestore realtime hook on the trips collection, ordered by start date app/trips/page.tsx — List page split into Upcoming / Past sections, card click navigates to detail, "+ New Trip" modal for admins with all fields including date range and maps link app/trips/[id]/page.tsx — Detail page fetched via C2 API (gets trip + events in one call), day-by-day itinerary with time, location, maps link, notes, and Discord attendees. Add Event modal (date constrained to trip range). Admin-only delete trip + remove event. components/Nav.tsx — Trips link added to the nav |
||
|
|
fb096d582d |
feat: add /trip slash commands + add trips & itinerary system
New /trips router with full CRUD, attendee management, and nested events. Events validate date is within parent trip range and inherit trip location when not explicitly set. Leaving a trip cascades removal from all its events. New TripCommands cog with /trip create, list, view, delete, join, leave and /trip event add, remove, join, leave. Event autocomplete is scoped to the selected trip. Enforces must-be-on-trip rule for event joins with a clear error message. |
||
|
|
a4962d7b0e | map fixes | ||
|
|
4e0e0fc79f |
Backend (incident_correlator.py):
- Create path (line ~1274): title only uses "at {location}" when location_coords is also set
- Update path (line ~1226): same guard — best_coords must be truthy alongside best_location
Frontend (MapView.tsx):
- Desktop sidebar: cards with location_coords → <button> fly-to; cards without → <a href> that navigates to the incident page with "View details →" text
- Mobile drawer: same split — with coords fly-to+close, without coords navigate via <a>
- Removed the "no coords" italic placeholder text; the card behavior itself makes it clear
|