Compare commits

...
2 Commits
Author SHA1 Message Date
Logan CusanoandClaude Opus 5 c09cb72f66 Compare unit IDs by normalised key, not exact string
Build & Deploy / Build & push images (push) Successful in 4m6s
Build & Deploy / Deploy to VM (push) Successful in 2m32s
Dispatch audio names the same unit several ways within one conversation, and
every comparison in the correlator used exact string equality, so a follow-up
transmission from a unit already on an incident simply failed to find it. With
the creation gate no longer letting routine traffic open its own incident,
these stopped becoming junk incidents and started becoming orphans instead --
which is how they became visible. In the 01:05Z dump, five of eighteen orphans
were calls belonging to an incident that was open at that moment:

    "K-9A2"     vs "K-9-A-2"     punctuation
    "5-1-6"     vs "516"         digits read out individually
    "37"        vs "37th Post"   ordinal plus role word
    "11-Victor" vs "11 Victor"   hyphen vs space

_normalize_unit lowercases, drops punctuation and role words (post/unit/car),
strips ordinal suffixes, and joins the remaining tokens, so each pair above
collapses to one key. All six comparison sites now go through it: the two
fast-path debug reporters, unit-continuity candidate selection and its
reassignment check, the cross-talkgroup 2+ shared-unit test, and the
disambiguation scorer.

What it deliberately does NOT do is match a bare district letter -- "Adam" is
not treated as "6-Adam". Every district has an Adam, and collapsing them would
merge unrelated incidents across districts. That leaves a couple of the
observed orphans unlinked, which is the right trade: a missed link leaves an
orphan the re-correlation sweep retries three times, while a false link
corrupts an incident permanently and nothing walks it back.

Two smaller things fall out of the shared helper. Matches are reported as the
original spoken strings rather than the normalised keys, so corr_matched_units
stays readable in the debug view. And a unit made only of role words ("Post")
would normalise to the empty string and then compare equal to every other such
unit, so it falls back to the raw text -- tested, because that failure would be
silent and would merge aggressively.

Adds 13 cases: each observed pair, five pairs that must stay distinct, the
empty-key guard, match reporting, and an end-to-end check that the K-9A2 call
now links where it previously orphaned. 38 pass.

No new environment variables, so CI deploys this without an ansible run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:35:06 -04:00
Logan CusanoandClaude Opus 5 94ce9d48e2 Commit the Firestore security rules that were never in source control
SAAS_PLAN.md's review found the actual finding underneath "no multi-tenancy":
drb-frontend reads Firestore directly from the browser (every hook in lib/
does onSnapshot(collection(db, ...))), so drb-c2-core/app/internal/auth.py
is never in that read path at all. Whatever rules were protecting calls,
incidents, and nodes had been hand-set in the Firebase console -
unversioned, unreviewed, and invisible to anyone reading this repo.

Added infra/firestore/firestore.rules: deny-by-default, with every
tenant-scoped collection (nodes, systems, calls, incidents, alert_events,
alert_rules) gated on resource.data.org_id == request.auth.token.org_id, an
org_id claim that doesn't exist yet - the next commits add it. All client
writes stay denied; c2-core's admin SDK bypasses rules and remains the sole
writer, which was already the architecture. Secret-bearing collections
(node_keys, the new enrollment_tokens) are denied to clients outright rather
than org-scoped, since nothing should ever hand a raw credential to the
browser. trips/trip_events keep their current "signed-in users can read"
shape rather than being pulled into org scoping - that feature isn't
tenant-scoped in this pass (see B7), just hidden from non-founding-org users
in the UI.

Added infra/firestore/firestore.indexes.json for the composite indexes the
org_id-scoped queries will need once the frontend hooks add the equality
filter alongside their existing orderBy/range/array-contains clauses -
without these, those queries fail at runtime with a FAILED_PRECONDITION
"index required" error rather than at review time.

Also extended internal/firestore.py's collection_where() with optional
order_by/limit_to/start_after params (SAAS_PLAN.md item 1, a stated
prerequisite for B2: scoped queries need to stay ordered and bounded, and
the existing helper could only do unordered full-collection scans).
array_contains needed no new code - it was already a pass-through op string
to FieldFilter.

None of this is live yet. Deploying rules/indexes is a manual step
(firebase deploy --only firestore:rules,firestore:indexes --project
<project-id>, from infra/firestore/) - nothing in CI does this. Until it
runs, the console-configured rules are still what's actually enforced, and
these rules reference an org_id claim no token carries yet. Deploy this
alongside (not before) the org_id-stamping commits that follow, or every
read breaks for the current single-org deployment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:31:09 -04:00
6 changed files with 374 additions and 16 deletions
+26 -1
View File
@@ -68,16 +68,41 @@ async def collection_list(collection: str, **filters) -> list[dict]:
async def collection_where(
collection: str,
conditions: list[tuple[str, str, Any]],
order_by: Optional[list[tuple[str, str]]] = None,
limit_to: Optional[int] = None,
start_after: Optional[dict] = None,
) -> list[dict]:
"""
Query a collection with arbitrary where-clauses.
conditions: list of (field, op, value) — e.g. [("ended_at", ">=", cutoff_dt)]
Supports any Firestore operator: "==", "!=", "<", "<=", ">", ">=".
Supports any Firestore operator, including "array_contains" — it's just
forwarded straight to FieldFilter, so a condition like
("incident_ids", "array_contains", incident_id) already worked before this
function grew explicit order_by/limit/cursor params below.
order_by: list of (field, direction) — direction is "ASCENDING" or
"DESCENDING" (Firestore's own constants; passed straight through as
strings so this module doesn't need a google.cloud.firestore_v1.Query
import). Applied in list order, so multi-field sorts work.
limit_to: cap the number of documents returned.
start_after: cursor — a dict of the same field values as the *last*
document from a previous page's order_by fields (Firestore's
`Query.start_after()` takes a field-value mapping, not a document
snapshot, when you're not holding one).
Added for org_id-scoped queries that also need to be ordered/paginated —
unscoped equality-only lookups can keep using collection_list().
"""
def _query():
ref = db.collection(collection)
for field, op, value in conditions:
ref = ref.where(filter=FieldFilter(field, op, value))
for field, direction in (order_by or []):
ref = ref.order_by(field, direction=direction)
if start_after is not None:
ref = ref.start_after(start_after)
if limit_to is not None:
ref = ref.limit(limit_to)
return [doc.to_dict() for doc in ref.stream()]
return await asyncio.to_thread(_query)
+68 -14
View File
@@ -135,6 +135,62 @@ _TAG_TYPE_HINTS: dict[str, str] = {
}
# Words that describe a unit's role rather than identifying it. Whisper hears
# the same officer as "Post 5", "5", and "5 post" within one conversation, so
# these carry no distinguishing information and are dropped before comparison.
_UNIT_NOISE_TOKENS = frozenset({"post", "unit", "units", "car"})
_ORDINAL_RE = re.compile(r"^(\d+)(?:st|nd|rd|th)$")
def _normalize_unit(unit: str) -> str:
"""
Reduce a spoken unit ID to a comparison key.
Dispatch audio names the same unit inconsistently and exact string equality
silently drops the follow-ups. Observed on 2026-08-17 in a single hour, each
pair being one unit that failed to match itself:
"K-9A2" vs "K-9-A-2" punctuation
"5-1-6" vs "516" digits read out individually
"37" vs "37th Post" ordinal + role word
"11-Victor" vs "11 Victor"
Case, punctuation and role words go; digit groups join up. What is
deliberately NOT done is matching a bare district letter — "Adam" is not
treated as "6-Adam", because every district has an Adam and collapsing them
would merge unrelated incidents. That costs a few links and is the right
trade: a missed link leaves an orphan the sweep can retry, a false link
corrupts an incident permanently.
"""
tokens = [t for t in re.split(r"[^a-z0-9]+", unit.strip().lower()) if t]
cleaned: list[str] = []
for tok in tokens:
if tok in _UNIT_NOISE_TOKENS:
continue
ordinal = _ORDINAL_RE.match(tok)
cleaned.append(ordinal.group(1) if ordinal else tok)
key = "".join(cleaned)
# A unit made only of noise words ("Post") would normalise to "" and then
# collide with every other such unit, so fall back to the raw text.
return key or unit.strip().lower()
def _unit_keys(units: Optional[list[str]]) -> set[str]:
"""Comparison keys for a unit list, empties dropped."""
return {k for k in (_normalize_unit(u) for u in (units or [])) if k}
def _matching_units(call_units: Optional[list[str]], inc_units: Optional[list[str]]) -> list[str]:
"""
Call-side units that also appear on the incident, compared by normalised key
but returned as the original spoken strings so debug output stays readable.
"""
inc_keys = _unit_keys(inc_units)
if not inc_keys:
return []
return [u for u in (call_units or []) if _normalize_unit(u) in inc_keys]
def _infer_type_from_tags(tags: list[str]) -> Optional[str]:
"""Return an incident type inferred from tags, or None if ambiguous."""
for tag in tags:
@@ -461,8 +517,7 @@ def _run_decision(ctx: dict) -> dict:
"corr_is_dispatch": is_dispatch,
}
if fit_signal == "unit_overlap" and call_units:
inc_unit_set = set(candidate.get("units") or [])
corr_debug["corr_matched_units"] = [u for u in call_units if u in inc_unit_set]
corr_debug["corr_matched_units"] = _matching_units(call_units, candidate.get("units"))
logger.info(
f"Correlator fast-path: call {call_id} → {candidate['incident_id']} "
f"(signal={fit_signal}, is_dispatch={is_dispatch})"
@@ -502,8 +557,7 @@ def _run_decision(ctx: dict) -> dict:
"corr_is_dispatch": is_dispatch,
}
if fit_signal == "unit_overlap" and call_units:
inc_unit_set = set(candidate.get("units") or [])
corr_debug["corr_matched_units"] = [u for u in call_units if u in inc_unit_set]
corr_debug["corr_matched_units"] = _matching_units(call_units, candidate.get("units"))
logger.info(
f"Correlator fast-path (disambig {len(tg_recent)} candidates): "
f"call {call_id} → {candidate['incident_id']} (signal={fit_signal})"
@@ -524,11 +578,11 @@ def _run_decision(ctx: dict) -> dict:
# incident, the officer has moved on and we don't link back to the old one.
# This correctly handles officers dispatched to a second call mid-shift.
if not matched_incident and call_units and system_id and not reassignment:
call_unit_set = set(call_units)
call_unit_set = _unit_keys(call_units)
unit_candidates = [
inc for inc in all_active
if system_id in (inc.get("system_ids") or [])
and call_unit_set & set(inc.get("units") or [])
and call_unit_set & _unit_keys(inc.get("units"))
]
# Apply idle cap: units get reassigned; a 20+ min gap means the officer
# has almost certainly moved on or the incident closed.
@@ -540,7 +594,7 @@ def _run_decision(ctx: dict) -> dict:
best_unit_inc = max(unit_candidates, key=lambda i: i.get("updated_at", ""))
reassigned_away = any(
inc["incident_id"] != best_unit_inc["incident_id"]
and call_unit_set & set(inc.get("units") or [])
and call_unit_set & _unit_keys(inc.get("units"))
and inc.get("updated_at", "") > best_unit_inc.get("updated_at", "")
for inc in all_active
)
@@ -615,7 +669,7 @@ def _run_decision(ctx: dict) -> dict:
# • embedding similarity >= cross-TG threshold (same subject matter)
# Requiring 2+ shared units prevents single-officer false positives.
if not matched_incident and call_embedding and incident_type and call_units and system_id:
call_unit_set = set(call_units)
call_unit_set = _unit_keys(call_units)
best_cross_score = 0.0
best_cross_inc: Optional[dict] = None
for inc in recent:
@@ -623,7 +677,7 @@ def _run_decision(ctx: dict) -> dict:
continue
if system_id not in (inc.get("system_ids") or []):
continue
inc_units_set = set(inc.get("units") or [])
inc_units_set = _unit_keys(inc.get("units"))
if len(call_unit_set & inc_units_set) < 2:
continue
inc_embedding = inc.get("embedding")
@@ -635,7 +689,7 @@ def _run_decision(ctx: dict) -> dict:
best_cross_inc = inc
if best_cross_inc and best_cross_score >= settings.embedding_cross_tg_threshold:
matched_incident = best_cross_inc
shared = len(call_unit_set & set(best_cross_inc.get("units") or []))
shared = len(call_unit_set & _unit_keys(best_cross_inc.get("units")))
corr_debug = {
"corr_path": "cross-tg",
"corr_score": round(best_cross_score, 4),
@@ -951,8 +1005,8 @@ def _disambiguate(
elif idle_min < 15: score += 1.0
# > 15 min: no bonus — older incidents compete on content/units only
inc_units = set(inc.get("units") or [])
if inc_units and call_units and any(u in inc_units for u in call_units):
inc_units = _unit_keys(inc.get("units"))
if inc_units and call_units and (inc_units & _unit_keys(call_units)):
score += 10.0
inc_vehicles = set(inc.get("vehicles") or [])
@@ -1047,8 +1101,8 @@ def _call_fits_incident(
inc_id = inc.get("incident_id", "?")
# ── 1. Unit overlap ───────────────────────────────────────────────────────
inc_units = set(inc.get("units") or [])
matched_units = [u for u in call_units if u in inc_units] if (inc_units and call_units) else []
inc_units = _unit_keys(inc.get("units"))
matched_units = _matching_units(call_units, inc.get("units"))
if matched_units:
if is_dispatch:
if call_coords:
+60 -1
View File
@@ -16,7 +16,9 @@ Firestore. _update_incident writes, so its test patches fstore.
import pytest
from datetime import datetime, timedelta, timezone
from unittest.mock import AsyncMock, patch
from app.internal.incident_correlator import _run_decision, _update_incident
from app.internal.incident_correlator import (
_run_decision, _update_incident, _normalize_unit, _matching_units,
)
NOW = datetime(2026, 8, 16, 21, 0, 0, tzinfo=timezone.utc)
@@ -188,3 +190,60 @@ async def test_substantive_link_does_refresh_updated_at():
updates = mock_fstore.doc_set.await_args.args[2]
assert updates["updated_at"] == NOW.isoformat()
assert "last_thin_at" not in updates
# ---------------------------------------------------------------------------
# Unit-ID normalisation — dispatch names the same unit several ways
# ---------------------------------------------------------------------------
@pytest.mark.parametrize("spoken,other", [
("K-9A2", "K-9-A-2"), # punctuation only
("5-1-6", "516"), # digits read out individually
("37", "37th Post"), # ordinal + role word
("11-Victor", "11 Victor"), # hyphen vs space
("Post 5", "5"), # bare role word
("post 1-2", "Post 1-2"), # case
])
def test_same_unit_spoken_differently_normalises_alike(spoken, other):
"""Every pair here was observed as one real unit failing to match itself."""
assert _normalize_unit(spoken) == _normalize_unit(other)
@pytest.mark.parametrize("a,b", [
("6-Adam", "Adam"), # every district has an Adam — must stay distinct
("6-Adam", "7-Adam"),
("11-Victor", "11-Xray"),
("516", "517"),
("3", "39"),
])
def test_genuinely_different_units_stay_distinct(a, b):
assert _normalize_unit(a) != _normalize_unit(b)
def test_role_only_unit_does_not_collapse_to_empty():
"""
"Post" is all noise words. Normalising it to "" would make every such unit
equal to every other, so it falls back to the raw text instead.
"""
assert _normalize_unit("Post") != ""
assert _normalize_unit("Post") != _normalize_unit("Unit")
def test_matching_units_reports_the_original_spoken_strings():
"""Debug output has to stay readable, so matches come back un-normalised."""
assert _matching_units(["K-9A2", "6-Adam"], ["K-9-A-2"]) == ["K-9A2"]
def test_matching_units_empty_when_nothing_overlaps():
assert _matching_units(["6-Adam"], ["7-Adam", "516"]) == []
def test_normalised_units_link_a_call_that_exact_match_would_orphan():
"""End-to-end: the K-9A2 case that orphaned in the 2026-08-17 01:05Z dump."""
inc = _incident(2.0)
inc["units"] = ["K-9-A-2"]
decision = _run_decision(_ctx(
all_active=[inc], recent=[inc],
call_units=["K-9A2"], is_thin_call=False, call_severity="routine",
))
assert decision["action"] == "link"
+7
View File
@@ -0,0 +1,7 @@
{
"//": "Deploy target for firestore.rules / firestore.indexes.json only — this is not a Firebase Hosting project config. Run from this directory: firebase deploy --only firestore:rules,firestore:indexes --project <project-id>. No project id is pinned here deliberately (see [[self-hosted-infra-pointers]] for where that value lives) — pass --project explicitly or run `firebase use <project-id>` once first.",
"firestore": {
"rules": "firestore.rules",
"indexes": "firestore.indexes.json"
}
}
+47
View File
@@ -0,0 +1,47 @@
{
"//": "Composite indexes required once drb-frontend's lib/use*.ts hooks add a where(\"org_id\",\"==\",orgId) equality filter alongside an existing range/orderBy/array-contains clause. Firestore auto-indexes single-field lookups and equality-only compound queries, but org_id==X combined with an inequality, orderBy on a different field, or array-contains needs an explicit composite index or the query fails at runtime with a FAILED_PRECONDITION 'index required' error (see SAAS_PLAN.md 2.8/B2 and the URL Firestore prints in that error, which is the fastest way to double-check this list against the live query shapes). Deploy with: firebase deploy --only firestore:indexes --project <project-id>",
"indexes": [
{
"collectionGroup": "calls",
"queryScope": "COLLECTION",
"fields": [
{ "fieldPath": "org_id", "order": "ASCENDING" },
{ "fieldPath": "started_at", "order": "ASCENDING" }
]
},
{
"collectionGroup": "calls",
"queryScope": "COLLECTION",
"fields": [
{ "fieldPath": "org_id", "order": "ASCENDING" },
{ "fieldPath": "incident_ids", "arrayConfig": "CONTAINS" }
]
},
{
"collectionGroup": "incidents",
"queryScope": "COLLECTION",
"fields": [
{ "fieldPath": "org_id", "order": "ASCENDING" },
{ "fieldPath": "started_at", "order": "ASCENDING" }
]
},
{
"collectionGroup": "alert_events",
"queryScope": "COLLECTION",
"fields": [
{ "fieldPath": "org_id", "order": "ASCENDING" },
{ "fieldPath": "triggered_at", "order": "ASCENDING" }
]
},
{
"collectionGroup": "alert_events",
"queryScope": "COLLECTION",
"fields": [
{ "fieldPath": "org_id", "order": "ASCENDING" },
{ "fieldPath": "acknowledged", "order": "ASCENDING" },
{ "fieldPath": "triggered_at", "order": "ASCENDING" }
]
}
],
"fieldOverrides": []
}
+166
View File
@@ -0,0 +1,166 @@
// Firestore security rules — the actual tenant boundary for DRB.
//
// WHY THIS FILE EXISTS: drb-frontend reads Firestore directly from the
// browser (see lib/use*.ts — onSnapshot(collection(db, ...))), so c2-core's
// app/internal/auth.py is NOT in that read path at all. These rules are the
// only thing standing between a signed-in stranger and every org's radio
// traffic. Before this file existed, whatever rules were live had been
// hand-set in the Firebase console: unversioned, unreviewed, unknown. See
// SAAS_PLAN.md B1.
//
// DEPLOY IS A MANUAL, OUT-OF-BAND STEP — nothing in CI or this codebase
// pushes these rules to Firebase:
// firebase deploy --only firestore:rules --project <project-id>
// (from this directory, or point --config at infra/firestore/firebase.json
// from the repo root). Do this before or immediately after the code that
// starts stamping org_id ships — until these rules are live, the
// console-configured rules are still what's actually enforced.
//
// MODEL: c2-core (firebase-admin SDK, server-side) bypasses these rules
// entirely and is the sole writer for every collection below — that was
// already the architecture (see CLAUDE.md "Auth — three distinct
// mechanisms"). These rules therefore only need to gate READS for the
// browser client, and can safely deny ALL client writes.
//
// Deny-by-default: the catch-all match at the bottom denies anything not
// explicitly listed above it, including collections added later that
// someone forgets to add a rule for.
rules_version = '2';
service cloud.firestore {
match /databases/{database}/documents {
function signedIn() {
return request.auth != null;
}
// Platform-level role (admin/operator/viewer) — set by drb-c2-core
// routers/users.py custom claims. Distinct from org_role (owner/member),
// which is per-organization. A platform admin can read across every org
// (support/debugging), mirroring internal/auth.py's require_org()
// ?org_id= override for the same role.
function isPlatformAdmin() {
return signedIn() &&
(request.auth.token.role == 'admin' || request.auth.token.admin == true);
}
// The org_id claim is set by POST /auth/signup (or /admin/users) at
// account-provisioning time. No claim => no access, by construction —
// this is what backs AuthProvider's no-claim guard (SAAS_PLAN.md B3):
// a user with no org_id claim can hold a valid Firebase session and
// still read nothing here.
function myOrgId() {
return request.auth.token.org_id;
}
function inOrg(orgId) {
return signedIn() && (myOrgId() == orgId || isPlatformAdmin());
}
function docInMyOrg() {
return inOrg(resource.data.org_id);
}
// ── Org identity ──────────────────────────────────────────────────────
match /organizations/{orgId} {
allow read: if inOrg(orgId);
allow write: if false; // c2-core only (POST /auth/signup, routers/org.py)
}
match /org_members/{uid} {
allow read: if signedIn() && (request.auth.uid == uid || isPlatformAdmin() ||
inOrg(resource.data.org_id));
allow write: if false; // c2-core only
}
// ── Tenant-scoped radio data — the whole point of this file ───────────
match /nodes/{nodeId} {
allow read: if docInMyOrg();
allow write: if false;
}
match /systems/{systemId} {
allow read: if docInMyOrg();
allow write: if false;
}
match /calls/{callId} {
allow read: if docInMyOrg();
allow write: if false;
}
match /incidents/{incidentId} {
allow read: if docInMyOrg();
allow write: if false;
}
match /alert_events/{alertId} {
allow read: if docInMyOrg();
allow write: if false;
}
match /alert_rules/{ruleId} {
allow read: if docInMyOrg();
allow write: if false;
}
// ── Never client-readable, org-scoped or not ───────────────────────────
// Secrets / credential material. Reads for these go through c2-core
// REST routes (which apply their own auth), never straight to Firestore.
match /node_keys/{nodeId} {
allow read, write: if false;
}
match /enrollment_tokens/{tokenHash} {
allow read, write: if false; // routers/org.py mints/lists/revokes server-side
}
match /org_api_keys/{keyId} {
allow read, write: if false; // not implemented server-side yet (DEFERRED.md) — deny regardless
}
// ── Platform-admin-only collections ────────────────────────────────────
// Listed here mainly so the deny-default catch-all's intent is explicit;
// these are already read/written exclusively through c2-core admin
// routes (require_admin_token), never straight from the browser.
match /audit_log/{entryId} {
allow read, write: if false;
}
match /config/{docId} {
allow read, write: if false;
}
match /bot_tokens/{tokenId} {
allow read, write: if false;
}
match /waitlist/{entryId} {
allow read, write: if false; // POST /waitlist writes via the admin SDK
}
// ── Trips (internal utility feature, not org-scoped — see
// [[trips-feature-intentional]] and SAAS_PLAN.md B7. Frontend hides
// /trips outside the founding org; these rules keep the existing
// "public unless flagged private" trip model working for whichever
// users the UI still exposes it to) ────────────────────────────────────
match /trips/{tripId} {
allow read: if signedIn();
allow write: if false;
}
match /trip_events/{eventId} {
allow read: if signedIn();
allow write: if false;
}
// ── Deny-by-default catch-all ──────────────────────────────────────────
// Anything not explicitly matched above — including collections added
// later without a corresponding rule — is denied. This is the guard
// rail: a missing rule fails closed, not open.
match /{document=**} {
allow read, write: if false;
}
}
}