intelligence: traffic stops and self-initiated activity open incidents #178

Merged
logan merged 1 commits from feat/traffic-stops into main 2026-09-26 20:22:05 -04:00
2 changed files with 55 additions and 1 deletions
Showing only changes of commit badfe28823 - Show all commits
+39 -1
View File
@@ -67,7 +67,7 @@ Rules:
- tags: describe WHAT happened, not WHERE. Specific, lowercase, hyphenated. Do not use location names, road names, talkgroup names, or place names as tags (wrong: "lower-macy's", "canvas-route-6", "route-202"; right: "suspect-search", "shoplifting", "vehicle-pursuit"). Do not repeat incident_type as a tag. - tags: describe WHAT happened, not WHERE. Specific, lowercase, hyphenated. Do not use location names, road names, talkgroup names, or place names as tags (wrong: "lower-macy's", "canvas-route-6", "route-202"; right: "suspect-search", "shoplifting", "vehicle-pursuit"). Do not repeat incident_type as a tag.
- units: ONLY identifiers that appear verbatim in the transcript. Use speaker role inference to distinguish units being dispatched from units acknowledging — both should be included. Never infer or guess unit IDs not present in the text. If a unit ID format is given below, use it to recognise a unit spoken in a shortened or partial form (e.g. just the phonetic name alone) as the same unit — but still only extract what is actually said, never fabricate the full form. - units: ONLY identifiers that appear verbatim in the transcript. Use speaker role inference to distinguish units being dispatched from units acknowledging — both should be included. Never infer or guess unit IDs not present in the text. If a unit ID format is given below, use it to recognise a unit spoken in a shortened or partial form (e.g. just the phonetic name alone) as the same unit — but still only extract what is actually said, never fabricate the full form.
- Do not invent details not present in the transcript. - Do not invent details not present in the transcript.
- incident_type: FIRST decide whether this transmission has any incident behind it at all, using the same bar as the "routine" severity rule below — pure administrative/status traffic with nothing describable happening: post/unit check-ins, roll call, bare acknowledgements ("10-4", "copy", "received"), records/report exchanges, "show me admin"/"show me available", a status ten-code with no event attached. If it is administrative/status-only, return "unknown" — this applies on EVERY channel, including a police channel; do not let the channel default override it (server-26#138: forcing a channel default onto content-free chatter is what let radio housekeeping open incidents). Only once real event content is present, let the talkgroup channel be your primary signal for WHICH type. Use "fire" ONLY if the talkgroup is clearly a fire/rescue channel OR the transcript explicitly describes active fire, smoke, flames, or structure fire activation. Police or EMS referencing a fire scene → use "police" or "ems". When the channel is a police channel, a real event is present, and nothing in the transcript contradicts it, return "police". Reserve "other" for a real event that genuinely belongs to no emergency service (rail operations, public works, utility coordination) — not for administrative chatter, which is "unknown" per above regardless of channel. Also reserve "unknown" for transcripts too garbled to place at all. - incident_type: FIRST decide whether this transmission has any incident behind it at all, using the same bar as the "routine" severity rule below — pure administrative/status traffic with nothing describable happening: post/unit check-ins, roll call, bare acknowledgements ("10-4", "copy", "received"), records/report exchanges, "show me admin"/"show me available", a status ten-code with no event attached. If it is administrative/status-only, return "unknown" — this applies on EVERY channel, including a police channel; do not let the channel default override it (server-26#138: forcing a channel default onto content-free chatter is what let radio housekeeping open incidents). Only once real event content is present, let the talkgroup channel be your primary signal for WHICH type. Use "fire" ONLY if the talkgroup is clearly a fire/rescue channel OR the transcript explicitly describes active fire, smoke, flames, or structure fire activation. Police or EMS referencing a fire scene → use "police" or "ems". When the channel is a police channel, a real event is present, and nothing in the transcript contradicts it, return "police". Reserve "other" for a real event that genuinely belongs to no emergency service (rail operations, public works, utility coordination) — not for administrative chatter, which is "unknown" per above regardless of channel. Also reserve "unknown" for transcripts too garbled to place at all. A unit reporting its OWN activity is a real event, not status traffic: "on a stop" / traffic stop / car stop, "out with a vehicle", "put me out with a pedestrian/subject" — return "police", tag it (e.g. "traffic-stop", "pedestrian-assist"), severity at least "minor". The plate/license lookups for that stop belong to it.
- severity: ALWAYS return one of the four values. Judge the underlying event, not how dramatic the words sound. - severity: ALWAYS return one of the four values. Judge the underlying event, not how dramatic the words sound.
"routine" — administrative/status traffic with no incident behind it: mileage and transport logging, radio checks, acknowledgements, shift changes, track block/power requests, records lookups. "routine" — administrative/status traffic with no incident behind it: mileage and transport logging, radio checks, acknowledgements, shift changes, track block/power requests, records lookups.
"minor" — a real but low-stakes call: lift assist, parking complaint, past-tense larceny report, noise complaint, welfare check. "minor" — a real but low-stakes call: lift assist, parking complaint, past-tense larceny report, noise complaint, welfare check.
@@ -418,6 +418,10 @@ async def extract_scenes(
transcript, segments, segment_indices, transcript_corrected transcript, segments, segment_indices, transcript_corrected
) )
tags, incident_type, severity = _self_initiated_backstop(
scene_transcript or transcript, tags, incident_type, severity
)
processed.append({ processed.append({
"tags": tags, "tags": tags,
"incident_type": incident_type, "incident_type": incident_type,
@@ -477,6 +481,40 @@ async def extract_scenes(
return processed return processed
# Self-initiated activity: a unit putting itself "on a stop" or "out with" a
# vehicle/pedestrian. Replay of 09-22 (server-26#170): every traffic stop on
# the Ch 1 channel ("45 Adam on a stop, Eastbound Central Express", "CM2 on
# the stop, southbound") came back untyped/untagged/routine from extraction —
# read as status traffic — so the creation gate never opened an incident and
# the stop was visible only in the archive. The prompt now says so too; this
# is the deterministic backstop, because a tag is what the creation gate
# counts as substance (incident_correlator.has_event_substance).
_SELF_INITIATED = (
(re.compile(r"\b(on (a|the) (traffic |car |vehicle )?stop|traffic stop|car stop|vehicle stop|"
r"pull(ed|ing)? (a car |a vehicle |him |her |them )?over)\b", re.IGNORECASE),
"traffic-stop"),
(re.compile(r"\b((put|show) me out with|out with (a|one) (pedestrian|vehicle|disabled|male|female|"
r"subject|party|juvenile))\b", re.IGNORECASE),
"self-initiated"),
)
_NEGATED = re.compile(r"\b(not|don't|dont|no|never)\s+(\w+\s+){0,2}$", re.IGNORECASE)
def _self_initiated_backstop(
text: str, tags: list, incident_type: Optional[str], severity: str,
) -> tuple[list, Optional[str], str]:
for pattern, tag in _SELF_INITIATED:
m = pattern.search(text or "")
if not m or _NEGATED.search(text[: m.start()]):
continue
if tag not in tags:
tags = [*tags, tag]
incident_type = incident_type or "police"
if severity == "routine":
severity = "minor"
return tags, incident_type, severity
# "45-9, I'm clear." / "Vehicle 1, clear." / "Car 12 10-8" — a unit reporting # "45-9, I'm clear." / "Vehicle 1, clear." / "Car 12 10-8" — a unit reporting
# itself back in service is the one signal that ends an incident, and it is # itself back in service is the one signal that ends an incident, and it is
# almost always five words or fewer, which is exactly the population the # almost always five words or fewer, which is exactly the population the
@@ -99,3 +99,19 @@ def test_reopen_only_for_a_call_after_the_close():
closed = {"resolved_at": "2026-09-22T15:00:00+00:00"} closed = {"resolved_at": "2026-09-22T15:00:00+00:00"}
assert ic._after_close(closed, datetime(2026, 9, 22, 15, 5, tzinfo=timezone.utc)) assert ic._after_close(closed, datetime(2026, 9, 22, 15, 5, tzinfo=timezone.utc))
assert not ic._after_close(closed, datetime(2026, 9, 22, 14, 55, tzinfo=timezone.utc)) assert not ic._after_close(closed, datetime(2026, 9, 22, 14, 55, tzinfo=timezone.utc))
def test_traffic_stops_become_events():
from app.internal.intelligence import _self_initiated_backstop as b
for t in ("45 Adam on a stop, Eastbound Central Express.",
"11-0. CM2 on the stop, southbound, KFLA on the right.",
"Car 7, traffic stop, Route 9 at Main"):
tags, typ, sev = b(t, [], None, "routine")
assert "traffic-stop" in tags and typ == "police" and sev == "minor", t
tags, typ, sev = b("Charlie 1. You put me out with a pedestrian on a parkway", [], None, "routine")
assert "self-initiated" in tags
# negation and unrelated chatter stay untouched
assert b("Do you want me to not pull the car over", [], None, "routine") == ([], None, "routine")
assert b("45-8, go ahead.", [], None, "routine") == ([], None, "routine")
# an existing type/severity is never downgraded
assert b("on a stop", ["dwi"], "police", "moderate") == (["dwi", "traffic-stop"], "police", "moderate")