Transcription is the top of the pipeline and it fails soft: any exception logs
a WARNING, returns None, and upload.py carries on. That is the right behaviour
for a network blip and exactly the wrong behaviour for an unpayable account,
because with no transcript there is no extraction, no correlation and no
incident -- the system keeps accepting calls and quietly stores empty ones,
which looks like quiet radio traffic rather than an outage.
This is the third instance of the same failure mode today. The Gemini
correlator was down first on a retired model ID and then on a depleted
balance, and in both cases the only signal was a per-call WARNING that read as
noise. The OpenAI balance is low enough that this one is a matter of when.
Billing-shaped errors (insufficient_quota, billing, credit, quota exceeded)
now log once at ERROR, name what is dead downstream, and link the top-up page.
Everything else keeps the existing per-call WARNING.
No new environment variables, so CI deploys this without an ansible run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two independent sources of garbage in the AI pipeline, both visible in the
2026-08-16 correlation dump.
1. Hallucinated transcripts. The Whisper prompt opened with an enumerated run
of ten-codes: 10-4, 10-23, 10-20, 10-97 and so on. Whisper treats prompt
text as preceding transcript, so on noisy or silent audio it continued the
series, emitting transcripts that count upward from 10-4 to 10-99. The
existing no_speech_prob filter could not catch these: the model is highly
confident in text it invented by continuing a pattern.
The prompt no longer contains a series to extend, and _is_degenerate()
rejects the three shapes this failure takes: ascending ten-code runs, one
phrase looping, and near-identical segments across a whole recording.
Verified against 13 transcripts from production: all four known
hallucinations rejected, all nine real ones kept, including terse traffic
containing legitimate codes.
2. Duplicate recordings. node-002 and node-PI-2 both cover TG 9048 and both
uploaded the same transmissions, ~1.1s apart. Nine pairs appeared in one
dump. Each was transcribed, billed and correlated twice, and the resulting
incident listed two units where there was one.
Canonical selection is by earliest started_at, tie-broken on call_id, NOT
by upload order: upload order varies with encode time and network latency,
so it would make the authoritative recording non-deterministic. Call
documents are created from MQTT call_start before uploads arrive, so both
nodes independently reach the same verdict. The loser keeps its audio (it
may be the cleaner capture) but is excluded from STT, correlation, the
re-correlation sweep and the orphan debug view.
Also fixes _sync_transcribe returning a bare None when OPENAI_API_KEY is
missing, where the caller unpacks two values. A missing key surfaced as a
misleading "Transcription failed" instead of the real warning.
Adds tests/test_dedup.py (15 cases). dedup.py reaches Firestore through an
injected callable so it stays importable without firebase-admin present.
- correlator: unit_overlap on dispatch channels now applies content
divergence check when the call has geocoded coords but the incident
doesn't; previously this gap caused unrelated calls to merge into
stale incidents (e.g. patrol officer at a second scene 70 min later)
- STT: switch default model from gpt-4o-transcribe to whisper-1, which
faithfully transcribes all exchanges in multi-PTT recordings; gpt-4o
was silently dropping utterances, starving the correlation engine
- STT: remove vocabulary from the Whisper prompt; whisper-1 echoes
prompted terms into noise/silence, skewing extracted incident data;
vocabulary context is now applied exclusively in the GPT extraction
step (build_gpt_vocab_block) where it is used as reference only
Upload 404 warning doc_set(merge=True) in upload.py — creates doc if missing
MQTT call_end 404 error doc_set(merge=True) in mqtt_handler.py — same root cause
Transcription 404 (saving transcript to nonexistent doc) doc_set(merge=True) in transcription.py
Transcription ADC credentials error Explicit service_account.Credentials from gcp-key.json in _sync_transcribe — same pattern as storage.py