Correlation has only ever been measured through live AI windows: days of
wall time per change, and the 09-20→22 window was invalidated outright by
unfunded AI accounts (#169). Recordings are kept regardless of AI, so the
traffic to measure against already exists.
- internal/replay.py: runs a time range of real calls through the live
pipeline code in original order, clock pinned per call, into
replay_runs/{run_id}/calls|incidents. Modes: audio (re-transcribe),
transcripts (re-extract), reuse (correlation only from a prior run's
scenes). Simulates the idle-resolve and orphan-recorrelation sweeps on
virtual time. No alerts, summaries, vocab, AI-health alerts or pending
terms. One run at a time, <=5000 calls, <=7 days.
- firestore.py: ContextVar sandbox redirect for calls/incidents.
- clock.py: ContextVar-pinnable now(), used on the correlation path.
- feature_flags.py: ContextVar flag override so replay runs with live AI off.
- upload.py: scene loop extracted to _extract_and_correlate, shared by the
live pipeline and replay so replay measures the code that runs live.
- resolved_via on every incident resolve, so a real clear can be told
from the idle timeout — live and in replay.
- routers/replay.py + /admin Replay tab: estimate, start, compare runs,
drill into incidents with audio.
Reviewed by drb-correlation-review; its leak and fidelity findings are
fixed and covered by tests. c2-core: 456 pass. Frontend typecheck not run
(no Node on the authoring box).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Three AI dependency failures in one night (retired Gemini model IDs,
depleted Gemini balance, unpayable OpenAI account) each surfaced only
as a single ERROR log line that nobody was watching. Add
app/internal/ai_health.py, a shared in-memory registry that
transcription.py and llm_correlator.py report into on every call
(success and failure), distinguishing permanent conditions (dead
model, dead billing) which alert immediately from transient ones
(rate limits, network blips) which only alert after they persist.
Alerts POST once per degradation episode and once on recovery to an
optional Discord webhook (AI_ALERT_WEBHOOK_URL), reusing alerter.py's
httpx pattern. State is exposed unauthenticated at GET /health/ai
alongside the existing /health.
Closeslogan/server-26#14