deploy.yml ran `compose up -d` before the health check and never reverted on failure. A build that passes tests, returns 200 on /health with the right git_sha, but has a live logic bug (exactly the class of bug the correlator instrumentation exists to catch) would stay live indefinitely - notify-failure would even claim production was "still running the previous build", which is false in that scenario. Deploy step now reads /opt/drb/.last_good_tag (written only after a prior deploy's own health check confirmed its SHA) to capture the previously- verified tag before switching, and emits it as a step output. Health check is unchanged in shape (bounded 20x5s retry, still requires the polled git_sha to match) but now persists the new SHA as the rollback target only once confirmed live. A new Rollback step runs on any failure above, re-deploys the previous tag, and re-verifies via the same git_sha check rather than trusting mere liveness - then fails the job loudly either way, since the push itself was still bad. notify-failure now reports what actually happened (rollback succeeded/failed/skipped and to which SHA) instead of the old unconditional claim. This unblocks #62 decision 9: autonomous pushes to incident_correlator.py, llm_correlator.py, intelligence.py and routers/upload.py were frozen until this rollback path landed. Refs #65, #62, #60, #57.