Files
server-26/.gitea/workflows/deploy.yml
T
Logan Cusano 861ea41cec
Build & Deploy / Build & push images (push) Successful in 4m4s
Build & Deploy / Deploy to VM (push) Successful in 1m26s
Build & Deploy / Report a failed deploy (push) Skipped
Deploy the commit's own images instead of :latest
The build-stamp health check added in 8fbfe7d worked on its first run, and
what it caught was not a stale container -- it was a race. Runs 544 and 545
overlapped; both deployed :latest, 545's images won, and 544's health check
correctly reported that the build serving traffic was not the one it had just
deployed.

That is a real hazard, not a false positive: with :latest, two pushes landing
close together means whichever finishes last silently wins for BOTH, and
neither run's log tells you which code is actually live. Pushes land close
together constantly here.

docker-compose.yml already resolved images as ${TAG:-latest}, so the fix is to
export TAG=<commit sha> for the deploy. Each run now pulls and starts exactly
the images it built, rollback becomes "deploy a different tag", and the health
check's assertion becomes meaningful rather than order-dependent. A manual
`docker compose up -d` on the VM with no TAG set still falls back to :latest,
which is the intended escape hatch.

Also replaces the health check's single `sleep 20` with a poll of up to 100s
that stops as soon as the expected SHA appears. A fixed sleep is either too
short -- flaky red runs -- or wastes time on every deploy, and a check that
cries wolf gets ignored, which is exactly the failure this job exists to stop.

Refs logan/server-26#21
2026-08-23 01:38:09 -04:00

182 lines
6.9 KiB
YAML

name: Build & Deploy
on:
push:
branches: [main]
env:
# REGISTRY secret = "git.vpn.cusano.net/logan" (full image prefix)
REGISTRY: ${{ secrets.REGISTRY }}
jobs:
build:
name: Build & push images
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Log in to Gitea registry
uses: docker/login-action@v3
with:
registry: git.vpn.cusano.net
username: ${{ secrets.REGISTRY_USER }}
password: ${{ secrets.BUILD_TOKEN }}
- name: Build & push c2-core
uses: docker/build-push-action@v5
with:
context: ./drb-c2-core
push: true
build-args: |
GIT_SHA=${{ gitea.sha }}
tags: |
${{ env.REGISTRY }}/c2-core:latest
${{ env.REGISTRY }}/c2-core:${{ gitea.sha }}
- name: Build & push discord-bot
uses: docker/build-push-action@v5
with:
context: ./drb-server-discord-bot
push: true
tags: |
${{ env.REGISTRY }}/discord-bot:latest
${{ env.REGISTRY }}/discord-bot:${{ gitea.sha }}
- name: Build & push frontend
uses: docker/build-push-action@v5
with:
context: ./drb-frontend
push: true
tags: |
${{ env.REGISTRY }}/frontend:latest
${{ env.REGISTRY }}/frontend:${{ gitea.sha }}
build-args: |
NEXT_PUBLIC_C2_URL=https://api.${{ secrets.DRB_DOMAIN }}
NEXT_PUBLIC_FIREBASE_API_KEY=${{ secrets.FIREBASE_API_KEY }}
NEXT_PUBLIC_FIREBASE_AUTH_DOMAIN=${{ secrets.FIREBASE_AUTH_DOMAIN }}
NEXT_PUBLIC_FIREBASE_PROJECT_ID=${{ secrets.FIREBASE_PROJECT_ID }}
NEXT_PUBLIC_FIREBASE_STORAGE_BUCKET=${{ secrets.FIREBASE_STORAGE_BUCKET }}
NEXT_PUBLIC_FIREBASE_MESSAGING_SENDER_ID=${{ secrets.FIREBASE_MESSAGING_SENDER_ID }}
NEXT_PUBLIC_FIREBASE_APP_ID=${{ secrets.FIREBASE_APP_ID }}
NEXT_PUBLIC_FIRESTORE_DATABASE=${{ secrets.FIRESTORE_DATABASE }}
deploy:
name: Deploy to VM
needs: build
runs-on: ubuntu-latest
steps:
- name: Check runner outbound IP
run: curl -s ifconfig.me
- name: Write SSH key
run: |
printf '%s\n' "${{ secrets.SSH_PRIVATE_KEY }}" > /tmp/deploy_key
chmod 600 /tmp/deploy_key
ssh-keygen -l -f /tmp/deploy_key
- name: Deploy
run: |
ssh -o StrictHostKeyChecking=no \
-o HostKeyAlgorithms=ssh-ed25519,rsa-sha2-256,rsa-sha2-512 \
-o ConnectTimeout=15 \
-v \
-i /tmp/deploy_key \
drb@${{ secrets.SERVER_IP }} << 'ENDSSH'
set -e
cd /opt/drb
# Update compose files + mosquitto config
git pull origin main
# Deploy THIS commit's images, not :latest. Overlapping runs are
# normal here, and with :latest whichever finishes last wins for
# both -- run 544 asserted its own SHA and found run 545's build
# already serving. compose already supports ${TAG:-latest}, so
# pinning makes each deploy deterministic and a rollback just a
# different tag. A later manual `up -d` on the VM without TAG set
# still falls back to :latest, which is the intended escape hatch.
export TAG=${{ gitea.sha }}
# Pull pre-built images and restart (no build on the VM).
#
# The retry is not defensive padding: this exact step failed fifteen
# deploys in a row (2026-08-18 to 08-20) with containerd unable to
# extract a layer -- "failed to Lchown ... no such file or directory"
# -- a corrupted entry in the snapshot store. Pruning clears the bad
# layer and the second pull succeeds. If it fails again after a
# prune that is a real problem (check the VM's disk) and should stop
# the deploy rather than be retried forever.
COMPOSE="docker compose -f docker-compose.yml -f docker-compose.prod.yml"
if ! $COMPOSE pull; then
echo "image pull failed - pruning and retrying once"
docker image prune -af
$COMPOSE pull
fi
$COMPOSE up -d --remove-orphans
docker image prune -f
ENDSSH
- name: Health check
run: |
# Poll rather than sleep-once: the container has to finish starting,
# and a fixed sleep is either too short (flaky red) or wastes time on
# every deploy. A health check that cries wolf gets ignored, which is
# the failure mode this whole job exists to prevent.
BODY=""
for _ in $(seq 1 20); do
sleep 5
BODY=$(curl -fsS https://api.${{ secrets.DRB_DOMAIN }}/health) || continue
case "$BODY" in *"${{ gitea.sha }}"*) break ;; esac
done
if [ -z "$BODY" ]; then
echo "Health check failed: /health never responded"; exit 1
fi
echo "$BODY"
# Liveness alone is not enough. A deploy can report success while the
# PREVIOUS container keeps serving -- that is how production ran
# 08-18 code for two days without a single red run. Assert that the
# build which answered is the commit we just pushed.
RUNNING=$(printf '%s' "$BODY" | tr ',' '\n' | grep git_sha | cut -d'"' -f4)
if [ "$RUNNING" != "${{ gitea.sha }}" ]; then
echo "Deployed build is '$RUNNING', expected '${{ gitea.sha }}'."
echo "The container was not actually replaced."
exit 1
fi
notify-failure:
name: Report a failed deploy
needs: [build, deploy]
if: failure()
runs-on: ubuntu-latest
steps:
- name: Post to Discord
# A red run in Gitea is only visible to someone who opens Gitea, and
# nobody did for two days. Same shape as an AI tier dying quietly,
# which is why both now push a message out of the box instead of
# waiting to be discovered. No webhook configured => skip quietly
# rather than fail, since not every deployment will set one.
env:
WEBHOOK: ${{ secrets.DEPLOY_ALERT_WEBHOOK }}
RUN_URL: ${{ gitea.server_url }}/${{ gitea.repository }}/actions/runs/${{ gitea.run_number }}
SHA: ${{ gitea.sha }}
run: |
if [ -z "$WEBHOOK" ]; then
echo "DEPLOY_ALERT_WEBHOOK is not set - skipping notification."
exit 0
fi
python3 - <<'PY' > /tmp/payload.json
import json, os
print(json.dumps({"content":
"**DRB deploy failed** on `%s`\n%s\nProduction is still running the previous build."
% (os.environ["SHA"][:8], os.environ["RUN_URL"])}))
PY
curl -sS -X POST -H "Content-Type: application/json" \
--data @/tmp/payload.json "$WEBHOOK" || echo "notification POST failed"