Files
server-26/drb-c2-core/app/internal/auth.py
T
Logan CusanoandClaude Opus 5 865b5b4317
Build & Deploy / Build & push images (push) Successful in 4m15s
Build & Deploy / Deploy to VM (push) Successful in 1m56s
Build & Deploy / Report a failed deploy (push) Skipped
Close the /admin/features side-door that needed a container shell to flip AI spend
Board minutes #62 Decision 2 (server-26#64), due 2026-08-31. CTO draft #60
finding 1 and CISO draft #61 finding 3 reached this independently.

GET/PUT /admin/features accepted only a Firebase admin token, so the unattended
runbook had no headless path and SSHed into the c2-core container to write
config/ai_features with the admin SDK. Moving a platform-wide AI cost switch
required a full container shell, and set_flags() wrote no audit entry either
way, so a flag flip was unattributable however it happened.

- New agent_service_key (AGENT_SERVICE_KEY), deliberately separate from the
  Discord bot's service_key. Sharing one key would collapse two principals into
  a single unattributable identity in every log line, and the bot has no
  business flipping AI flags regardless.
- require_agent_key_or_admin accepts the agent key or a Firebase admin, and
  rejects the Discord key. The "key is configured" guard is load-bearing:
  compare_digest("", "") is a match, so a deployment that never set the key
  would otherwise accept an empty credential.
- set_flags() writes an audit_log entry with before/after values and the actor,
  wrapped so an audit failure cannot lose the flag write or 500 the route.
- Cascade helper sets the global doc and every system carrying an ai_flags
  override in one call. A global False already beats everything, but a system
  False beats a global True, so turning AI *on* could half-apply and leave a
  radio system hot after shutoff. It scans for the override rather than
  hardcoding the two known system IDs, so a new system cannot silently defeat
  it.
- cascade defaults to False. PUT /systems/{id}/ai-flags and the AiFlagsPanel
  toggle mean a per-system override is deliberate operator intent; cascading by
  default would erase it on any unrelated global flip. The runbook opts in.

Issue items 5 and 6 (retiring the SSH path from drb-worksession.md) are NOT
done here and the runbook is untouched. The credential does not exist in
production yet, so the SSH path is still the only one that works; retiring it
now would break the next unattended run. Owner activation is recorded on #64.

Tests 273 -> 289.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-30 02:52:00 -04:00

333 lines
15 KiB
Python

import secrets
import time
from collections import defaultdict, deque
from typing import Optional
from fastapi import HTTPException, Security
from fastapi.security import HTTPBearer, HTTPAuthorizationCredentials
from firebase_admin import auth as firebase_auth
from app.config import settings
_bearer = HTTPBearer(auto_error=False)
async def require_firebase_token(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Verify a Firebase ID token from the Authorization: Bearer header."""
if not credentials:
raise HTTPException(status_code=401, detail="Missing authorization token")
try:
return firebase_auth.verify_id_token(credentials.credentials)
except Exception:
raise HTTPException(status_code=401, detail="Invalid or expired token")
async def require_service_or_firebase_token(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Accept either a Firebase ID token or the internal service key."""
if not credentials:
raise HTTPException(status_code=401, detail="Missing authorization token")
token = credentials.credentials
if settings.service_key and secrets.compare_digest(token, settings.service_key):
return {"service": True}
try:
return firebase_auth.verify_id_token(token)
except Exception:
raise HTTPException(status_code=401, detail="Invalid or expired token")
async def require_node_service_or_firebase_token(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Accept a node's own API key in addition to a service key / Firebase token.
Edge nodes need to read ``/systems`` to build their OP25 config, but they
hold neither a Firebase token nor the shared service key — only the
per-node api_key that ``/upload`` already trusts. Without this they got a
flat 401 and silently fell back to their stale offline cache, so a system
edited in the UI never reached the node.
Unlike ``/upload``, the node sends no node_id alongside the bearer token,
so the key is matched by querying ``node_keys`` for the value rather than
fetching a known document. Mutating routes are unaffected: they carry
their own ``require_admin_token`` dependency, so widening the router-level
gate grants nodes read access only.
"""
if not credentials:
raise HTTPException(status_code=401, detail="Missing authorization token")
token = credentials.credentials
if settings.service_key and secrets.compare_digest(token, settings.service_key):
return {"service": True}
try:
return firebase_auth.verify_id_token(token)
except Exception:
pass
# Deferred import: app.internal.firestore initialises firebase-admin at
# import time, and auth.py is imported from module scope in the routers.
from app.internal import firestore as fstore
matches = await fstore.collection_list("node_keys", api_key=token)
if matches:
return {"node": True, "node_id": matches[0].get("node_id")}
raise HTTPException(status_code=401, detail="Invalid or expired token")
def get_role(decoded: dict) -> str:
"""Extract the effective role from a decoded Firebase token.
Checks the granular ``role`` claim first, then falls back to the legacy
``admin`` boolean so existing tokens continue to work during the transition.
"""
if decoded.get("role") == "admin" or decoded.get("admin"):
return "admin"
role = decoded.get("role", "viewer")
return role if role in ("admin", "operator", "viewer") else "viewer"
# ---------------------------------------------------------------------------
# Tenancy — org_id / org_role claims, set by POST /auth/signup (routers/links.py)
# ---------------------------------------------------------------------------
# `role` above is platform-level (admin/operator/viewer — unrelated to which
# org a user belongs to). `org_role` is the customer-facing one: "owner" or
# "member" of the org named by the `org_id` claim. See SAAS_PLAN.md B2/B4.
def get_org_role(decoded: dict) -> Optional[str]:
org_role = decoded.get("org_role")
return org_role if org_role in ("owner", "member") else None
def require_org(decoded: dict) -> str:
"""Return the caller's org_id claim, or 403 if they don't have one.
A Firebase token with no org_id claim is a real, valid session (the user
signed in) that is nonetheless provisioned into nothing — see
AuthProvider's no-claim guard (SAAS_PLAN.md B3). Every org-scoped route
depends on this rather than trusting a client-supplied org_id, so a
caller can never read/write outside the org their own token names.
"""
org_id = decoded.get("org_id")
if not org_id:
raise HTTPException(403, "This account is not associated with an organization.")
return org_id
def resolve_org_scope(decoded: dict, org_id_override: Optional[str] = None) -> str:
"""Return the org_id a request should be scoped to.
Platform admins (role == "admin") may pass ?org_id=<id> to cross into
another org's data for support/debugging — the one exception to "you can
only ever see your own org's data" called out in SAAS_PLAN.md B2. Every
other caller is locked to their own token's org_id claim regardless of
what (if anything) they pass.
"""
if org_id_override and get_role(decoded) == "admin":
return org_id_override
return require_org(decoded)
async def resolve_caller_org_id(decoded: dict) -> Optional[str]:
"""
Resolve the org_id a caller should be scoped to, across every credential
shape this file's dependencies can produce (service key, node api_key,
Firebase user) — a single helper so read routes gated by
require_service_or_firebase_token / require_node_service_or_firebase_token
don't each need their own caller-shape switch.
Returns None for callers that should see across every org: the internal
service key (the Discord bot — a single fleet-wide principal, see
CLAUDE.md's auth section) and platform admins, matching
require_admin_token's existing "admin sees everything" behaviour. A
route that wants admins scoped too should check get_role() itself rather
than relying on this function to do it.
"""
if decoded.get("service"):
return None
if decoded.get("node"):
# Deferred import — same reasoning as require_node_service_or_firebase_token
# above: app.internal.firestore initialises firebase-admin at import
# time, and auth.py is imported from module scope in the routers.
from app.internal import firestore as fstore
node = await fstore.doc_get_cached("nodes", decoded.get("node_id") or "")
return (node or {}).get("org_id")
if get_role(decoded) == "admin":
return None
return require_org(decoded)
async def require_org_owner_token(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Verify a Firebase ID token AND require org_role == "owner" (or platform admin).
Used for org-administrative actions a regular member shouldn't be able to
do on their own org — minting/revoking enrollment tokens, renaming the
org. Platform admins pass through regardless of org_role so support can
act on an org that has no reachable owner.
"""
decoded = await require_firebase_token(credentials)
require_org(decoded)
if get_org_role(decoded) != "owner" and get_role(decoded) != "admin":
raise HTTPException(status_code=403, detail="Organization owner access required.")
return decoded
async def require_admin_token(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Verify a Firebase ID token AND require the admin role.
Accepts both the legacy ``admin: True`` boolean claim and the newer
``role: "admin"`` claim so tokens issued before the role migration still work.
"""
decoded = await require_firebase_token(credentials)
if get_role(decoded) != "admin":
raise HTTPException(status_code=403, detail="Admin access required")
return decoded
async def require_service_key(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Accept only the internal service key — used for bot-only endpoints."""
if not credentials:
raise HTTPException(status_code=401, detail="Missing authorization token")
if not settings.service_key:
raise HTTPException(status_code=503, detail="Service key not configured")
if not secrets.compare_digest(credentials.credentials, settings.service_key):
raise HTTPException(status_code=403, detail="Service key required")
return {"service": True}
async def require_service_key_or_admin(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Accept either the internal service key or a Firebase admin token.
Used for endpoints that the Discord bot (service key) and dashboard admins
(Firebase + admin claim) both need to call, but regular Firebase users must not.
"""
if not credentials:
raise HTTPException(status_code=401, detail="Missing authorization token")
token = credentials.credentials
if settings.service_key and secrets.compare_digest(token, settings.service_key):
return {"service": True}
try:
decoded = firebase_auth.verify_id_token(token)
except Exception:
raise HTTPException(status_code=401, detail="Invalid or expired token")
if get_role(decoded) != "admin":
raise HTTPException(status_code=403, detail="Admin access required")
return decoded
# ---------------------------------------------------------------------------
# Automation / agent principal
# ---------------------------------------------------------------------------
# Identity written into audit_log when the agent key is what authenticated a
# request. A Firebase admin gets their own uid/email instead, so the two are
# always distinguishable after the fact — which is the point.
AGENT_PRINCIPAL_UID = "agent-service"
AGENT_PRINCIPAL_EMAIL = "agent-service@drb.internal"
async def require_agent_key_or_admin(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Accept either the agent service key or a Firebase admin token.
Deliberately does NOT accept ``settings.service_key``. That key belongs to
the Discord bot, and honouring it here would collapse two principals into
one unattributable identity in every log line and audit entry — the exact
thing server-26#64 exists to end. The bot has no business flipping
platform-wide AI flags either way.
Exists so the unattended runbook can flip AI flags over HTTP instead of
SSHing into the container and writing ``config/ai_features`` with the admin
SDK, which needs a full container shell to move a cost switch.
The ``settings.agent_service_key and ...`` guard is load-bearing, not
stylistic: ``secrets.compare_digest("", "")`` is a MATCH, so any form of
``compare_digest(token, settings.agent_service_key or "")`` would turn a
deployment that never configured the key into one that accepts an empty
credential. Check the key is configured first and never substitute a
placeholder. (``require_service_key`` states the same intent by raising
503 when unset; both are correct, this one just stays open to admins.)
"""
if not credentials:
raise HTTPException(status_code=401, detail="Missing authorization token")
token = credentials.credentials
if settings.agent_service_key and secrets.compare_digest(token, settings.agent_service_key):
return {
"service": True,
"principal": "agent",
"uid": AGENT_PRINCIPAL_UID,
"email": AGENT_PRINCIPAL_EMAIL,
}
try:
decoded = firebase_auth.verify_id_token(token)
except Exception:
raise HTTPException(status_code=401, detail="Invalid or expired token")
if get_role(decoded) != "admin":
raise HTTPException(status_code=403, detail="Admin access required")
return decoded
def describe_actor(principal: dict) -> tuple[str, str]:
"""Return ``(actor_uid, actor_email)`` for an audit entry.
Works for any credential shape the dependencies above produce, so an audit
call site never has to switch on principal type itself.
"""
if principal.get("principal") == "agent":
return AGENT_PRINCIPAL_UID, AGENT_PRINCIPAL_EMAIL
if principal.get("service"):
return "service", "service@drb.internal"
if principal.get("node"):
node_id = principal.get("node_id") or "unknown"
return f"node:{node_id}", ""
return principal.get("uid") or "unknown", principal.get("email") or ""
# ---------------------------------------------------------------------------
# Simple in-memory sliding-window rate limiter
# ---------------------------------------------------------------------------
# Not persistent across restarts; good enough for a single-instance deployment.
# Key format is caller-defined (e.g. "{uid}:{endpoint}").
class _RateLimiter:
def __init__(self, max_calls: int, window_seconds: int):
self.max_calls = max_calls
self.window = window_seconds
self._log: dict[str, deque] = defaultdict(deque)
def check(self, key: str) -> None:
now = time.monotonic()
q = self._log[key]
while q and now - q[0] > self.window:
q.popleft()
if len(q) >= self.max_calls:
raise HTTPException(
status_code=429,
detail="Rate limit exceeded. Please wait before trying again.",
)
q.append(now)
# Shared limiter instances
# trip chat: 20 requests per user per 5 minutes
trip_chat_limiter = _RateLimiter(max_calls=20, window_seconds=300)
# per-incident summarize: 5 per incident per 10 minutes
summarize_limiter = _RateLimiter(max_calls=5, window_seconds=600)
# vocabulary bootstrap: 2 per system per hour
bootstrap_limiter = _RateLimiter(max_calls=2, window_seconds=3600)
# per-call reprocess: 3 per call per 10 minutes — reprocess re-runs the full
# Whisper + Gemini pipeline, which is real spend per call; this is now also
# admin-only (see routers/calls.py) but the limiter stays as a second guard
# against a compromised/careless admin session looping it. Keyed by call_id,
# same pattern as summarize_limiter.
reprocess_limiter = _RateLimiter(max_calls=3, window_seconds=600)
# public waitlist submissions: 5 per source IP per hour — POST /waitlist has
# no auth at all by design (SAAS_PLAN.md B6), so this is the only thing
# standing between it and being spammed.
waitlist_limiter = _RateLimiter(max_calls=5, window_seconds=3600)