Close the /admin/features side-door that needed a container shell to flip AI spend
Build & Deploy / Build & push images (push) Successful in 4m15s
Build & Deploy / Deploy to VM (push) Successful in 1m56s
Build & Deploy / Report a failed deploy (push) Skipped

Board minutes #62 Decision 2 (server-26#64), due 2026-08-31. CTO draft #60
finding 1 and CISO draft #61 finding 3 reached this independently.

GET/PUT /admin/features accepted only a Firebase admin token, so the unattended
runbook had no headless path and SSHed into the c2-core container to write
config/ai_features with the admin SDK. Moving a platform-wide AI cost switch
required a full container shell, and set_flags() wrote no audit entry either
way, so a flag flip was unattributable however it happened.

- New agent_service_key (AGENT_SERVICE_KEY), deliberately separate from the
  Discord bot's service_key. Sharing one key would collapse two principals into
  a single unattributable identity in every log line, and the bot has no
  business flipping AI flags regardless.
- require_agent_key_or_admin accepts the agent key or a Firebase admin, and
  rejects the Discord key. The "key is configured" guard is load-bearing:
  compare_digest("", "") is a match, so a deployment that never set the key
  would otherwise accept an empty credential.
- set_flags() writes an audit_log entry with before/after values and the actor,
  wrapped so an audit failure cannot lose the flag write or 500 the route.
- Cascade helper sets the global doc and every system carrying an ai_flags
  override in one call. A global False already beats everything, but a system
  False beats a global True, so turning AI *on* could half-apply and leave a
  radio system hot after shutoff. It scans for the override rather than
  hardcoding the two known system IDs, so a new system cannot silently defeat
  it.
- cascade defaults to False. PUT /systems/{id}/ai-flags and the AiFlagsPanel
  toggle mean a per-system override is deliberate operator intent; cascading by
  default would erase it on any unrelated global flip. The runbook opts in.

Issue items 5 and 6 (retiring the SSH path from drb-worksession.md) are NOT
done here and the runbook is untouched. The credential does not exist in
production yet, so the SSH path is still the only one that works; retiring it
now would break the next unattended run. Owner activation is recorded on #64.

Tests 273 -> 289.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Logan Cusano
2026-08-30 02:52:00 -04:00
co-authored by Claude Opus 5
parent 0635de8dac
commit 865b5b4317
9 changed files with 574 additions and 10 deletions
+15
View File
@@ -141,6 +141,21 @@ class Settings(BaseSettings):
# Internal service key — allows server-side services (discord bot) to call C2 without Firebase
service_key: Optional[str] = None
# Automation/agent service key — the unattended work-session agent's own
# credential for the headless routes it needs (currently GET/PUT
# /admin/features).
#
# DELIBERATELY SEPARATE from service_key above, not a second consumer of
# it. service_key is the Discord bot's, and it is handed to a process that
# relays radio traffic to a chat server; sharing it here would make "the
# bot" and "the agent" the same principal in every log line and audit
# entry, so a global AI-cost flag flip could never be attributed to whoever
# actually made it. Two keys, two identities (server-26#64 item 1).
#
# Unset means the agent path is simply closed — the routes still accept a
# Firebase admin token. Generate with: openssl rand -hex 32
agent_service_key: Optional[str] = None
# Fleet-wide token edge nodes present to POST /nodes/enroll on first boot.
# Not a per-node secret — see routers/enrollment.py for why a leaked copy
# of this alone can't steal an already-approved node's key.
+68
View File
@@ -220,6 +220,74 @@ async def require_service_key_or_admin(
return decoded
# ---------------------------------------------------------------------------
# Automation / agent principal
# ---------------------------------------------------------------------------
# Identity written into audit_log when the agent key is what authenticated a
# request. A Firebase admin gets their own uid/email instead, so the two are
# always distinguishable after the fact — which is the point.
AGENT_PRINCIPAL_UID = "agent-service"
AGENT_PRINCIPAL_EMAIL = "agent-service@drb.internal"
async def require_agent_key_or_admin(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Accept either the agent service key or a Firebase admin token.
Deliberately does NOT accept ``settings.service_key``. That key belongs to
the Discord bot, and honouring it here would collapse two principals into
one unattributable identity in every log line and audit entry — the exact
thing server-26#64 exists to end. The bot has no business flipping
platform-wide AI flags either way.
Exists so the unattended runbook can flip AI flags over HTTP instead of
SSHing into the container and writing ``config/ai_features`` with the admin
SDK, which needs a full container shell to move a cost switch.
The ``settings.agent_service_key and ...`` guard is load-bearing, not
stylistic: ``secrets.compare_digest("", "")`` is a MATCH, so any form of
``compare_digest(token, settings.agent_service_key or "")`` would turn a
deployment that never configured the key into one that accepts an empty
credential. Check the key is configured first and never substitute a
placeholder. (``require_service_key`` states the same intent by raising
503 when unset; both are correct, this one just stays open to admins.)
"""
if not credentials:
raise HTTPException(status_code=401, detail="Missing authorization token")
token = credentials.credentials
if settings.agent_service_key and secrets.compare_digest(token, settings.agent_service_key):
return {
"service": True,
"principal": "agent",
"uid": AGENT_PRINCIPAL_UID,
"email": AGENT_PRINCIPAL_EMAIL,
}
try:
decoded = firebase_auth.verify_id_token(token)
except Exception:
raise HTTPException(status_code=401, detail="Invalid or expired token")
if get_role(decoded) != "admin":
raise HTTPException(status_code=403, detail="Admin access required")
return decoded
def describe_actor(principal: dict) -> tuple[str, str]:
"""Return ``(actor_uid, actor_email)`` for an audit entry.
Works for any credential shape the dependencies above produce, so an audit
call site never has to switch on principal type itself.
"""
if principal.get("principal") == "agent":
return AGENT_PRINCIPAL_UID, AGENT_PRINCIPAL_EMAIL
if principal.get("service"):
return "service", "service@drb.internal"
if principal.get("node"):
node_id = principal.get("node_id") or "unknown"
return f"node:{node_id}", ""
return principal.get("uid") or "unknown", principal.get("email") or ""
# ---------------------------------------------------------------------------
# Simple in-memory sliding-window rate limiter
# ---------------------------------------------------------------------------
+123 -4
View File
@@ -63,18 +63,137 @@ async def get_flags() -> dict[str, bool]:
return dict(_cache)
async def set_flags(updates: dict[str, bool]) -> dict[str, bool]:
"""Write flag updates to Firestore and invalidate the cache."""
global _cache, _cache_ts
async def _cascade_to_systems(clean: dict[str, bool]) -> tuple[list[dict], list[dict]]:
"""Clear per-system ``ai_flags`` overrides for the keys just set globally.
Returns ``(changes, errors)``.
Why clearing rather than overwriting with the new value: an override that
stays present, merely agreeing with the global switch for now, defeats the
NEXT flip exactly the same way. Removing it makes the system inherit, which
is the same semantics the human-facing route already offers
(``PUT /systems/{id}/ai-flags`` with null → "clear override, inherit
global").
Systems are discovered by scanning for documents that actually carry an
``ai_flags`` map — never a hardcoded id list. Two systems carry overrides
today; a third added tomorrow would silently defeat a global shutoff if
this were pinned to the current pair.
"""
changes: list[dict] = []
errors: list[dict] = []
systems = await fstore.collection_list("systems")
for system in systems:
sid = system.get("system_id")
ai_flags = system.get("ai_flags")
# Only documents that actually carry the map. A system with no
# overrides already inherits, so there is nothing to cascade to.
if not sid or not isinstance(ai_flags, dict) or not ai_flags:
continue
removed = {k: ai_flags[k] for k in clean if k in ai_flags}
if not removed:
continue
remaining = {k: v for k, v in ai_flags.items() if k not in clean}
try:
await fstore.doc_update("systems", sid, {"ai_flags": remaining})
except Exception as e:
# Report rather than swallow: a half-applied cascade is the exact
# failure mode this helper exists to prevent, so it must be visible
# in the log and the audit entry.
logger.error(f"Feature flags: cascade to system '{sid}' failed ({e})")
errors.append({"system_id": sid, "error": str(e)})
continue
changes.append({
"system_id": sid,
"cleared_overrides": removed,
"now_inherits": {k: clean[k] for k in removed},
})
return changes, errors
async def set_flags(
updates: dict[str, bool],
actor: tuple[str, str] | None = None,
cascade: bool = False,
) -> dict[str, bool]:
"""Write flag updates to Firestore, invalidate the cache, and audit it.
``actor`` is ``(actor_uid, actor_email)`` — see auth.describe_actor. It is
optional so existing callers keep working; an unattributed flip is logged
as "unknown" rather than not logged at all.
``cascade`` also clears the matching per-system ``ai_flags`` overrides, so
one call is a total flip. Defaults to False deliberately — see the route's
comment in routers/admin.py.
Returns the resulting global flags dict, unchanged in shape: the admin UI
(drb-frontend/lib/c2api.ts setFeatureFlags) types the response as
Record<string, boolean>, so cascade/audit detail goes to the log and the
audit entry rather than into this payload.
"""
global _cache_ts
clean = {k: bool(v) for k, v in updates.items() if k in _DEFAULTS}
if not clean:
raise ValueError(f"No recognised flag keys in update: {list(updates)}")
# Force a fresh read for the "before" side of the audit entry: the TTL
# cache can be up to _TTL seconds stale, and a wrong previous value in an
# audit log is worse than none.
_cache_ts = 0.0
before = await get_flags()
await fstore.doc_set(_COLLECTION, _DOC_ID, clean)
_cache_ts = 0.0 # force re-read on next get_flags()
logger.info(f"Feature flags updated: {clean}")
return await get_flags()
cascaded: list[dict] = []
cascade_errors: list[dict] = []
if cascade:
cascaded, cascade_errors = await _cascade_to_systems(clean)
logger.info(
f"Feature flags: cascaded {list(clean)} to {len(cascaded)} system(s), "
f"{len(cascade_errors)} error(s)"
)
after = await get_flags()
# The audit entry is a record OF the write, never a precondition for it.
# audit_log lives in the same Firestore that just accepted the flag write,
# so a failure here is nearly always transient — losing the flip (or 500ing
# a route that already succeeded, which invites a retry that flips it back)
# would be a far worse outcome than an unrecorded flip that is still in the
# service log above.
try:
# Deferred import: app.internal.audit pulls in firestore, and this
# module is imported from router module scope.
from app.internal import audit
actor_uid, actor_email = actor or ("unknown", "")
changed = {
k: {"from": before.get(k), "to": after.get(k)}
for k in clean
if before.get(k) != after.get(k)
}
await audit.write_audit(
actor_uid=actor_uid,
actor_email=actor_email,
action="feature_flags.update",
details={
"requested": clean,
"changed": changed,
"before": before,
"after": after,
"cascade": cascade,
"cascaded_systems": cascaded,
"cascade_errors": cascade_errors,
},
)
except Exception as e:
logger.error(f"Feature flags: audit write failed ({e}) — flag change stands")
return after
async def resolve_flags(system_id: str | None):
+39 -5
View File
@@ -1,7 +1,7 @@
import asyncio
from datetime import datetime, timezone, timedelta
from fastapi import APIRouter, Depends, Query
from app.internal.auth import require_admin_token
from app.internal.auth import require_admin_token, require_agent_key_or_admin, describe_actor
from app.internal.feature_flags import get_flags, set_flags
from app.internal import firestore as fstore
from app.config import settings
@@ -25,20 +25,54 @@ router = APIRouter(prefix="/admin", tags=["admin"])
@router.get("/features")
async def get_feature_flags(_=Depends(require_admin_token)):
async def get_feature_flags(_=Depends(require_agent_key_or_admin)):
"""
Return the current AI feature flag state. Admin-only (SAAS_PLAN.md B2c) —
was previously any authenticated user via require_firebase_token, which
handed platform-wide AI configuration state to every signed-in viewer
regardless of org.
Also reachable with the agent service key (server-26#64) so the unattended
runbook can read the switch over HTTP instead of shelling into the
container. Note this is require_agent_key_or_admin, NOT the Discord bot's
service key — see internal/auth.py.
"""
return await get_flags()
@router.put("/features")
async def update_feature_flags(body: dict, _=Depends(require_admin_token)):
"""Update one or more AI feature flags. Admin only."""
return await set_flags(body)
async def update_feature_flags(
body: dict,
cascade: bool = Query(
False,
description=(
"Also clear per-system ai_flags overrides for the keys being set, "
"so the flip applies to every radio system."
),
),
principal: dict = Depends(require_agent_key_or_admin),
):
"""Update one or more AI feature flags. Admin or agent service key.
``cascade`` defaults to **False**, deliberately.
The tempting default is True: feature_flags.resolve_flags lets a
system-level False beat a global True, so turning AI back ON globally can
half-apply and leave a system dark, and cascade-by-default would make every
flip total. That reasoning holds only if per-system ai_flags are set
exclusively by hand. They are not — PUT /systems/{system_id}/ai-flags
(routers/systems.py) is a real admin route and drb-frontend's AiFlagsPanel
(app/systems/page.tsx) is a real toggle in the UI. So an override is a
deliberate operator decision that is visible in the interface, and
cascading by default would silently erase it on the next unrelated global
flip, with the operator's own UI still showing what they set until reload.
Silently destroying operator intent is the worse failure, so the caller
says when it means "everywhere": the runbook passes cascade=true on the
shutoff, and the admin UI (which does not pass it) keeps its per-system
overrides.
"""
return await set_flags(body, actor=describe_actor(principal), cascade=cascade)
@router.get("/debug/correlation")