Close the /admin/features side-door that needed a container shell to flip AI spend
Board minutes #62 Decision 2 (server-26#64), due 2026-08-31. CTO draft #60 finding 1 and CISO draft #61 finding 3 reached this independently. GET/PUT /admin/features accepted only a Firebase admin token, so the unattended runbook had no headless path and SSHed into the c2-core container to write config/ai_features with the admin SDK. Moving a platform-wide AI cost switch required a full container shell, and set_flags() wrote no audit entry either way, so a flag flip was unattributable however it happened. - New agent_service_key (AGENT_SERVICE_KEY), deliberately separate from the Discord bot's service_key. Sharing one key would collapse two principals into a single unattributable identity in every log line, and the bot has no business flipping AI flags regardless. - require_agent_key_or_admin accepts the agent key or a Firebase admin, and rejects the Discord key. The "key is configured" guard is load-bearing: compare_digest("", "") is a match, so a deployment that never set the key would otherwise accept an empty credential. - set_flags() writes an audit_log entry with before/after values and the actor, wrapped so an audit failure cannot lose the flag write or 500 the route. - Cascade helper sets the global doc and every system carrying an ai_flags override in one call. A global False already beats everything, but a system False beats a global True, so turning AI *on* could half-apply and leave a radio system hot after shutoff. It scans for the override rather than hardcoding the two known system IDs, so a new system cannot silently defeat it. - cascade defaults to False. PUT /systems/{id}/ai-flags and the AiFlagsPanel toggle mean a per-system override is deliberate operator intent; cascading by default would erase it on any unrelated global flip. The runbook opts in. Issue items 5 and 6 (retiring the SSH path from drb-worksession.md) are NOT done here and the runbook is untouched. The credential does not exist in production yet, so the SSH path is still the only one that works; retiring it now would break the next unattended run. Owner activation is recorded on #64. Tests 273 -> 289. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
0635de8dac
commit
865b5b4317
@@ -220,6 +220,74 @@ async def require_service_key_or_admin(
|
||||
return decoded
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Automation / agent principal
|
||||
# ---------------------------------------------------------------------------
|
||||
# Identity written into audit_log when the agent key is what authenticated a
|
||||
# request. A Firebase admin gets their own uid/email instead, so the two are
|
||||
# always distinguishable after the fact — which is the point.
|
||||
AGENT_PRINCIPAL_UID = "agent-service"
|
||||
AGENT_PRINCIPAL_EMAIL = "agent-service@drb.internal"
|
||||
|
||||
|
||||
async def require_agent_key_or_admin(
|
||||
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
|
||||
) -> dict:
|
||||
"""Accept either the agent service key or a Firebase admin token.
|
||||
|
||||
Deliberately does NOT accept ``settings.service_key``. That key belongs to
|
||||
the Discord bot, and honouring it here would collapse two principals into
|
||||
one unattributable identity in every log line and audit entry — the exact
|
||||
thing server-26#64 exists to end. The bot has no business flipping
|
||||
platform-wide AI flags either way.
|
||||
|
||||
Exists so the unattended runbook can flip AI flags over HTTP instead of
|
||||
SSHing into the container and writing ``config/ai_features`` with the admin
|
||||
SDK, which needs a full container shell to move a cost switch.
|
||||
|
||||
The ``settings.agent_service_key and ...`` guard is load-bearing, not
|
||||
stylistic: ``secrets.compare_digest("", "")`` is a MATCH, so any form of
|
||||
``compare_digest(token, settings.agent_service_key or "")`` would turn a
|
||||
deployment that never configured the key into one that accepts an empty
|
||||
credential. Check the key is configured first and never substitute a
|
||||
placeholder. (``require_service_key`` states the same intent by raising
|
||||
503 when unset; both are correct, this one just stays open to admins.)
|
||||
"""
|
||||
if not credentials:
|
||||
raise HTTPException(status_code=401, detail="Missing authorization token")
|
||||
token = credentials.credentials
|
||||
if settings.agent_service_key and secrets.compare_digest(token, settings.agent_service_key):
|
||||
return {
|
||||
"service": True,
|
||||
"principal": "agent",
|
||||
"uid": AGENT_PRINCIPAL_UID,
|
||||
"email": AGENT_PRINCIPAL_EMAIL,
|
||||
}
|
||||
try:
|
||||
decoded = firebase_auth.verify_id_token(token)
|
||||
except Exception:
|
||||
raise HTTPException(status_code=401, detail="Invalid or expired token")
|
||||
if get_role(decoded) != "admin":
|
||||
raise HTTPException(status_code=403, detail="Admin access required")
|
||||
return decoded
|
||||
|
||||
|
||||
def describe_actor(principal: dict) -> tuple[str, str]:
|
||||
"""Return ``(actor_uid, actor_email)`` for an audit entry.
|
||||
|
||||
Works for any credential shape the dependencies above produce, so an audit
|
||||
call site never has to switch on principal type itself.
|
||||
"""
|
||||
if principal.get("principal") == "agent":
|
||||
return AGENT_PRINCIPAL_UID, AGENT_PRINCIPAL_EMAIL
|
||||
if principal.get("service"):
|
||||
return "service", "service@drb.internal"
|
||||
if principal.get("node"):
|
||||
node_id = principal.get("node_id") or "unknown"
|
||||
return f"node:{node_id}", ""
|
||||
return principal.get("uid") or "unknown", principal.get("email") or ""
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Simple in-memory sliding-window rate limiter
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@@ -63,18 +63,137 @@ async def get_flags() -> dict[str, bool]:
|
||||
return dict(_cache)
|
||||
|
||||
|
||||
async def set_flags(updates: dict[str, bool]) -> dict[str, bool]:
|
||||
"""Write flag updates to Firestore and invalidate the cache."""
|
||||
global _cache, _cache_ts
|
||||
async def _cascade_to_systems(clean: dict[str, bool]) -> tuple[list[dict], list[dict]]:
|
||||
"""Clear per-system ``ai_flags`` overrides for the keys just set globally.
|
||||
|
||||
Returns ``(changes, errors)``.
|
||||
|
||||
Why clearing rather than overwriting with the new value: an override that
|
||||
stays present, merely agreeing with the global switch for now, defeats the
|
||||
NEXT flip exactly the same way. Removing it makes the system inherit, which
|
||||
is the same semantics the human-facing route already offers
|
||||
(``PUT /systems/{id}/ai-flags`` with null → "clear override, inherit
|
||||
global").
|
||||
|
||||
Systems are discovered by scanning for documents that actually carry an
|
||||
``ai_flags`` map — never a hardcoded id list. Two systems carry overrides
|
||||
today; a third added tomorrow would silently defeat a global shutoff if
|
||||
this were pinned to the current pair.
|
||||
"""
|
||||
changes: list[dict] = []
|
||||
errors: list[dict] = []
|
||||
|
||||
systems = await fstore.collection_list("systems")
|
||||
for system in systems:
|
||||
sid = system.get("system_id")
|
||||
ai_flags = system.get("ai_flags")
|
||||
# Only documents that actually carry the map. A system with no
|
||||
# overrides already inherits, so there is nothing to cascade to.
|
||||
if not sid or not isinstance(ai_flags, dict) or not ai_flags:
|
||||
continue
|
||||
removed = {k: ai_flags[k] for k in clean if k in ai_flags}
|
||||
if not removed:
|
||||
continue
|
||||
remaining = {k: v for k, v in ai_flags.items() if k not in clean}
|
||||
try:
|
||||
await fstore.doc_update("systems", sid, {"ai_flags": remaining})
|
||||
except Exception as e:
|
||||
# Report rather than swallow: a half-applied cascade is the exact
|
||||
# failure mode this helper exists to prevent, so it must be visible
|
||||
# in the log and the audit entry.
|
||||
logger.error(f"Feature flags: cascade to system '{sid}' failed ({e})")
|
||||
errors.append({"system_id": sid, "error": str(e)})
|
||||
continue
|
||||
changes.append({
|
||||
"system_id": sid,
|
||||
"cleared_overrides": removed,
|
||||
"now_inherits": {k: clean[k] for k in removed},
|
||||
})
|
||||
|
||||
return changes, errors
|
||||
|
||||
|
||||
async def set_flags(
|
||||
updates: dict[str, bool],
|
||||
actor: tuple[str, str] | None = None,
|
||||
cascade: bool = False,
|
||||
) -> dict[str, bool]:
|
||||
"""Write flag updates to Firestore, invalidate the cache, and audit it.
|
||||
|
||||
``actor`` is ``(actor_uid, actor_email)`` — see auth.describe_actor. It is
|
||||
optional so existing callers keep working; an unattributed flip is logged
|
||||
as "unknown" rather than not logged at all.
|
||||
|
||||
``cascade`` also clears the matching per-system ``ai_flags`` overrides, so
|
||||
one call is a total flip. Defaults to False deliberately — see the route's
|
||||
comment in routers/admin.py.
|
||||
|
||||
Returns the resulting global flags dict, unchanged in shape: the admin UI
|
||||
(drb-frontend/lib/c2api.ts setFeatureFlags) types the response as
|
||||
Record<string, boolean>, so cascade/audit detail goes to the log and the
|
||||
audit entry rather than into this payload.
|
||||
"""
|
||||
global _cache_ts
|
||||
|
||||
clean = {k: bool(v) for k, v in updates.items() if k in _DEFAULTS}
|
||||
if not clean:
|
||||
raise ValueError(f"No recognised flag keys in update: {list(updates)}")
|
||||
|
||||
# Force a fresh read for the "before" side of the audit entry: the TTL
|
||||
# cache can be up to _TTL seconds stale, and a wrong previous value in an
|
||||
# audit log is worse than none.
|
||||
_cache_ts = 0.0
|
||||
before = await get_flags()
|
||||
|
||||
await fstore.doc_set(_COLLECTION, _DOC_ID, clean)
|
||||
_cache_ts = 0.0 # force re-read on next get_flags()
|
||||
logger.info(f"Feature flags updated: {clean}")
|
||||
return await get_flags()
|
||||
|
||||
cascaded: list[dict] = []
|
||||
cascade_errors: list[dict] = []
|
||||
if cascade:
|
||||
cascaded, cascade_errors = await _cascade_to_systems(clean)
|
||||
logger.info(
|
||||
f"Feature flags: cascaded {list(clean)} to {len(cascaded)} system(s), "
|
||||
f"{len(cascade_errors)} error(s)"
|
||||
)
|
||||
|
||||
after = await get_flags()
|
||||
|
||||
# The audit entry is a record OF the write, never a precondition for it.
|
||||
# audit_log lives in the same Firestore that just accepted the flag write,
|
||||
# so a failure here is nearly always transient — losing the flip (or 500ing
|
||||
# a route that already succeeded, which invites a retry that flips it back)
|
||||
# would be a far worse outcome than an unrecorded flip that is still in the
|
||||
# service log above.
|
||||
try:
|
||||
# Deferred import: app.internal.audit pulls in firestore, and this
|
||||
# module is imported from router module scope.
|
||||
from app.internal import audit
|
||||
actor_uid, actor_email = actor or ("unknown", "")
|
||||
changed = {
|
||||
k: {"from": before.get(k), "to": after.get(k)}
|
||||
for k in clean
|
||||
if before.get(k) != after.get(k)
|
||||
}
|
||||
await audit.write_audit(
|
||||
actor_uid=actor_uid,
|
||||
actor_email=actor_email,
|
||||
action="feature_flags.update",
|
||||
details={
|
||||
"requested": clean,
|
||||
"changed": changed,
|
||||
"before": before,
|
||||
"after": after,
|
||||
"cascade": cascade,
|
||||
"cascaded_systems": cascaded,
|
||||
"cascade_errors": cascade_errors,
|
||||
},
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Feature flags: audit write failed ({e}) — flag change stands")
|
||||
|
||||
return after
|
||||
|
||||
|
||||
async def resolve_flags(system_id: str | None):
|
||||
|
||||
Reference in New Issue
Block a user