Close the /admin/features side-door that needed a container shell to flip AI spend
Board minutes #62 Decision 2 (server-26#64), due 2026-08-31. CTO draft #60 finding 1 and CISO draft #61 finding 3 reached this independently. GET/PUT /admin/features accepted only a Firebase admin token, so the unattended runbook had no headless path and SSHed into the c2-core container to write config/ai_features with the admin SDK. Moving a platform-wide AI cost switch required a full container shell, and set_flags() wrote no audit entry either way, so a flag flip was unattributable however it happened. - New agent_service_key (AGENT_SERVICE_KEY), deliberately separate from the Discord bot's service_key. Sharing one key would collapse two principals into a single unattributable identity in every log line, and the bot has no business flipping AI flags regardless. - require_agent_key_or_admin accepts the agent key or a Firebase admin, and rejects the Discord key. The "key is configured" guard is load-bearing: compare_digest("", "") is a match, so a deployment that never set the key would otherwise accept an empty credential. - set_flags() writes an audit_log entry with before/after values and the actor, wrapped so an audit failure cannot lose the flag write or 500 the route. - Cascade helper sets the global doc and every system carrying an ai_flags override in one call. A global False already beats everything, but a system False beats a global True, so turning AI *on* could half-apply and leave a radio system hot after shutoff. It scans for the override rather than hardcoding the two known system IDs, so a new system cannot silently defeat it. - cascade defaults to False. PUT /systems/{id}/ai-flags and the AiFlagsPanel toggle mean a per-system override is deliberate operator intent; cascading by default would erase it on any unrelated global flip. The runbook opts in. Issue items 5 and 6 (retiring the SSH path from drb-worksession.md) are NOT done here and the runbook is untouched. The credential does not exist in production yet, so the SSH path is still the only one that works; retiring it now would break the next unattended run. Owner activation is recorded on #64. Tests 273 -> 289. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
0635de8dac
commit
865b5b4317
@@ -1,7 +1,7 @@
|
||||
import asyncio
|
||||
from datetime import datetime, timezone, timedelta
|
||||
from fastapi import APIRouter, Depends, Query
|
||||
from app.internal.auth import require_admin_token
|
||||
from app.internal.auth import require_admin_token, require_agent_key_or_admin, describe_actor
|
||||
from app.internal.feature_flags import get_flags, set_flags
|
||||
from app.internal import firestore as fstore
|
||||
from app.config import settings
|
||||
@@ -25,20 +25,54 @@ router = APIRouter(prefix="/admin", tags=["admin"])
|
||||
|
||||
|
||||
@router.get("/features")
|
||||
async def get_feature_flags(_=Depends(require_admin_token)):
|
||||
async def get_feature_flags(_=Depends(require_agent_key_or_admin)):
|
||||
"""
|
||||
Return the current AI feature flag state. Admin-only (SAAS_PLAN.md B2c) —
|
||||
was previously any authenticated user via require_firebase_token, which
|
||||
handed platform-wide AI configuration state to every signed-in viewer
|
||||
regardless of org.
|
||||
|
||||
Also reachable with the agent service key (server-26#64) so the unattended
|
||||
runbook can read the switch over HTTP instead of shelling into the
|
||||
container. Note this is require_agent_key_or_admin, NOT the Discord bot's
|
||||
service key — see internal/auth.py.
|
||||
"""
|
||||
return await get_flags()
|
||||
|
||||
|
||||
@router.put("/features")
|
||||
async def update_feature_flags(body: dict, _=Depends(require_admin_token)):
|
||||
"""Update one or more AI feature flags. Admin only."""
|
||||
return await set_flags(body)
|
||||
async def update_feature_flags(
|
||||
body: dict,
|
||||
cascade: bool = Query(
|
||||
False,
|
||||
description=(
|
||||
"Also clear per-system ai_flags overrides for the keys being set, "
|
||||
"so the flip applies to every radio system."
|
||||
),
|
||||
),
|
||||
principal: dict = Depends(require_agent_key_or_admin),
|
||||
):
|
||||
"""Update one or more AI feature flags. Admin or agent service key.
|
||||
|
||||
``cascade`` defaults to **False**, deliberately.
|
||||
|
||||
The tempting default is True: feature_flags.resolve_flags lets a
|
||||
system-level False beat a global True, so turning AI back ON globally can
|
||||
half-apply and leave a system dark, and cascade-by-default would make every
|
||||
flip total. That reasoning holds only if per-system ai_flags are set
|
||||
exclusively by hand. They are not — PUT /systems/{system_id}/ai-flags
|
||||
(routers/systems.py) is a real admin route and drb-frontend's AiFlagsPanel
|
||||
(app/systems/page.tsx) is a real toggle in the UI. So an override is a
|
||||
deliberate operator decision that is visible in the interface, and
|
||||
cascading by default would silently erase it on the next unrelated global
|
||||
flip, with the operator's own UI still showing what they set until reload.
|
||||
|
||||
Silently destroying operator intent is the worse failure, so the caller
|
||||
says when it means "everywhere": the runbook passes cascade=true on the
|
||||
shutoff, and the admin UI (which does not pass it) keeps its per-system
|
||||
overrides.
|
||||
"""
|
||||
return await set_flags(body, actor=describe_actor(principal), cascade=cascade)
|
||||
|
||||
|
||||
@router.get("/debug/correlation")
|
||||
|
||||
Reference in New Issue
Block a user