Close the /admin/features side-door that needed a container shell to flip AI spend
Build & Deploy / Build & push images (push) Successful in 4m15s
Build & Deploy / Deploy to VM (push) Successful in 1m56s
Build & Deploy / Report a failed deploy (push) Skipped

Board minutes #62 Decision 2 (server-26#64), due 2026-08-31. CTO draft #60
finding 1 and CISO draft #61 finding 3 reached this independently.

GET/PUT /admin/features accepted only a Firebase admin token, so the unattended
runbook had no headless path and SSHed into the c2-core container to write
config/ai_features with the admin SDK. Moving a platform-wide AI cost switch
required a full container shell, and set_flags() wrote no audit entry either
way, so a flag flip was unattributable however it happened.

- New agent_service_key (AGENT_SERVICE_KEY), deliberately separate from the
  Discord bot's service_key. Sharing one key would collapse two principals into
  a single unattributable identity in every log line, and the bot has no
  business flipping AI flags regardless.
- require_agent_key_or_admin accepts the agent key or a Firebase admin, and
  rejects the Discord key. The "key is configured" guard is load-bearing:
  compare_digest("", "") is a match, so a deployment that never set the key
  would otherwise accept an empty credential.
- set_flags() writes an audit_log entry with before/after values and the actor,
  wrapped so an audit failure cannot lose the flag write or 500 the route.
- Cascade helper sets the global doc and every system carrying an ai_flags
  override in one call. A global False already beats everything, but a system
  False beats a global True, so turning AI *on* could half-apply and leave a
  radio system hot after shutoff. It scans for the override rather than
  hardcoding the two known system IDs, so a new system cannot silently defeat
  it.
- cascade defaults to False. PUT /systems/{id}/ai-flags and the AiFlagsPanel
  toggle mean a per-system override is deliberate operator intent; cascading by
  default would erase it on any unrelated global flip. The runbook opts in.

Issue items 5 and 6 (retiring the SSH path from drb-worksession.md) are NOT
done here and the runbook is untouched. The credential does not exist in
production yet, so the SSH path is still the only one that works; retiring it
now would break the next unattended run. Owner activation is recorded on #64.

Tests 273 -> 289.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Logan Cusano
2026-08-30 02:52:00 -04:00
co-authored by Claude Opus 5
parent 0635de8dac
commit 865b5b4317
9 changed files with 574 additions and 10 deletions
+8
View File
@@ -19,6 +19,14 @@ services:
- mosquitto_data:/mosquitto/data
- mosquitto_certs:/mosquitto/certs
# c2-core takes ALL of its configuration from ./drb-c2-core/.env — there is
# deliberately no `environment:` block here. An entry in that block wins over
# env_file, so listing a key here (e.g. AGENT_SERVICE_KEY=${AGENT_SERVICE_KEY})
# would let an unset top-level .env silently blank out a value the owner had
# correctly pasted into drb-c2-core/.env. New settings go in
# drb-c2-core/.env.example and, for the VM, in
# infra/ansible/roles/deploy/templates/c2-core.env.j2 + vault.yml.
# AGENT_SERVICE_KEY (server-26#64) is configured that way.
c2-core:
image: ${REGISTRY}/c2-core:${TAG:-latest}
build: ./drb-c2-core
+12
View File
@@ -37,3 +37,15 @@ EMBEDDING_SIMILARITY_THRESHOLD=0.82
# (POST /nodes/enroll). Shared across every node — NOT a per-node secret.
# Generate with: openssl rand -hex 32
ENROLLMENT_TOKEN=
# Shared key the Discord bot presents to reach C2 without Firebase.
# Generate with: openssl rand -hex 32
SERVICE_KEY=
# Agent/automation key for the unattended work session's headless routes
# (GET/PUT /admin/features). DELIBERATELY a different value from SERVICE_KEY —
# reusing the bot's key would make both principals indistinguishable in
# audit_log, which is the whole point of server-26#64. Leave blank to keep the
# agent path closed; the routes still take a Firebase admin token either way.
# Generate with: openssl rand -hex 32
AGENT_SERVICE_KEY=
+15
View File
@@ -141,6 +141,21 @@ class Settings(BaseSettings):
# Internal service key — allows server-side services (discord bot) to call C2 without Firebase
service_key: Optional[str] = None
# Automation/agent service key — the unattended work-session agent's own
# credential for the headless routes it needs (currently GET/PUT
# /admin/features).
#
# DELIBERATELY SEPARATE from service_key above, not a second consumer of
# it. service_key is the Discord bot's, and it is handed to a process that
# relays radio traffic to a chat server; sharing it here would make "the
# bot" and "the agent" the same principal in every log line and audit
# entry, so a global AI-cost flag flip could never be attributed to whoever
# actually made it. Two keys, two identities (server-26#64 item 1).
#
# Unset means the agent path is simply closed — the routes still accept a
# Firebase admin token. Generate with: openssl rand -hex 32
agent_service_key: Optional[str] = None
# Fleet-wide token edge nodes present to POST /nodes/enroll on first boot.
# Not a per-node secret — see routers/enrollment.py for why a leaked copy
# of this alone can't steal an already-approved node's key.
+68
View File
@@ -220,6 +220,74 @@ async def require_service_key_or_admin(
return decoded
# ---------------------------------------------------------------------------
# Automation / agent principal
# ---------------------------------------------------------------------------
# Identity written into audit_log when the agent key is what authenticated a
# request. A Firebase admin gets their own uid/email instead, so the two are
# always distinguishable after the fact — which is the point.
AGENT_PRINCIPAL_UID = "agent-service"
AGENT_PRINCIPAL_EMAIL = "agent-service@drb.internal"
async def require_agent_key_or_admin(
credentials: Optional[HTTPAuthorizationCredentials] = Security(_bearer),
) -> dict:
"""Accept either the agent service key or a Firebase admin token.
Deliberately does NOT accept ``settings.service_key``. That key belongs to
the Discord bot, and honouring it here would collapse two principals into
one unattributable identity in every log line and audit entry — the exact
thing server-26#64 exists to end. The bot has no business flipping
platform-wide AI flags either way.
Exists so the unattended runbook can flip AI flags over HTTP instead of
SSHing into the container and writing ``config/ai_features`` with the admin
SDK, which needs a full container shell to move a cost switch.
The ``settings.agent_service_key and ...`` guard is load-bearing, not
stylistic: ``secrets.compare_digest("", "")`` is a MATCH, so any form of
``compare_digest(token, settings.agent_service_key or "")`` would turn a
deployment that never configured the key into one that accepts an empty
credential. Check the key is configured first and never substitute a
placeholder. (``require_service_key`` states the same intent by raising
503 when unset; both are correct, this one just stays open to admins.)
"""
if not credentials:
raise HTTPException(status_code=401, detail="Missing authorization token")
token = credentials.credentials
if settings.agent_service_key and secrets.compare_digest(token, settings.agent_service_key):
return {
"service": True,
"principal": "agent",
"uid": AGENT_PRINCIPAL_UID,
"email": AGENT_PRINCIPAL_EMAIL,
}
try:
decoded = firebase_auth.verify_id_token(token)
except Exception:
raise HTTPException(status_code=401, detail="Invalid or expired token")
if get_role(decoded) != "admin":
raise HTTPException(status_code=403, detail="Admin access required")
return decoded
def describe_actor(principal: dict) -> tuple[str, str]:
"""Return ``(actor_uid, actor_email)`` for an audit entry.
Works for any credential shape the dependencies above produce, so an audit
call site never has to switch on principal type itself.
"""
if principal.get("principal") == "agent":
return AGENT_PRINCIPAL_UID, AGENT_PRINCIPAL_EMAIL
if principal.get("service"):
return "service", "service@drb.internal"
if principal.get("node"):
node_id = principal.get("node_id") or "unknown"
return f"node:{node_id}", ""
return principal.get("uid") or "unknown", principal.get("email") or ""
# ---------------------------------------------------------------------------
# Simple in-memory sliding-window rate limiter
# ---------------------------------------------------------------------------
+123 -4
View File
@@ -63,18 +63,137 @@ async def get_flags() -> dict[str, bool]:
return dict(_cache)
async def set_flags(updates: dict[str, bool]) -> dict[str, bool]:
"""Write flag updates to Firestore and invalidate the cache."""
global _cache, _cache_ts
async def _cascade_to_systems(clean: dict[str, bool]) -> tuple[list[dict], list[dict]]:
"""Clear per-system ``ai_flags`` overrides for the keys just set globally.
Returns ``(changes, errors)``.
Why clearing rather than overwriting with the new value: an override that
stays present, merely agreeing with the global switch for now, defeats the
NEXT flip exactly the same way. Removing it makes the system inherit, which
is the same semantics the human-facing route already offers
(``PUT /systems/{id}/ai-flags`` with null → "clear override, inherit
global").
Systems are discovered by scanning for documents that actually carry an
``ai_flags`` map — never a hardcoded id list. Two systems carry overrides
today; a third added tomorrow would silently defeat a global shutoff if
this were pinned to the current pair.
"""
changes: list[dict] = []
errors: list[dict] = []
systems = await fstore.collection_list("systems")
for system in systems:
sid = system.get("system_id")
ai_flags = system.get("ai_flags")
# Only documents that actually carry the map. A system with no
# overrides already inherits, so there is nothing to cascade to.
if not sid or not isinstance(ai_flags, dict) or not ai_flags:
continue
removed = {k: ai_flags[k] for k in clean if k in ai_flags}
if not removed:
continue
remaining = {k: v for k, v in ai_flags.items() if k not in clean}
try:
await fstore.doc_update("systems", sid, {"ai_flags": remaining})
except Exception as e:
# Report rather than swallow: a half-applied cascade is the exact
# failure mode this helper exists to prevent, so it must be visible
# in the log and the audit entry.
logger.error(f"Feature flags: cascade to system '{sid}' failed ({e})")
errors.append({"system_id": sid, "error": str(e)})
continue
changes.append({
"system_id": sid,
"cleared_overrides": removed,
"now_inherits": {k: clean[k] for k in removed},
})
return changes, errors
async def set_flags(
updates: dict[str, bool],
actor: tuple[str, str] | None = None,
cascade: bool = False,
) -> dict[str, bool]:
"""Write flag updates to Firestore, invalidate the cache, and audit it.
``actor`` is ``(actor_uid, actor_email)`` — see auth.describe_actor. It is
optional so existing callers keep working; an unattributed flip is logged
as "unknown" rather than not logged at all.
``cascade`` also clears the matching per-system ``ai_flags`` overrides, so
one call is a total flip. Defaults to False deliberately — see the route's
comment in routers/admin.py.
Returns the resulting global flags dict, unchanged in shape: the admin UI
(drb-frontend/lib/c2api.ts setFeatureFlags) types the response as
Record<string, boolean>, so cascade/audit detail goes to the log and the
audit entry rather than into this payload.
"""
global _cache_ts
clean = {k: bool(v) for k, v in updates.items() if k in _DEFAULTS}
if not clean:
raise ValueError(f"No recognised flag keys in update: {list(updates)}")
# Force a fresh read for the "before" side of the audit entry: the TTL
# cache can be up to _TTL seconds stale, and a wrong previous value in an
# audit log is worse than none.
_cache_ts = 0.0
before = await get_flags()
await fstore.doc_set(_COLLECTION, _DOC_ID, clean)
_cache_ts = 0.0 # force re-read on next get_flags()
logger.info(f"Feature flags updated: {clean}")
return await get_flags()
cascaded: list[dict] = []
cascade_errors: list[dict] = []
if cascade:
cascaded, cascade_errors = await _cascade_to_systems(clean)
logger.info(
f"Feature flags: cascaded {list(clean)} to {len(cascaded)} system(s), "
f"{len(cascade_errors)} error(s)"
)
after = await get_flags()
# The audit entry is a record OF the write, never a precondition for it.
# audit_log lives in the same Firestore that just accepted the flag write,
# so a failure here is nearly always transient — losing the flip (or 500ing
# a route that already succeeded, which invites a retry that flips it back)
# would be a far worse outcome than an unrecorded flip that is still in the
# service log above.
try:
# Deferred import: app.internal.audit pulls in firestore, and this
# module is imported from router module scope.
from app.internal import audit
actor_uid, actor_email = actor or ("unknown", "")
changed = {
k: {"from": before.get(k), "to": after.get(k)}
for k in clean
if before.get(k) != after.get(k)
}
await audit.write_audit(
actor_uid=actor_uid,
actor_email=actor_email,
action="feature_flags.update",
details={
"requested": clean,
"changed": changed,
"before": before,
"after": after,
"cascade": cascade,
"cascaded_systems": cascaded,
"cascade_errors": cascade_errors,
},
)
except Exception as e:
logger.error(f"Feature flags: audit write failed ({e}) — flag change stands")
return after
async def resolve_flags(system_id: str | None):
+39 -5
View File
@@ -1,7 +1,7 @@
import asyncio
from datetime import datetime, timezone, timedelta
from fastapi import APIRouter, Depends, Query
from app.internal.auth import require_admin_token
from app.internal.auth import require_admin_token, require_agent_key_or_admin, describe_actor
from app.internal.feature_flags import get_flags, set_flags
from app.internal import firestore as fstore
from app.config import settings
@@ -25,20 +25,54 @@ router = APIRouter(prefix="/admin", tags=["admin"])
@router.get("/features")
async def get_feature_flags(_=Depends(require_admin_token)):
async def get_feature_flags(_=Depends(require_agent_key_or_admin)):
"""
Return the current AI feature flag state. Admin-only (SAAS_PLAN.md B2c) —
was previously any authenticated user via require_firebase_token, which
handed platform-wide AI configuration state to every signed-in viewer
regardless of org.
Also reachable with the agent service key (server-26#64) so the unattended
runbook can read the switch over HTTP instead of shelling into the
container. Note this is require_agent_key_or_admin, NOT the Discord bot's
service key — see internal/auth.py.
"""
return await get_flags()
@router.put("/features")
async def update_feature_flags(body: dict, _=Depends(require_admin_token)):
"""Update one or more AI feature flags. Admin only."""
return await set_flags(body)
async def update_feature_flags(
body: dict,
cascade: bool = Query(
False,
description=(
"Also clear per-system ai_flags overrides for the keys being set, "
"so the flip applies to every radio system."
),
),
principal: dict = Depends(require_agent_key_or_admin),
):
"""Update one or more AI feature flags. Admin or agent service key.
``cascade`` defaults to **False**, deliberately.
The tempting default is True: feature_flags.resolve_flags lets a
system-level False beat a global True, so turning AI back ON globally can
half-apply and leave a system dark, and cascade-by-default would make every
flip total. That reasoning holds only if per-system ai_flags are set
exclusively by hand. They are not — PUT /systems/{system_id}/ai-flags
(routers/systems.py) is a real admin route and drb-frontend's AiFlagsPanel
(app/systems/page.tsx) is a real toggle in the UI. So an override is a
deliberate operator decision that is visible in the interface, and
cascading by default would silently erase it on the next unrelated global
flip, with the operator's own UI still showing what they set until reload.
Silently destroying operator intent is the worse failure, so the caller
says when it means "everywhere": the runbook passes cascade=true on the
shutoff, and the admin UI (which does not pass it) keeps its per-system
overrides.
"""
return await set_flags(body, actor=describe_actor(principal), cascade=cascade)
@router.get("/debug/correlation")
@@ -0,0 +1,296 @@
"""
server-26#64 — a headless, attributable, total AI-flag flip.
Three things are held here:
* ``require_agent_key_or_admin`` is a DISTINCT principal. It takes the agent
service key or a Firebase admin token and refuses the Discord bot's
``service_key``, so an audit entry can name who flipped the switch.
* ``set_flags`` writes an ``audit_log`` entry carrying before/after values,
and an audit failure can neither lose the flag write nor 500 the route.
* ``cascade=True`` clears per-system ``ai_flags`` overrides for the keys
being set, so a flip cannot half-apply — discovered by scanning for
documents that carry the map, never a hardcoded system-id list.
The dependency is exercised directly rather than through TestClient: these are
assertions about the credential check, and routing them through the ASGI stack
would only add ways for the test to pass for the wrong reason.
"""
import pytest
from unittest.mock import AsyncMock, patch
from fastapi import HTTPException
from fastapi.security import HTTPAuthorizationCredentials
from app.config import settings
from app.internal import auth, feature_flags
from app.routers import admin
AGENT_KEY = "agent-key-aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
BOT_KEY = "bot-key-bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"
def _creds(token: str) -> HTTPAuthorizationCredentials:
return HTTPAuthorizationCredentials(scheme="Bearer", credentials=token)
@pytest.fixture
def keys(monkeypatch):
"""Both keys configured and different — the production shape."""
monkeypatch.setattr(settings, "agent_service_key", AGENT_KEY, raising=False)
monkeypatch.setattr(settings, "service_key", BOT_KEY, raising=False)
@pytest.fixture(autouse=True)
def _clear_flag_cache():
"""feature_flags keeps module-level cache state; don't leak it across tests."""
feature_flags._cache = {}
feature_flags._cache_ts = 0.0
yield
feature_flags._cache = {}
feature_flags._cache_ts = 0.0
# ── Item 1: the credential ────────────────────────────────────────────────────
@pytest.mark.asyncio
async def test_agent_key_is_accepted_and_identifies_itself(keys):
principal = await auth.require_agent_key_or_admin(_creds(AGENT_KEY))
assert principal["principal"] == "agent"
# The caller must be able to tell the agent from a human admin, or the
# audit entry in item 3 cannot name the actor.
assert auth.describe_actor(principal) == (
auth.AGENT_PRINCIPAL_UID, auth.AGENT_PRINCIPAL_EMAIL,
)
@pytest.mark.asyncio
async def test_discord_bot_service_key_is_rejected(keys):
"""The whole point of a second key: the bot's key must not open this door.
It falls through to the Firebase branch and fails there, so the bot gets a
401 rather than an unattributable flag flip.
"""
with patch.object(auth.firebase_auth, "verify_id_token", side_effect=Exception("not a token")):
with pytest.raises(HTTPException) as exc:
await auth.require_agent_key_or_admin(_creds(BOT_KEY))
assert exc.value.status_code == 401
@pytest.mark.asyncio
async def test_unset_agent_key_cannot_be_bypassed(monkeypatch):
"""An unconfigured key must match nothing — especially not an empty string.
``secrets.compare_digest("", "")`` is a match, so the guard has to be on
the key being configured, not on a ``or ""`` fallback.
"""
monkeypatch.setattr(settings, "agent_service_key", None, raising=False)
with patch.object(auth.firebase_auth, "verify_id_token", side_effect=Exception("not a token")):
for token in ("", " ", "None", "null", AGENT_KEY):
with pytest.raises(HTTPException) as exc:
await auth.require_agent_key_or_admin(_creds(token))
assert exc.value.status_code == 401, token
@pytest.mark.asyncio
async def test_empty_string_agent_key_cannot_be_bypassed(monkeypatch):
"""Same guarantee for a key set to "" by an empty env var."""
monkeypatch.setattr(settings, "agent_service_key", "", raising=False)
with patch.object(auth.firebase_auth, "verify_id_token", side_effect=Exception("not a token")):
with pytest.raises(HTTPException) as exc:
await auth.require_agent_key_or_admin(_creds(""))
assert exc.value.status_code == 401
@pytest.mark.asyncio
async def test_firebase_admin_token_still_works(keys):
decoded = {"uid": "u-1", "email": "admin@example.com", "role": "admin"}
with patch.object(auth.firebase_auth, "verify_id_token", return_value=decoded):
principal = await auth.require_agent_key_or_admin(_creds("firebase-id-token"))
assert principal == decoded
assert auth.describe_actor(principal) == ("u-1", "admin@example.com")
@pytest.mark.asyncio
async def test_non_admin_firebase_token_is_forbidden(keys):
decoded = {"uid": "u-2", "email": "viewer@example.com", "role": "viewer"}
with patch.object(auth.firebase_auth, "verify_id_token", return_value=decoded):
with pytest.raises(HTTPException) as exc:
await auth.require_agent_key_or_admin(_creds("firebase-id-token"))
assert exc.value.status_code == 403
@pytest.mark.asyncio
async def test_missing_credentials_is_401(keys):
with pytest.raises(HTTPException) as exc:
await auth.require_agent_key_or_admin(None)
assert exc.value.status_code == 401
def test_features_routes_use_the_agent_dependency_and_others_do_not():
"""Guards the wiring: only /admin/features moved off require_admin_token."""
def deps(path, method):
for r in admin.router.routes:
if r.path == path and method in r.methods:
return {d.call for d in r.dependant.dependencies}
raise AssertionError(f"no route {method} {path}")
assert auth.require_agent_key_or_admin in deps("/admin/features", "GET")
assert auth.require_agent_key_or_admin in deps("/admin/features", "PUT")
assert auth.require_admin_token in deps("/admin/audit", "GET")
assert auth.require_admin_token in deps("/admin/debug/correlation", "GET")
# ── Items 3 and 4: set_flags audits, and cascades on request ──────────────────
def _fstore_mock(stored: dict, systems: list[dict], updates_sink: list):
"""A Firestore stand-in for feature_flags: one config doc, N system docs."""
mock = AsyncMock()
async def doc_get(collection, doc_id):
return dict(stored) if collection == "config" else None
async def doc_set(collection, doc_id, data, merge=True):
stored.update(data)
async def collection_list(collection, **filters):
return systems if collection == "systems" else []
async def doc_update(collection, doc_id, data):
updates_sink.append((collection, doc_id, data))
mock.doc_get = AsyncMock(side_effect=doc_get)
mock.doc_set = AsyncMock(side_effect=doc_set)
mock.collection_list = AsyncMock(side_effect=collection_list)
mock.doc_update = AsyncMock(side_effect=doc_update)
return mock
def _systems():
return [
# Two systems carry overrides today; the ids are irrelevant to the
# helper and must stay that way.
{"system_id": "sys-a", "ai_flags": {"stt_enabled": False, "correlation_enabled": False}},
{"system_id": "sys-b", "ai_flags": {"stt_enabled": False}},
# Carries the map but not the key being flipped — must be left alone.
{"system_id": "sys-c", "ai_flags": {"summaries_enabled": False}},
# No overrides at all: already inherits, nothing to cascade to.
{"system_id": "sys-d"},
{"system_id": "sys-e", "ai_flags": {}},
]
async def _run_set_flags(updates, *, stored=None, systems=None, cascade=False, actor=None,
audit_side_effect=None):
stored = stored if stored is not None else {"stt_enabled": True, "correlation_enabled": True}
systems = systems if systems is not None else _systems()
updates_sink: list = []
audit_mock = AsyncMock(side_effect=audit_side_effect)
with patch.object(feature_flags, "fstore", _fstore_mock(stored, systems, updates_sink)), \
patch("app.internal.audit.write_audit", new=audit_mock):
result = await feature_flags.set_flags(updates, actor=actor, cascade=cascade)
return result, stored, updates_sink, audit_mock
@pytest.mark.asyncio
async def test_set_flags_is_backward_compatible_without_actor_or_cascade():
"""Existing call shape — set_flags({...}) — must keep working."""
result, stored, updates_sink, audit_mock = await _run_set_flags({"stt_enabled": False})
assert result["stt_enabled"] is False
assert stored["stt_enabled"] is False
assert updates_sink == [] # no cascade unless asked
assert audit_mock.await_count == 1 # but still audited
@pytest.mark.asyncio
async def test_audit_records_before_and_after_and_the_actor():
_, _, _, audit_mock = await _run_set_flags(
{"stt_enabled": False},
actor=(auth.AGENT_PRINCIPAL_UID, auth.AGENT_PRINCIPAL_EMAIL),
)
kwargs = audit_mock.await_args.kwargs
assert kwargs["action"] == "feature_flags.update"
assert kwargs["actor_uid"] == auth.AGENT_PRINCIPAL_UID
assert kwargs["actor_email"] == auth.AGENT_PRINCIPAL_EMAIL
details = kwargs["details"]
assert details["changed"]["stt_enabled"] == {"from": True, "to": False}
assert details["before"]["stt_enabled"] is True
assert details["after"]["stt_enabled"] is False
@pytest.mark.asyncio
async def test_audit_failure_neither_loses_the_write_nor_raises():
"""audit_log is a record OF the write, never a precondition for it."""
result, stored, _, audit_mock = await _run_set_flags(
{"stt_enabled": False},
audit_side_effect=RuntimeError("firestore down"),
)
assert audit_mock.await_count == 1
assert stored["stt_enabled"] is False # flag write survived
assert result["stt_enabled"] is False # and the route returns normally
@pytest.mark.asyncio
async def test_cascade_clears_matching_system_overrides_at_both_levels():
result, stored, updates_sink, audit_mock = await _run_set_flags(
{"stt_enabled": True}, cascade=True,
)
# Global level.
assert stored["stt_enabled"] is True
assert result["stt_enabled"] is True
# System level: only the two documents whose ai_flags carry stt_enabled.
written = {sid: data["ai_flags"] for _, sid, data in updates_sink}
assert set(written) == {"sys-a", "sys-b"}
# The flipped key is removed so the system inherits; unrelated overrides stay.
assert written["sys-a"] == {"correlation_enabled": False}
assert written["sys-b"] == {}
# And the cascade is recorded, per system, in the audit entry.
cascaded = audit_mock.await_args.kwargs["details"]["cascaded_systems"]
assert {c["system_id"] for c in cascaded} == {"sys-a", "sys-b"}
assert cascaded[0]["cleared_overrides"] == {"stt_enabled": False}
@pytest.mark.asyncio
async def test_cascade_finds_systems_by_shape_not_by_hardcoded_id():
"""A newly added system carrying an override must not defeat a flip."""
systems = _systems() + [{"system_id": "sys-new", "ai_flags": {"stt_enabled": False}}]
_, _, updates_sink, _ = await _run_set_flags(
{"stt_enabled": True}, systems=systems, cascade=True,
)
assert "sys-new" in {sid for _, sid, _ in updates_sink}
@pytest.mark.asyncio
async def test_cascade_off_leaves_every_system_override_intact():
"""The default path must not silently erase a deliberate per-system value."""
_, _, updates_sink, _ = await _run_set_flags({"stt_enabled": True}, cascade=False)
assert updates_sink == []
@pytest.mark.asyncio
async def test_cascade_error_on_one_system_does_not_stop_the_others():
systems = _systems()
stored = {"stt_enabled": True, "correlation_enabled": True}
updates_sink: list = []
fs = _fstore_mock(stored, systems, updates_sink)
real_update = fs.doc_update.side_effect
async def flaky(collection, doc_id, data):
if doc_id == "sys-a":
raise RuntimeError("write conflict")
return await real_update(collection, doc_id, data)
fs.doc_update = AsyncMock(side_effect=flaky)
audit_mock = AsyncMock()
with patch.object(feature_flags, "fstore", fs), \
patch("app.internal.audit.write_audit", new=audit_mock):
await feature_flags.set_flags({"stt_enabled": True}, cascade=True)
assert [sid for _, sid, _ in updates_sink] == ["sys-b"]
details = audit_mock.await_args.kwargs["details"]
assert [e["system_id"] for e in details["cascade_errors"]] == ["sys-a"]
@pytest.mark.asyncio
async def test_unrecognised_keys_still_raise():
with pytest.raises(ValueError):
await _run_set_flags({"not_a_flag": True})
@@ -30,6 +30,14 @@ GEMINI_API_KEY={{ vault_gemini_api_key }}
SERVICE_KEY={{ vault_service_key }}
ENROLLMENT_TOKEN={{ vault_enrollment_token }}
# Agent/automation key for the unattended work session's headless routes
# (GET/PUT /admin/features). MUST NOT equal vault_service_key: that one is the
# Discord bot's, and one shared value would make the bot and the agent the same
# unattributable principal in audit_log (server-26#64). default('') so a vault
# that predates this key still templates instead of failing the play; blank
# just leaves the agent path closed.
AGENT_SERVICE_KEY={{ vault_agent_service_key | default('') }}
# Bare domain, not app.<domain>: the frontend is served on {{ domain }} itself
# (see Caddyfile.j2 — only api. and the bare name have DNS records). This said
# app.{{ domain }} while the browser origin was https://{{ domain }}, so every
+5 -1
View File
@@ -22,7 +22,11 @@ vault_mqtt_c2_pass: "CHANGE_ME"
vault_mqtt_dynsec_admin_pass: "CHANGE_ME" # openssl rand -hex 32 — must be >=12 chars, plugin-enforced minimum
# ── C2 Core ───────────────────────────────────────────────────────────────────
vault_service_key: "" # openssl rand -hex 32
vault_service_key: "" # openssl rand -hex 32 — the Discord bot's key
# The work-session agent's own key for GET/PUT /admin/features. Generate a
# SEPARATE value — never a copy of vault_service_key, or a flag flip cannot be
# attributed to the agent vs the bot (server-26#64).
vault_agent_service_key: "" # openssl rand -hex 32
vault_enrollment_token: "" # openssl rand -hex 32 — fleet-wide, shared by every node's POST /nodes/enroll
vault_openai_api_key: ""
vault_google_maps_api_key: ""