fix(dashboard): close the indirection holes and the unnamed-profile 400s

Round 2 of the REST profile-scope pass. A call-graph audit (AST over every
router plus their imported helpers, following functools.partial bindings,
closures and callbacks handed to to_thread/_spawn_job) found the writes a
signature grep cannot see, and the SPA/Desktop callers that never named a
profile at all.

Handlers that reached a config write through indirection:
* PUT /api/dashboard/plugin-providers wrote memory.provider + context.engine
  through functools.partial(_write_config_value, ...) — the SAME key
  PUT /api/memory/provider scopes — into the launch profile's config.yaml.
* POST /api/local-models/quickstart is `activate` plus a download; only
  `activate` had been scoped, so _set_runtime_enabled/_assign_default landed
  in the launch profile.
* POST /api/local-models/runtime/install regenerates launch presets under
  get_hermes_home().
* GET /api/model/recommended-default lazily PERSISTS discovered custom-provider
  models via build_models_payload -> _save_discovered_models_to_config.

Policy gaps:
* POST /api/ops/hooks is now gated like DELETE: writing an arbitrary command
  into `hooks:` and, with approve, into the consent allowlist is strictly more
  privileged than removing one.
* POST /api/sessions/prune (non-dry-run), DELETE /api/sessions/empty and
  POST /api/sessions/bulk-delete join the same destructive class.
* PUT /api/memory/provider ran its readiness check OUTSIDE the scope, so it
  judged the launch profile and could write a broken setting into another.

Callers:
* /api/credentials was missing from the SPA's PROFILE_SCOPED_PREFIXES, so the
  credential-pool delete button 400'd unconditionally; so were
  /api/dashboard/plugin-providers, /api/local-models and the model route above.
* The management scope was empty until the switcher resolved, so every
  destructive route 400'd in working UI on any host with a second profile
  directory. The backend now injects the profile it itself serves
  (__HERMES_DASHBOARD_PROFILE__) — and only when that name provably resolves
  back to its own home, so the fallback can never retarget another profile.
* authedFetch went around the scope entirely; the Desktop /api/ops callers
  (doctor, security-audit, backup, debug-share) sent no profile while Electron
  pins that family to the shared primary backend.
This commit is contained in:
teknium1
2026-09-21 09:26:19 -07:00
committed by Teknium
parent 6a3dcc39ab
commit 2e497bead8
11 changed files with 155 additions and 27 deletions

View File

@@ -241,18 +241,29 @@ export function getGhAuthStatus(refresh = false): Promise<{ available: boolean;
// audit` / `hermes backup` / `hermes debug share` and the dashboard System
// page). All except debug share are spawn-based background actions tailed via
// getActionStatus().
//
// Every one carries the ambient profile: Electron pins the whole /api/ops
// family to the shared primary backend (connection-config's
// LOCAL_PRIMARY_SCOPED_ROUTES), so an unprofiled call acts on that backend's
// LAUNCH profile — and debug share uploads a home's logs and config.
// ---------------------------------------------------------------------------
export function runDoctor(): Promise<ActionResponse> {
return hermesApi<ActionResponse>({ path: '/api/ops/doctor', method: 'POST', body: {} })
return hermesApi<ActionResponse>({ ...profileScoped(), path: '/api/ops/doctor', method: 'POST', body: {} })
}
export function runSecurityAudit(): Promise<ActionResponse> {
return hermesApi<ActionResponse>({ path: '/api/ops/security-audit', method: 'POST', body: {} })
return hermesApi<ActionResponse>({
...profileScoped(),
path: '/api/ops/security-audit',
method: 'POST',
body: {}
})
}
export function runBackup(): Promise<ActionResponse & { archive?: string }> {
return hermesApi<ActionResponse & { archive?: string }>({
...profileScoped(),
path: '/api/ops/backup',
method: 'POST',
body: {}
@@ -261,6 +272,7 @@ export function runBackup(): Promise<ActionResponse & { archive?: string }> {
export function runDebugShare(): Promise<DebugShareResponse> {
return hermesApi<DebugShareResponse>({
...profileScoped(),
path: '/api/ops/debug-share',
method: 'POST',
body: {},

View File

@@ -58,16 +58,19 @@ async def config_scoped_to_thread(profile: Optional[str], fn: Callable[[], Any])
def destructive_profile(profile: Optional[str], route: str) -> Optional[str]:
"""The profile a DESTRUCTIVE route acts on, or 400 when it is ambiguous.
"""The profile a DESTRUCTIVE or PRIVILEGED route acts on, or 400 when it is ambiguous.
One backend serves every profile, so an omitted ``profile`` on a route that
deletes or overwrites profile-owned data is not a default — it silently meant
One backend serves every profile, so an omitted ``profile`` on a route that deletes,
overwrites or privileges profile-owned data is not a default — it silently meant
"whichever home this process launched with". Named profile: honoured. Omitted:
rejected as soon as the process hosts more than one profile
(``is_multiplex_active()``, decided once at boot by
``activate_multi_profile_hosting_eagerly``). A genuinely single-profile host has
nothing to confuse, so there an omitted profile keeps meaning the launch profile
and `curl` against a plain ``hermes serve`` is unchanged.
"Privileged" is the same class as "destructive": arming an auto-approved shell hook
in the wrong profile is at least as bad as removing one from it.
"""
if (profile or "").strip():
return profile

View File

@@ -276,15 +276,23 @@ async def delete_agent_plugin(request: Request, name: str):
@router.put("/api/dashboard/plugin-providers")
async def put_plugin_providers(request: Request, body: _PluginProvidersPutBody):
async def put_plugin_providers(request: Request, body: _PluginProvidersPutBody,
profile: Optional[str] = None):
"""Persist memory provider / context engine selection (writes config.yaml)."""
_require_token(request)
from hermes_cli.plugins_cmd import _save_context_engine, _save_memory_provider
def _run():
with _CONFIG_MUTATION_LOCK:
# ``_save_memory_provider``/``_save_context_engine`` are functools.partial over
# ``_write_config_value`` -> save_config: the write is one hop away and lands in
# whatever home the scope names. Unlike plugin INSTALLATION (host venv, pinned to the
# serving profile), these two keys are per-profile settings — `memory.provider` is the
# same key PUT /api/memory/provider scopes — so they follow ``?profile=``.
with _config_profile_scope(profile), _CONFIG_MUTATION_LOCK:
if body.memory_provider is not None:
memory_provider = _normalize_memory_provider_name(body.memory_provider)
# Readiness resolves through load_config(); inside the scope so the answer is
# about the profile being written, not the launch profile.
_require_memory_provider_ready(memory_provider)
_save_memory_provider(memory_provider)
if body.context_engine is not None:

View File

@@ -657,7 +657,7 @@ def _restart_on_new_tag(job: Dict[str, Any], tag: str, previous: list) -> bool:
@router.post("/api/local-models/runtime/install")
async def local_models_runtime_install(body: RuntimeInstallBody):
async def local_models_runtime_install(body: RuntimeInstallBody, profile: Optional[str] = None):
tag, backend = _runtime_target(body.backend)
plan = _resolve_assets_or_400(tag, backend)
job = _job("runtime-install", f"llama.cpp {tag} ({backend})")
@@ -665,9 +665,13 @@ async def local_models_runtime_install(body: RuntimeInstallBody):
def _run():
previous = binaries.installed_tags()
_step(job, "downloading", f"Fetching {len(plan.assets)} package(s) for {backend}")
binaries.ensure_runtime_installed(tag, backend, progress=_runtime_progress_hook(job))
# Restart failure is logged only: the new build is installed either way and the next boot serves it.
restarted = _quiet(lambda: _restart_on_new_tag(job, tag, previous), False, warn="post-update restart skipped: %s")
# The engine binaries are machine-global, but ensure_local_runtime also regenerates the
# launch presets under get_hermes_home() — scope so those land in the named profile.
with _config_profile_scope(profile):
binaries.ensure_runtime_installed(tag, backend, progress=_runtime_progress_hook(job))
# Restart failure is logged only: the new build is installed either way and the next boot serves it.
restarted = _quiet(lambda: _restart_on_new_tag(job, tag, previous), False,
warn="post-update restart skipped: %s")
# N-1 retention, only after the new tag verified: keep it + the newest previous build as the rollback pin target.
_quiet(lambda: binaries.prune_old_tags([tag] + [t for t in previous if t != tag][:1]), None,
warn="runtime prune skipped: %s")
@@ -748,7 +752,7 @@ def _quickstart_target(body: QuickstartBody, budget):
@router.post("/api/local-models/quickstart")
async def local_models_quickstart(body: QuickstartBody):
async def local_models_quickstart(body: QuickstartBody, profile: Optional[str] = None):
"""One job: install the runtime (if missing), download this machine's build of the recommended model (if
missing), make it the default. Each leg uses the same code as the individual setup routes.
Preflight rejects (no automatic recommendation or no servable choice) fail the POST
@@ -775,11 +779,14 @@ async def local_models_quickstart(body: QuickstartBody):
job["done_bytes"] = 0
job["total_bytes"] = download_bytes
_run_download_plan(job, download_plan, entry.display_name)
# Activate: same sequence as /activate's job body.
_ensure_server(job, _set_runtime_enabled(True), variant.model_id,
fail_detail="The local server could not start — open Local Models for details",
skip_msg="quickstart rescan check skipped")
_assign_default(job, variant.model_id)
# Activate: same sequence as /activate's job body, and the same scope. Quickstart IS
# `activate` plus a download: _set_runtime_enabled and _assign_default both reach
# save_config, so without this the config.yaml write lands in the launch profile.
with _config_profile_scope(profile):
_ensure_server(job, _set_runtime_enabled(True), variant.model_id,
fail_detail="The local server could not start — open Local Models for details",
skip_msg="quickstart rescan check skipped")
_assign_default(job, variant.model_id)
_finish(job, f"{entry.display_name} is ready — new chats use it")
_spawn_job(job, "lr-quickstart", _run, fail_msg="quickstart failed: %s", on_exit=_QUICKSTART_LOCK.release)

View File

@@ -137,7 +137,7 @@ def _nous_recommended_default() -> dict:
@router.get("/api/model/recommended-default")
def get_recommended_default_model(provider: str = ""):
def get_recommended_default_model(provider: str = "", profile: Optional[str] = None):
"""Recommended default model for a freshly-authenticated provider, mirroring
``hermes model``'s curation so GUI onboarding lands on a sensible default.
Nous honors the user's free/paid tier. Any other provider gets the preferred
@@ -159,7 +159,10 @@ def get_recommended_default_model(provider: str = ""):
from hermes_cli.inventory import build_models_payload, load_picker_context
from hermes_cli.models import pick_silent_default_model
payload = build_models_payload(load_picker_context())
# build_models_payload -> list_authenticated_providers -> _save_discovered_models_to_config:
# this GET lazily PERSISTS discovered custom-provider models, so it needs the scope too.
with _config_profile_scope(profile):
payload = build_models_payload(load_picker_context())
for row in payload.get("providers", []):
if str(row.get("slug", "")).lower() == slug:
models = [str(m) for m in (row.get("models") or [])]

View File

@@ -499,8 +499,12 @@ async def set_memory_provider(body: MemoryProviderSelect, profile: Optional[str]
provider = _normalize_memory_provider_name(body.provider)
def _run():
_require_memory_provider_ready(provider)
# Readiness resolves through load_config()/_discover_memory_provider_statuses(), so it
# MUST run inside the scope: outside it a provider configured only in the target profile
# reads as "not ready" (refused) and one configured only in the launch profile reads as
# ready and gets written into the target as a broken setting.
with config_write_scope(profile):
_require_memory_provider_ready(provider)
cfg = load_config()
if not isinstance(cfg.get("memory"), dict):
cfg["memory"] = {}
@@ -708,6 +712,11 @@ async def create_hook(body: HookCreate, profile: Optional[str] = None):
"""
from agent import shell_hooks
# Creating an auto-approved shell hook is strictly more privileged than removing one:
# it writes an arbitrary command into `hooks:` and, with `approve`, into that profile's
# consent allowlist. Same rule as DELETE — an unnamed profile is refused while several
# are served rather than silently arming the launch profile.
profile = destructive_profile(profile, "POST /api/ops/hooks")
event, command = _hook_body_fields(body)
valid_hooks = None
with contextlib.suppress(Exception):

View File

@@ -22,7 +22,7 @@ from hermes_cli.web_server_gateway import _strip_session_list_rows
from hermes_cli.web_server_sessions import _maybe_auto_archive_for_profile, _session_latest_descendant
from hermes_cli.web_models import (
BulkDeleteSessions, SessionImport, SessionOwnerBackfill, SessionPrune, SessionRename)
from hermes_cli.web_routers._common import log as _log, http_failure
from hermes_cli.web_routers._common import log as _log, destructive_profile, http_failure
from hermes_state import is_malformed_db_error
from hermes_state_errors import is_transient_sqlite_error
@@ -409,8 +409,9 @@ async def bulk_delete_sessions_endpoint(body: BulkDeleteSessions):
# Hard cap so a runaway selection can't lock the writer for long.
if len(body.ids) > 500:
raise HTTPException(status_code=400, detail="ids must contain at most 500 entries")
profile = destructive_profile(body.profile, "POST /api/sessions/bulk-delete")
deleted = await asyncio.to_thread(
_with_db, body.profile, lambda db: db.delete_sessions(body.ids), read_only=False)
_with_db, profile, lambda db: db.delete_sessions(body.ids), read_only=False)
return {"ok": True, "deleted": deleted}
@@ -458,7 +459,8 @@ async def delete_empty_sessions_endpoint(profile: Optional[str] = None):
parents are orphaned, not cascade-deleted. See #95868.
"""
deleted = await asyncio.to_thread(
_with_db, profile, lambda db: db.delete_empty_sessions(), read_only=False)
_with_db, destructive_profile(profile, "DELETE /api/sessions/empty"),
lambda db: db.delete_empty_sessions(), read_only=False)
return {"ok": True, "deleted": deleted}
@@ -760,6 +762,11 @@ async def export_session_endpoint(session_id: str, profile: Optional[str] = None
@manage_router.post("/api/sessions/prune")
async def prune_sessions_endpoint(body: SessionPrune):
"""Delete ended sessions matching filters without blocking the event loop."""
if not body.dry_run:
# Same destructive rule as the rest of the family; a dry run deletes nothing, so it
# keeps working unnamed (it is the preview the confirm dialog reads).
body = body.model_copy(update={
"profile": destructive_profile(body.profile, "POST /api/sessions/prune")})
return await asyncio.to_thread(_prune_sessions, body)

View File

@@ -156,12 +156,19 @@ def mount_spa(application: FastAPI):
# Launcher-preselected profile (``--open-profile``): the SPA's fallback scope when the URL
# omits ``?profile=`` (#73085). ``</`` escaped so a hostile name cannot close the script tag.
initial_profile_js = json.dumps(str(getattr(application.state, "initial_profile", "") or "")).replace("</", "<\\/")
# This backend's OWN profile name (empty when it cannot be named unambiguously). The SPA
# falls back to it when neither the URL nor --open-profile names one, so requests carry an
# explicit scope from the first paint: destructive routes 400 on an unnamed profile as soon
# as the host serves more than one, and the switcher shows the same profile it writes.
from hermes_cli.web_server_profiles import serving_profile_name as _serving_profile_name
serving_profile_js = json.dumps(_serving_profile_name()).replace("</", "<\\/")
bootstrap_script = (
f"<script>{token_js}"
f"window.__HERMES_DASHBOARD_EMBEDDED_CHAT__={chat_js};"
f'window.__HERMES_BASE_PATH__="{prefix}";'
f"window.__HERMES_AUTH_REQUIRED__={'true' if gated else 'false'};"
f"window.__HERMES_INITIAL_PROFILE__={initial_profile_js};"
f"window.__HERMES_DASHBOARD_PROFILE__={serving_profile_js};"
f"</script>"
)
if prefix:

View File

@@ -46,6 +46,28 @@ def _hermes_home_scope(path) -> Any:
reset_hermes_home_override(token)
def serving_profile_name() -> str:
"""This process's OWN profile name — but only when that name provably resolves back
to the process home.
The dashboard SPA needs an explicit scope for requests it fires before the profile
switcher has resolved: a destructive route now 400s on an unnamed profile as soon as
the host serves more than one, and "" would otherwise mean "whichever home this
process launched with" anyway. Naming it is only safe if the name cannot resolve
ELSEWHERE, so a custom HERMES_HOME outside ``profiles/`` (``get_active_profile_name()``
answers ``"custom"``) returns "" and keeps the old unnamed behaviour rather than
risking a wrong-profile write.
"""
from hermes_cli import profiles as profiles_mod
try:
name = (profiles_mod.get_active_profile_name() or "").strip()
if not name or name == "custom":
return ""
return name if profiles_mod.profile_matches_home(name, get_process_hermes_home()) else ""
except Exception:
return ""
def _is_other_profile(profile: Optional[str]) -> bool:
"""True when ``profile`` names a profile other than this process's own."""
if _is_current_profile(profile):

View File

@@ -4,6 +4,8 @@ import {
type ModelOptionsResult,
} from "@hermes/shared";
import { dashboardServingProfile } from "./profile-bootstrap";
// The dashboard can be served either at the root of its host (e.g.
// https://kanban.tilos.com/) or under a URL prefix when reverse-proxied
// (e.g. https://mission-control.tilos.com/hermes/). The Python backend
@@ -64,8 +66,18 @@ export function setManagementProfile(name: string): void {
_managementProfile = (name || "").trim();
}
/**
* The profile every management call targets: the switcher's selection, or —
* before it has resolved / on a host with no switcher interaction — the profile
* this backend itself serves.
*
* The fallback is not a guess: the backend injects a name only when it provably
* resolves back to its own home, so it names exactly the home an unnamed request
* already reached. Without it the dashboard sends no `?profile=` at all and every
* destructive route 400s as soon as the host has a second profile directory.
*/
export function getManagementProfile(): string {
return _managementProfile;
return _managementProfile || dashboardServingProfile();
}
// Endpoint families that honor ?profile= on the backend (web_server.py
@@ -108,18 +120,33 @@ const PROFILE_SCOPED_PREFIXES = [
"/api/ops",
"/api/logs",
"/api/portal",
// Pool entries live in the profile's home, and DELETE /api/credentials/pool/{provider}/{index}
// is destructive — without this prefix the dashboard's remove button never named a profile and
// a multi-profile host refused it outright.
"/api/credentials",
// Not covered by "/api/dashboard/plugins": this one writes memory.provider + context.engine
// into the named profile's config.yaml (same key as PUT /api/memory/provider).
"/api/dashboard/plugin-providers",
// Model/runtime activation persists into config.yaml; the read routes ignore an extra param.
"/api/model/recommended-default",
"/api/local-models",
"/api/dashboard/theme",
"/api/dashboard/font",
"/api/dashboard/plugins",
];
// The dashboard's own profile when nothing else named one. The backend injects it only
// when it provably resolves back to the serving home, so this can never retarget another
// profile — it just says out loud what an unnamed request already meant. Without it every
// destructive route 400s on a host that merely HAS a second profile directory.
function withManagementProfile(url: string): string {
if (!_managementProfile) return url;
const scope = getManagementProfile();
if (!scope) return url;
if (url.includes("profile=")) return url; // explicit param wins
const path = url.split("?")[0];
if (!PROFILE_SCOPED_PREFIXES.some((p) => path.startsWith(p))) return url;
const sep = url.includes("?") ? "&" : "?";
return `${url}${sep}profile=${encodeURIComponent(_managementProfile)}`;
return `${url}${sep}profile=${encodeURIComponent(scope)}`;
}
export async function fetchJSON<T>(
@@ -281,6 +308,11 @@ export async function authedFetch(
url: string,
init?: RequestInit,
): Promise<Response> {
// Same management scope as fetchJSON: a binary endpoint under a profile-scoped
// family (``/api/ops/backup/download``) must read the SELECTED profile's archive,
// not the launch profile's, and an unprofiled back door beside a family that now
// 400s is exactly how the next hole gets in.
url = withManagementProfile(url);
const headers = new Headers(init?.headers);
const token = window.__HERMES_SESSION_TOKEN__;
if (token) {

View File

@@ -1,5 +1,6 @@
declare global {
interface Window {
__HERMES_DASHBOARD_PROFILE__?: string;
__HERMES_INITIAL_PROFILE__?: string;
}
}
@@ -9,12 +10,29 @@ export function dashboardInitialProfile(): string {
return window.__HERMES_INITIAL_PROFILE__ ?? "";
}
/**
* The profile this backend process itself serves, injected by the server and
* empty when it cannot be named unambiguously (custom HERMES_HOME).
*
* It is the LAST fallback for the management scope: without it the dashboard
* sends no `?profile=` at all, which a multi-profile host now refuses (400) on
* every destructive route. It is never a guess — the backend only emits a name
* that provably resolves back to its own home, so it targets exactly the home
* an unnamed request used to reach.
*/
export function dashboardServingProfile(): string {
if (typeof window === "undefined") return "";
return window.__HERMES_DASHBOARD_PROFILE__ ?? "";
}
export function initialProfileScope(
searchParams: URLSearchParams,
bootstrapProfile = dashboardInitialProfile(),
servingProfile = dashboardServingProfile(),
): string {
const urlProfile = searchParams.get("profile");
return urlProfile === null ? bootstrapProfile : urlProfile;
if (urlProfile !== null) return urlProfile;
return bootstrapProfile || servingProfile;
}
export function shouldAdoptActiveProfile(